News / #pricing Tag Pricing changes 338 articles archived under #pricing · RSS Sign in to follow Hacker News — AI on Front Page community 14d ago Advancing the price-performance frontier with GPT‑5.6 Article URL: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/ Comments URL: https://news.ycombinator.com/item?id=49112867 Points: 262 # Comments: 162 27 r/LocalLLaMA community 14d ago How close are we to local llama robotics for consumer price point? I'm guessing 3 years, what do you think? In other words: many of us will be able to afford a general purpose robot in 3 years to experiment with in the home. Cost roughly $5k? Probably small size, but hopefully still able to do the dishes and operate a vacuum.   submitted by… 13 OpenAI official-blog 14d ago Advancing the price-performance frontier with GPT-5.6 Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale. 35 Smol AI News news-outlet 15d ago not much happened today **OpenAI** aggressively cut prices for **GPT-5.6 Luna** by 80% and **Terra** by 20%, introducing a faster **Sol Fast** tier with up to 2.5× lower latency at double the price, improving agent workflow costs by roughly 10×. The **ARC-AGI-3** debate highlighted that the complete… 14 arXiv — Machine Learning research 15d ago Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes arXiv:2607.27188v1 Announce Type: new Abstract: Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation,… 30 Vercel — AI dev-tools 15d ago AI Gateway: GPT-5.6 pricing and speed updates On AI Gateway , GPT-5.6 Luna and GPT-5.6 Terra are now cheaper and GPT-5.6 Sol is faster. AI Gateway adds no markup on token pricing, so these changes reach you at the upstream rate. The changes apply to both short and long context pricing. Model Change Input: Short context (per… 4 TechCrunch — AI news-outlet 15d ago Mark Zuckerberg predicts that billions of people will have personal AI agents in five years As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the price. 19 Dwarkesh Podcast news-outlet 15d ago Why compute might get 10x+ more expensive in coming years If a human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot price. 19 r/LocalLLaMA community 16d ago dropped 4k on a spark, am I crazy? Saw that the Asus Ascent 1tb was going for $3,950 from a few sources, couldn't stop thinking about it, finally just went ahead and did it. Am I completely insane? Will I regret this? I can't imagine the price will go down any time soon, so it seems like a good idea and I… 34 r/LocalLLaMA community 16d ago Nvidia is expected to raise GeForce RTX GPU prices again by up to 30%   submitted by   /u/ab2377 [link]   [comments] 21 r/LocalLLaMA community 16d ago I've been tracking RTX 5090 prices across EU stores since March, it's up €1,061 and still climbing Been running a GPU price tracker ( https://www.pricesquirrel.com ) since March, covering 20+ EU stores, recently added RAM, SSDs and CPUs too. Every GPU tier has gotten cheaper since launch. The RTX 5090 has done the exact opposite. The data: The ASUS TUF Gaming RTX 5090 OC was… 34 TechCrunch — AI news-outlet 17d ago Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing Cursor says India is now its third-largest market globally and plans to expand local hiring and enterprise sales. 23 arXiv — Machine Learning research 17d ago Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs arXiv:2607.22786v1 Announce Type: new Abstract: In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. The financial time series are price series such… 38 arXiv — Machine Learning research 17d ago Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features arXiv:2607.23370v1 Announce Type: new Abstract: Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fat-tailed returns, non-stationary dynamics, and a price discovery process influenced by social discourse on Reddit… 23 r/LocalLLaMA community 17d ago First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.   submitted by   /u/fulgencio_batista [link]   [comments] 37 r/LocalLLaMA community 18d ago I want to run Kimi K3 at home, so I’m trying to make 2.8T-scale experimentation cheaper Hey r/LocalLLaMA , I’m a retired engineer with a background in distributed computing, currently running a 1-person startup. Like many people here, I’d love to experiment with 2T+ MoE models locally. The problem is that I don’t have an H100 cluster in my living room. So I’ve been… 12 Smol AI News news-outlet 18d ago not much happened today **Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with… 33 r/LocalLLaMA community 18d ago Will prices finally go down? I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments into AI data centers and had no use for then, had to rent them, same thing with… 38 r/LocalLLaMA community 18d ago Do people building local LLM rigs track RTX Ada/workstation card prices, or just consumer cards like the 5090? curious how people here approach buying high-end/workstation cards (RTX 6000 Ada, 5000 Ada, etc) for local LLM work, do you actively watch pricing/timing on these specifically, or is the consumer 5090 usually enough for most builds? also wondering if price alerts/tracking tools… 15 Simon Willison community 18d ago An Inside Look at the Relay Market Powering Token Resellers and Fraud An Inside Look at the Relay Market Powering Token Resellers and Fraud Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources. This looks to be mostly a thing in China.… 18 r/LocalLLaMA community 19d ago Any use cases for RTX PRO 4500? At its price point, PRO 4500 doesn’t offer as much raw performance due to its lower power draw at 300W. The 5090 can perform up to 60-70% in short spurts with 600W, but can also be undervolted down to 400W. Are there legitimate reasons other than 24/7 usage and lower power draw… 27 r/LocalLLaMA community 19d ago Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding? I am a long time iOS app developer. In the last year I have been using Cursor+Claude/others to assist with app development. I am concerned that the current low pricing will disappear eventually. I am pricing out a new laptop with the intention of using local models instead. New… 35 Latent.Space news-outlet 20d ago [AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) ain't nobody beats Anthropic at distilling Fable! 19 Ars Technica — AI news-outlet 20d ago Anthropic's Opus 5 is about token efficiency, not a capability leap Models are improving quickly, but the cheaper options are often good enough. 28 TechCrunch — AI news-outlet 20d ago Anthropic launches Opus 5 Opus 5 will be both cheaper and less restrictive than Fable, likely making it preferable in most use cases 27 r/MachineLearning community 20d ago I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P] Built an open-source AI coding agent that was 7%–75% cheaper than a cold "claude -p" run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: $6.83, 207 turns AutoDev Studio: ~$1.70 for the same bug The full benchmark (including… 14 arXiv — Machine Learning research 21d ago Adaptive Multi-Horizon Reinforcement Learning arXiv:2607.20656v1 Announce Type: new Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which… 30 r/LocalLLaMA community 21d ago How much are RTX PRO 6000s going for in your country/state? Hello guys, hoping you're doing fine! On the last 2-3 months, price of the RTX 6000 PRO seem to have gone insane. I will start on the price here on my country, Chile: RTX 6000 PRO Workstation Edition: 21382 USD post 19% tax. RTX 6000 PRO MaxQ Workstation Edition: 20669 USD post… 29 r/MachineLearning community 22d ago Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R] Ran 10 realistic product tasks (classification, RAG QA, multi turn conversation, an agentic plan then execute task, etc) against the live APIs of OpenAI, Anthropic, Gemini and Kimi, using each provider's cost optimized tier. Total cost spread was 10.6x despite published rates… 15 Latent.Space news-outlet 22d ago [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" a quiet day lets us highlight a new neolab win. 22 arXiv — NLP / Computation & Language research 22d ago PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference arXiv:2607.20327v1 Announce Type: new Abstract: Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware… 21 r/LocalLLaMA community 22d ago Got these baddies in the mail today (2X 3080 20GB) About to plug them in. Currently running a single 3090. I got these for less than the price of a single 3090. 24GB wasn't enough for my use case, so 40 GB should be an upgrade. Going to throw my 3090 on ebay very likely. I feel like it's a perfect time to sell since the prices… 10 arXiv — Machine Learning research 23d ago Exposure-Based Reinforcement Learning to Rank arXiv:2607.18689v1 Announce Type: new Abstract: Reinforcement learning (RL) methods for learning-to-rank (LTR) can optimize (almost) any ranking goal, e.g., from precision or discounted cumulative gain to fairness-of-exposure or ranking distillation. However, standard RL is… 26 arXiv — NLP / Computation & Language research 23d ago The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation arXiv:2607.19226v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the… 9 r/LocalLLaMA community 23d ago Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro Model Size Terminal-Bench 2.1 SWE-bench Multilingual SWE-Bench Pro (Public Dataset) DeepSWE SWE Atlas (Codebase QnA) Toolathlon Verified Laguna S 2.1 118B-A8B 70.2% 78.5% 59.4% 40.4% 46.2% 49.7% Finally the banger we've been waiting from Laguna. probably will be great for 64GB+… 33 Ars Technica — AI news-outlet 23d ago Google reveals faster and cheaper Gemini 3.6 Flash, says 3.5 Pro is still in testing There are new 3.6 and 3.5 models today, but Google is already training Gemini 4. 35 arXiv — Machine Learning research 24d ago RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce arXiv:2607.16230v1 Announce Type: new Abstract: Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversion. In practice, shipping cost is shaped not only by distance but also by destination demand… 37 arXiv — NLP / Computation & Language research 24d ago PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer arXiv:2607.17620v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) makes finetuning large language models cheaper by adding to each weight matrix a trainable low-rank update parameterized as the product of two matrices. These matrices are usually trained with Adam,… 23 Vercel — AI dev-tools 24d ago Vercel MCP now supports purchases Vercel MCP now supports purchasing Vercel products. You can: Upgrade your team to the Pro plan Add prepaid credits for v0 (requires a paid v0 plan) or AI Gateway Purchase the SIEM add-on (requires an Enterprise plan) Purchase and register a domain Vercel MCP quotes the price,… 32 Simon Willison community 24d ago Reverse-engineering is cheap now I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes. I think this is an interesting illustration of the impact of the reduced cost of writing code. Prior to agents, it was entirely possible to reverse-engineer home… 33 r/LocalLLaMA community 24d ago So what happened with OpenClaw? It had an insanely meteoritic rise. It felt like it was the only thing anyone had been talking about for months. Then just, everyone stopped talking about it. Usage based pricing inevitably came and it seems like it was killed over night. Competitors were also rushing to get… 22 r/LocalLLaMA community 24d ago DeepSeek v4 flash release version appears to have been activated on api. Open weights imminent? https://np.reddit.com/r/DeepSeek/s/skO7urrE2C DS4 sort of came and went from the spotlight. The consensus seemed to be that its most notable feature is its price, and then we got distracted by the next big releases. However, people seem to forget this was only the preview… 32 arXiv — Machine Learning research 25d ago Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching arXiv:2607.15516v1 Announce Type: new Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware… 33 Hacker News — AI on Front Page community 26d ago Better and Cheaper Than IPTV Article URL: https://github.com/stupside/castor Comments URL: https://news.ycombinator.com/item?id=48964015 Points: 226 # Comments: 63 25 r/LocalLLaMA community 26d ago How are y’all stomaching the “AI Boom” prices? I am in the middle of considering an upgrade to my Home Server. I want to get a decent GPU for LocalAI. I mean, I have an RTX 3060 TI 16GB, I know that is much more than most people have, but even though it has a lot of cram the bus width is really slowing it down - I get ~23… 26 r/LocalLLaMA community 26d ago What kind of dark magic is Deepseek using? I was taking a look at Kimi K3 scores on the Artificial analysis leaderboard and was quite baffled when I saw this chart. Granted, Deepseek has always been the king of price to performance, but this is still incredible. Is it just API subsidization or have they optimized their… 4 TechCrunch — AI news-outlet 27d ago AI-driven memory crunch jolts India’s smartphone market India's smartphone slowdown highlights how the AI boom is reshaping consumer electronics, from pricing and demand to corporate strategy. 4 r/LocalLLaMA community 27d ago Kimi K3 one-shotting a racing game Kimi k3 is the goat, especially in UI and web dev. It provided a result matching fable 5, 70% cheaper. Opencode and Kimi just became my default setup now. This is my first time actually coding outside Claude Code and Codex. Any reason to go back, or am I missing something?… 11 r/LocalLLaMA community 28d ago Any news about MiMo V3? It's truly the most exciting model. Kimi K3 and GLM 5.2 look impressive and will likely be quite helpful for distilling models. But for everyday work, paying for APIs, I would only use something that is an order of magnitude cheaper, like MiMo and DS.   submitted by  … 9 arXiv — NLP / Computation & Language research 28d ago Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel arXiv:2607.14431v1 Announce Type: new Abstract: We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later… 10 Page 2 of 7 · 338 articles ← Newer Older →