News / #pricing Tag Pricing changes 500 articles archived under #pricing · RSS Sign in to follow r/LocalLLaMA community 1mo ago DS4 Flash incoming price increase "we've been able to reproduce their current prices even on rented GPUs" https://preview.redd.it/kvfk26z2uwhh1.png?width=598&format=png&auto=webp&s=356a8793a6c31bc563d552aaa5a73112ced7372e https://preview.redd.it/xthbu87auwhh1.png?width=598&format=png&auto=webp&s=08f686fee339905a33609a0346f13163aedc2671 Hello, I've seen these tweets from dax… 16 arXiv — NLP / Computation & Language research 1mo ago AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either… 8 TechCrunch — AI news-outlet 1mo ago OpenAI’s new AI smart speaker will reportedly sell for between $300-$400 Additional details about OpenAI's mysterious new AI device make it sound like a pricey smart speaker. 35 TechCrunch — AI news-outlet 1mo ago OpenAI’s new AI smart speaker will reportedly sell for between $300 and $400 Additional details about OpenAI's mysterious new AI device make it sound like a pricey smart speaker. 17 r/LocalLLaMA community 1mo ago I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLM I'm the author, so discount the enthusiasm accordingly. This is an unaffiliated community port, not endorsed by the vLLM project, which it uses to verify its correctness. What started it: I love vLLM, but a vLLM install here is 9.1 GiB of virtualenv, and I wanted to embed… 36 r/LocalLLaMA community 1mo ago They almost catched up on Frontier performance, so now catching up on prices Users also report that the free version was significantly downgraded after the release of the new models this is very important for us when considering local hosting. A lot of people decided not to buy expensive hardware because DeepSeek’s prices made it very difficult to break… 25 r/LocalLLaMA community 1mo ago I get that AI labs need to make money, but zero-warning price spikes are a nightmare for production builds Seen a ton of posts today about the DeepSeek API price hike. Half the feed is doom-posting, the other half is explaining basic GPU economics. Honestly, I get the cost side. Sub-cent tokens were never gonna last forever. But what actually sucks is the zero-day notice. Dropping a… 23 Smol AI News news-outlet 1mo ago not much happened today **Meta's Muse Spark 1.2** rapidly rose to frontier-tier with top 5 ranking on Vals Index at **$0.69/test**, being **3x cheaper than Kimi** and **10x+ cheaper than Fable, Opus, and 5.6 Sol**. It achieved **gold-medal-level performance in five STEM Olympiads** with perfect theory… 20 arXiv — Machine Learning research 1mo ago Out-Of-The-Loop Multi-Fidelity Bayesian Optimization arXiv:2608.04113v1 Announce Type: new Abstract: Black-box optimization is a ubiquitous problem in science and engineering, often dealing with expensive objective functions with cheaper lower-fidelity proxies available. Multi-fidelity Bayesian optimization (MF-BO) is a principled… 28 arXiv — Machine Learning research 1mo ago Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints arXiv:2608.04669v1 Announce Type: new Abstract: Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long-run usage of every resource must… 20 arXiv — Machine Learning research 1mo ago Monsoon Mayhem to Market Waves: Forecasting Fisheries Resilience in Sri Lanka arXiv:2608.04023v1 Announce Type: cross Abstract: Sri Lanka's fisheries sector is important for jobs and food supply. Between 2019 and 2025, it faced several major problems at the same time, and how these events together affected fish production and prices is still not well… 37 r/LocalLLaMA community 1mo ago Get AI max+ 395 laptop or wait for rtx spark? So I can either pull the trigger on a 128gb AI max+ 395 laptop or wait for RTX Spark for LLMs. Maybe I get it now and the price of the spark is super high so it's a good purchase or maybe the Spark shocks everyone with a low price and I forever regret my purchasing decision.… 37 Hacker News — AI on Front Page community 1mo ago Beating GPT-5.6 Sol on retrieval with 100x cheaper open models Article URL: https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency Comments URL: https://news.ycombinator.com/item?id=49186762 Points: 226 # Comments: 47 29 TechCrunch — AI news-outlet 1mo ago Hark previews its browser use agent for completing tasks Hark claims that its browser use agent is faster and cheaper than competition. 9 arXiv — Machine Learning research 1mo ago Neural Networks with Local Converging Inputs for Efficient Options Pricing Models arXiv:2608.02778v1 Announce Type: new Abstract: We present a novel application of Neural Networks with Local Converging Inputs (NNLCI) to improve the efficiency of existing numerical methods for pricing multi-asset options. The most concise input format for NNLCI has been… 29 Vercel — AI dev-tools 1mo ago Full Sandbox egress firewall now available on Hobby plan All Vercel Sandbox firewall features are now available on the Hobby plan. This brings the same network isolation that protects production workloads to the free tier, giving Hobby builders control over exactly what leaves the sandbox while keeping secrets out of the code… 31 r/LocalLLaMA community 1mo ago SK hynix, In Collaboration With SanDisk, Unveils The New High Bandwidth Flash (HBF) Standard, Helping To Resolve AI Inference Bottlenecks, Targeting Up To 3TB/s Bandwidth Hopefully this would let us have faster local models....but it will probably be out of our price range.   submitted by   /u/giveen [link]   [comments] 12 arXiv — Machine Learning research 1mo ago FedChronos: Federated Fine-Tuning of Time-Series Foundation Models for Privacy-Preserving Commodity Price Forecasting arXiv:2608.01290v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) such as Chronos have demonstrated strong forecasting capabilities across domains, yet adapting them to institutionally fragmented settings, where data cannot be centralized due to regulatory,… 18 arXiv — Machine Learning research 1mo ago When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design arXiv:2608.01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full… 20 arXiv — NLP / Computation & Language research 1mo ago Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this… 5 Dwarkesh Podcast news-outlet 1mo ago Why smarter AI models could drive up compute prices 10x The end of cheap compute? 29 r/LocalLLaMA community 1mo ago 70-class VRAM stagnation been thinking about how the desktop 70-class has sat at 12GB for two generations now, 4070, 4070 super, 5070, all 12GB. the 1070 gave you 8GB back in 2016 and it felt generous for the price. ten years later the jump is... 4GB. and the thing is these chips arent even weak. the… 32 Smol AI News news-outlet 1mo ago Qwen 3.8 Max **Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with… 21 arXiv — NLP / Computation & Language research 1mo ago Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation arXiv:2607.29250v1 Announce Type: new Abstract: Small language models (SLMs) are attractive for agentic deployment due to low latency, reduced cost, and on-device privacy, yet they struggle with tool-use tasks where training data is scarce and noisy. Unlike larger models, SLMs… 13 r/MachineLearning community 1mo ago [D] Self-Promotion Thread Please post your personal projects, startups, product placements, collaboration needs, blogs etc. Please mention the payment and pricing requirements for products and services. Please do not post link shorteners, link aggregator websites , or auto-subscribe links. -- Any abuse… 7 r/LocalLLaMA community 1mo ago DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026 March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models available to run locally on <8K USD (us prices - just guestimating/not exact)… 19 r/LocalLLaMA community 1mo ago Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper https://preview.redd.it/h7zv5tb3tmgh1.png?width=2854&format=png&auto=webp&s=507380e8f862c18f10f7c5c84da9e8d1c59139b0 Deepseek's new flash model is unexpectedly cheap and high-performing across useful benchmarks. It's priced at $0.09 / $0.18 per 1M. Truly "intelligence too cheap… 23 r/LocalLLaMA community 1mo ago Translation: We had to cut our price by 80% because a open waits model with 284B and 13B active parameter called DeepSeek v4 flash just price/performance mocked us again.   submitted by   /u/InternationalGap3698 [link]   [comments] 38 Hacker News — AI on Front Page community 1mo ago DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis Article URL: https://artificialanalysis.ai/models/deepseek-v4-flash-ga Comments URL: https://news.ycombinator.com/item?id=49120299 Points: 285 # Comments: 139 9 Hacker News — AI on Front Page community 1mo ago DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis Article URL: https://artificialanalysis.ai/models/deepseek-v4-flash Comments URL: https://news.ycombinator.com/item?id=49120299 Points: 423 # Comments: 229 20 Latent.Space news-outlet 1mo ago [AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization Distillation is all you need! 28 arXiv — NLP / Computation & Language research 1mo ago Using Large Language Models for Idea Generation in Innovation arXiv:2607.27553v1 Announce Type: cross Abstract: This research evaluates the efficacy of large language models (LLMs) in generating new product ideas. To do so, we compare three pools of ideas for new products targeted toward college students and priced at 50 dollars or less.… 31 Simon Willison community 1mo ago Advancing the price-performance frontier with GPT‑5.6 Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling this: in How GPT‑5.6 fuses frontier intelligence with frontier efficiency they… 20 TechCrunch — AI news-outlet 1mo ago Friend, the lonely AI wearable, returns with a new voice and a much bigger price tag Friend, the AI wearable, can now talk to its users — for an enhanced price. 38 Hacker News — AI on Front Page community 1mo ago Advancing the price-performance frontier with GPT‑5.6 Article URL: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/ Comments URL: https://news.ycombinator.com/item?id=49112867 Points: 262 # Comments: 162 27 r/LocalLLaMA community 1mo ago How close are we to local llama robotics for consumer price point? I'm guessing 3 years, what do you think? In other words: many of us will be able to afford a general purpose robot in 3 years to experiment with in the home. Cost roughly $5k? Probably small size, but hopefully still able to do the dishes and operate a vacuum.   submitted by… 13 OpenAI official-blog 2mo ago Advancing the price-performance frontier with GPT-5.6 Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale. 35 Smol AI News news-outlet 2mo ago not much happened today **OpenAI** aggressively cut prices for **GPT-5.6 Luna** by 80% and **Terra** by 20%, introducing a faster **Sol Fast** tier with up to 2.5× lower latency at double the price, improving agent workflow costs by roughly 10×. The **ARC-AGI-3** debate highlighted that the complete… 14 arXiv — Machine Learning research 2mo ago Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes arXiv:2607.27188v1 Announce Type: new Abstract: Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation,… 30 Vercel — AI dev-tools 2mo ago AI Gateway: GPT-5.6 pricing and speed updates On AI Gateway , GPT-5.6 Luna and GPT-5.6 Terra are now cheaper and GPT-5.6 Sol is faster. AI Gateway adds no markup on token pricing, so these changes reach you at the upstream rate. The changes apply to both short and long context pricing. Model Change Input: Short context (per… 4 TechCrunch — AI news-outlet 2mo ago Mark Zuckerberg predicts that billions of people will have personal AI agents in five years As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the price. 19 Dwarkesh Podcast news-outlet 2mo ago Why compute might get 10x+ more expensive in coming years If a human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot price. 19 r/LocalLLaMA community 2mo ago dropped 4k on a spark, am I crazy? Saw that the Asus Ascent 1tb was going for $3,950 from a few sources, couldn't stop thinking about it, finally just went ahead and did it. Am I completely insane? Will I regret this? I can't imagine the price will go down any time soon, so it seems like a good idea and I… 34 r/LocalLLaMA community 2mo ago Nvidia is expected to raise GeForce RTX GPU prices again by up to 30%   submitted by   /u/ab2377 [link]   [comments] 21 r/LocalLLaMA community 2mo ago I've been tracking RTX 5090 prices across EU stores since March, it's up €1,061 and still climbing Been running a GPU price tracker ( https://www.pricesquirrel.com ) since March, covering 20+ EU stores, recently added RAM, SSDs and CPUs too. Every GPU tier has gotten cheaper since launch. The RTX 5090 has done the exact opposite. The data: The ASUS TUF Gaming RTX 5090 OC was… 34 TechCrunch — AI news-outlet 2mo ago Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing Cursor says India is now its third-largest market globally and plans to expand local hiring and enterprise sales. 23 arXiv — Machine Learning research 2mo ago Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs arXiv:2607.22786v1 Announce Type: new Abstract: In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. The financial time series are price series such… 38 arXiv — Machine Learning research 2mo ago Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features arXiv:2607.23370v1 Announce Type: new Abstract: Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fat-tailed returns, non-stationary dynamics, and a price discovery process influenced by social discourse on Reddit… 23 r/LocalLLaMA community 2mo ago First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.   submitted by   /u/fulgencio_batista [link]   [comments] 37 r/LocalLLaMA community 2mo ago I want to run Kimi K3 at home, so I’m trying to make 2.8T-scale experimentation cheaper Hey r/LocalLLaMA , I’m a retired engineer with a background in distributed computing, currently running a 1-person startup. Like many people here, I’d love to experiment with 2T+ MoE models locally. The problem is that I don’t have an H100 cluster in my living room. So I’ve been… 12 Page 5 of 10 · 500 articles ← Newer Older →