News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 13d ago New DeepSeek V4 Flash 0731 vs ChatGPT Luna comparison   submitted by   /u/perelmanych [link]   [comments] 23 r/LocalLLaMA community 13d ago Deepseek V4 Flash 0731. LM Studio loading only into RAM. The model refuses to load into VRAM and uses only RAM. What can be an issue? Q2_K_XL from Unsloth if that changes something.   submitted by   /u/esw123 [link]   [comments] 29 r/LocalLLaMA community 13d ago DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026 March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models available to run locally on <8K USD (us prices - just guestimating/not exact)… 19 r/LocalLLaMA community 13d ago Are 1B LLMs Going Away in 2026? I don't know much about llms aside from downloading them through a frontend and running them on my laptop or potato phone. Google released gemma 4, but unlike gemma 3, there isn't a 1b model this time. Llama also had a 1b model before, but there doesn't seem to be a new one.… 21 r/LocalLLaMA community 13d ago Qwen 3.6 27B Q5 on 3x2080ti: 55tps with llama.cpp. Can I squeeze out more? CPU: Threadripper 3970X RAM: 128GB DDR4 GPUs: 3x2080ti 11GB The current best parameters to run it: llama-server \ --model Qwen3.6-27B-Q5_K_S.gguf \ --n-gpu-layers 999 \ --split-mode tensor \ --flash-attn on \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --ctx-size 16384 \… 8 r/LocalLLaMA community 13d ago What speeds are everyone getting with deepseek v4 flash 0731? What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context window of 128000, -ub/-b at 4096, “q8” unsloth’s lossless quant   submitted by… 36 r/LocalLLaMA community 13d ago Me: Worn out from all the new model drops this week, but still hyped for all the great new releases. I mean seriously y’all, what an amazing past few days. So many awesome new models to test out in the mid range model sizes.   submitted by   /u/Porespellar [link]   [comments] 15 Latent.Space news-outlet 13d ago [AINews] not much happened today apart from DeepSeek V4-Flash 0731, a quiet day. 4 r/LocalLLaMA community 13d ago DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf Antirez stealthily uploaded the new weights in the old folder... and there we were tapping our fingers. https://huggingface.co/antirez/deepseek-v4-gguf/tree/main   submitted by   /u/challis88ocarina [link]   [comments] 14 r/LocalLLaMA community 13d ago [audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox . It is closer to prompt-directed voice acting. DramaBox is built on the LTX-2.3 audio architecture, and prompts can control emotion, delivery, laughs, sighs, pauses, transitions, and speaker behavior. Example… 15 Simon Willison community 13d ago deepseek-ai/DeepSeek-V4-Flash-0731 deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on Hugging Face - but it appears to punch well above its weight. Artificial Analysis rank it ahead of MiniMax M3… 13 r/LocalLLaMA community 13d ago DeepSeek v4 Flash for DS4 (DwarfStar) GGUF w/ DSpark MTP Head I'm an avid user of Deepseek v4 Flash via antirez's DS4 DwarfStar inference engine , and so when the new checkpoint dropped, the first thing I did was rent a cloud box and spin up a quantization for use in my DS4 deployment. Props to Unsloth for getting their GGUFs out so… 7 Simon Willison community 13d ago Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp) Tuesday was Stateless MCP day - the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant change to the MCP spec since it first launched, and has also served to reignite my personal… 10 Simon Willison community 13d ago llm-mcp-client 0.1a0 Release: llm-mcp-client 0.1a0 See this blog entry . Tags: llm , model-context-protocol 5 r/LocalLLaMA community 13d ago Now Suddenly too many choices for DGX Spark with Qwen 3.5 122B . What would be the next upgrade? Laguna 2.1 at NVFP4 Deepseek v4 at Q2 Inkling-Small at IQ3 Which models you guys running now ? How it compares to 122b? Upcoming in few days : Ling 3.0 124B (Could be new king) LongCat 69B A3B ( very interesting worker model) What else?   submitted by   /u/Voxandr [link]… 10 r/LocalLLaMA community 13d ago We've gotten some great medium sized models lately (DSV4 Flash 0731, Inkling Small, Laguna S 2.1, Step 3.7 Flash) but does anybody else want to see some new 70-80b contenders? I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B. On my hardware I get like 500 to 800 tok/s prefill and 16 to 22 tok/s gen on ~120B class models, which is not the worst but it does get a bit annoying on agentic coding tasks. If we… 8 r/LocalLLaMA community 13d ago Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper https://preview.redd.it/h7zv5tb3tmgh1.png?width=2854&format=png&auto=webp&s=507380e8f862c18f10f7c5c84da9e8d1c59139b0 Deepseek's new flash model is unexpectedly cheap and high-performing across useful benchmarks. It's priced at $0.09 / $0.18 per 1M. Truly "intelligence too cheap… 23 Ars Technica — AI news-outlet 13d ago Claude published malicious code to the Internet and attacked 3 real companies Had the hacks used conventional methods, someone would likely go to prison. 6 TechCrunch — AI news-outlet 13d ago Google nixes its Earth AI feature one day after launch, amid criticism it would spread misinformation A tool that allowed anyone to generate fake AI-generated imagery and superimpose it over real Google Earth maps quickly spurred backlash. 16 r/LocalLLaMA community 13d ago Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ? Hello guys, I'm curious about running DeepSeek-V4-Flash-0731 locally. Since it’s a Mixture of Experts (MoE) model with only 13B active parameters, I was hoping the VRAM requirements might be manageable. Did someone tried out in some reasonable GPU sizes up to 48GB VRAM? Thanks… 37 llama.cpp releases dev-tools 13d ago b10212 llama : load MTP tensors only if they are really used ( #26296 ) llama : load MTP tensors only if they are really used llama : skip loading MTP (if not used) in remaining models that support MTP Co-authored-by: Stanisław Szymczyk [email protected] Website: https://llama.app… 35 r/LocalLLaMA community 13d ago Translation: We had to cut our price by 80% because a open waits model with 284B and 13B active parameter called DeepSeek v4 flash just price/performance mocked us again.   submitted by   /u/InternationalGap3698 [link]   [comments] 38 r/LocalLLaMA community 13d ago With release of Deepseek V4 I wanted see how the model sizes are trending over time. The trend is that by this time next year, we probably will have Opus 4.5 level models on consumer grade laptops! I was surprised to see that Deepseek V4 Flash is extremely smart and small enough to fit in setup that can be built with < $50,000. Expensive, but not a datacenter. So I wanted to see the trend over time of model sizes and their scores and created above plots using Opus/Sonnet… 32 r/LocalLLaMA community 13d ago DeepSeek V4 Flash unsloth quants are out! 4-bit - 155gb 8-bit - 162gb "Smaller ones are coming"   submitted by   /u/RunawayPeeko [link]   [comments] 30 r/LocalLLaMA community 13d ago DeepSeek V4 Flash GA ranks the same as Sonnet 5 and Grok 4.5 on DeepSWE Source: https://x.com/deepseek_ai/status/2083084415157022911 & https://deepswe.datacurve.ai/ just combined data view. DeepSeek claims, not verified by DeepSWE yet.   submitted by   /u/sdexca [link]   [comments] 35 r/LocalLLaMA community 13d ago DeepSeek-V4-Flash-0731 unsloth gguf on A100 A100 with 40gb VRAM: 162GB Q8_K_XL ~16.1 tok/s generation Only 15.8GB of 40GB VRAM used with all experts on CPU NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the single 40GB A100 at 17.7 tok/s with 6 experts loaded into VRAM, with Codex driving… 24 r/LocalLLaMA community 13d ago Deepseek flash 0731 reasoning is hilarious Tell me this is not funny - " OH MY GOD. I THINK I FINALLY SEE IT!!! The black pixels are at the QUAD CENTERS because of the mipmapping of the UV derivative at the quad DIAGONAL ... no. Hmm." I have never seen a reasoning trace say "OH MY GOD," lol.   submitted by  … 4 r/LocalLLaMA community 13d ago Uncensored Multi-Model Releases, LongCat-Flash-Lite with MTPs, Jamba2-Mini, Qwen3.5-9B-Nikusui-v1 with MTPs and Qwen3.5-27B-Nikusui-v1 with MTPs, Available in Safetensors and GGUF Formats! Been working hard for the past month to bring to the community some interesting curios, so for starters we have LongCat-Flash-Lite Uncensored Heretic with MTPs which has never before been uncensored, it is a 69B-A3B model, I spent a long time working on it to make it work with… 35 r/MachineLearning community 13d ago Learning path to fully understand the Kimi K3 technical report?[D] Hi everyone, Can anyone suggest a learning path to fully understand the technical report for Kimi K3? My background: - I've taken a graduate-level deep learning course. - I understand the Transformer architecture, attention, and the basics of LLMs. - I'm familiar with DeepSeek's… 15 r/LocalLLaMA community 14d ago Deepseek v4 flash MXFP4 (original quality) ggufs   submitted by   /u/Antique_Archer_7110 [link]   [comments] 8 r/LocalLLaMA community 14d ago A lesson about retries, hidden in the DeepSeek-V4 paper   submitted by   /u/pmigdal [link]   [comments] 6 r/LocalLLaMA community 14d ago Deepseek V4 Flash on SlopCodeBench While waiting for some of the quants to drop, I load the API with $50 and ran it on SlopCodeBench Just vibe reading the results it seems like Opus 4.8 < Deepseek < Opus 5 https://github.com/michaelasper/benchmarks/blob/main/deepseek-v4-flash-on-slop-code-bench.md I was mostly… 12 r/LocalLLaMA community 14d ago SenseNova U1.5 Lite preview just dropped SenseNova released U1.5-Lite-Preview Benchmarks: Qwen-Image-Bench from 47.14 to 55.20. ImgEdit-Bench from 3.90 to 4.37. GEdit-Bench-en from 7.47 to 8.17. Key updates: 4K native generation with better texture, material, and lighting detail Improved Chinese and English text… 15 r/LocalLLaMA community 14d ago Unsloth Deepseek V4 0731 GGUF's are UP!   submitted by   /u/BlackBeardAI [link]   [comments] 12 r/LocalLLaMA community 14d ago Meituan just dropped LongCat-Flash-Lite-Sparse It’s an MoE with ~3B active params and a 30B n-gram lookup table offloaded to RAM for fast 256k context on a 24GB GPU. Reminds me of Gemma 4’s PLE trick. Initial analysis suggest it wont be replacing my Qwen 3.6 27b.   submitted by   /u/Gohab2001 [link]   [comments] 16 llama.cpp releases dev-tools 14d ago b10206 llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized ( #25871 ) llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized llama : enforce the same K and V cache types for MLA models Co-authored-by:… 23 Simon Willison community 14d ago datasette-agent 0.4a0 Release: datasette-agent 0.4a0 New await context.browser_task() mechanism allowing agent tools to run code directly in the user's browser. #33 This is an exciting new capability: it makes it easy for Datasette Agent plugins to provide tools that execute custom JavaScript in the… 22 r/LocalLLaMA community 14d ago The Chinese LLM release carousel never stops. Place your bets for MiniMax next week.   submitted by   /u/Mountain_Patience231 [link]   [comments] 28 r/LocalLLaMA community 14d ago Rule Suggestion: "Open" models without weight releases should be tagged [no weights] A lot of recent models are being announced with promised open weights, but the weights are either weeks away, or in some cases (looking at you Meta) not being released at all. This sub is about local LLMs - not "maybe local in the future" llms. These models are still useful to… 26 r/LocalLLaMA community 14d ago I ported TurboFieldfare to Qwen 3.6 35B and it runs in 1.4 GB of RAM Was playing around with TurboFieldfare , a Mac engine that runs Gemma 4 26B in ~2 GB by streaming MoE experts off SSD instead of loading them. It only supported that one model, so I added support for Qwen 3.6 35B-A3B. Comparatively, Qwen needs lesser memory. ~1.4 GB vs ~2.1 GB… 11 r/LocalLLaMA community 14d ago deepseek-ai/DeepSeek-V4-Flash-0731 on Huggingface https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731   submitted by   /u/cgs019283 [link]   [comments] 22 r/LocalLLaMA community 14d ago DeepSeek-V4-Flash-0731 now far surpassing the DeepSeek-V4-Pro-Preview in benchmarks   submitted by   /u/SnooBunnies8392 [link]   [comments] 28 Hacker News — AI on Front Page community 14d ago DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis Article URL: https://artificialanalysis.ai/models/deepseek-v4-flash-ga Comments URL: https://news.ycombinator.com/item?id=49120299 Points: 285 # Comments: 139 9 Hacker News — AI on Front Page community 14d ago DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis Article URL: https://artificialanalysis.ai/models/deepseek-v4-flash Comments URL: https://news.ycombinator.com/item?id=49120299 Points: 423 # Comments: 229 20 r/LocalLLaMA community 14d ago New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna   submitted by   /u/MagicZhang [link]   [comments] 25 Vercel — AI dev-tools 14d ago DeepSeek V4 Flash now runs updated weights on AI Gateway DeepSeek V4 Flash now runs on updated weights by default on AI Gateway, with notably stronger agentic capabilities. On Terminal-Bench, it scores 82.7, up 25.8 points from 56.9 in the April preview. Requests to deepseek/deepseek-v4-flash pick up the new weights automatically,… 23 r/LocalLLaMA community 14d ago DeepSeek-V4-Flash-0731 is going to cause another market crash. Beats GLM 5.2, and is the same cost as the previous one.   submitted by   /u/Potential_Top_4669 [link]   [comments] 28 r/LocalLLaMA community 14d ago DeepSeek v4 Flash has a nice bump in Capability DeepSeek V4 Flash: Preview → 2026-07-31 Benchmark Preview 0731 Δ Terminal Bench* 56.9 82.7 +25.8 Toolathlon 51.8 70.3 +18.5 NL2Repo — 54.2 new Cybergym — 76.7 new DeepSWE — 54.4 new Agent Last Exam — 25.2 new Automation Bench — 25.1 new DSBench-FullStack — 68.7 new DSBench-Hard… 28 Hacker News — AI on Front Page community 14d ago DeepSeek-V4-Flash Update Article URL: https://api-docs.deepseek.com/updates/ Comments URL: https://news.ycombinator.com/item?id=49119559 Points: 312 # Comments: 142 13 r/LocalLLaMA community 14d ago The official release Deepseek V4 flash is live on the API   submitted by   /u/mineyevfan [link]   [comments] 34 Page 10 of 10 · 500 articles ← Newer