News / #version-bump Tag Version Bump 381 articles archived under #version-bump · RSS Sign in to follow r/LocalLLaMA community 3h ago It's actually crazy how good DSv4 Flash 0731 is I didn't think we'd get here so quickly. I can run this shit on a computer I spent less than $2k for (back before prices exploded). Crazy world Source: Artificial Analysis Intelligence Index v4.1.1   submitted by   /u/Master-Meal-77 [link]   [comments] 10 Ollama releases dev-tools 11h ago v0.32.11 launch: add DeepSeek Harness integration ( #17733 ) 23 ComfyUI releases dev-tools 13h ago v0.33.1 ComfyUI v0.33.1 27 Anthropic SDK (Python) releases dev-tools 15h ago v0.122.0 0.122.0 (2026-08-13) Full Changelog: v0.121.0...v0.122.0 Features api: add output_behavior to dream creation (create a new memory store or update the input store in place) ( 852c4bb ) Bug Fixes bedrock,aws: run SigV4 signing off the event loop in async clients ( #334 ) ( 2bae6c8… 23 ComfyUI releases dev-tools 16h ago v0.33.0 ComfyUI v0.33.0 33 r/LocalLLaMA community 20h ago GitHub - deepseek-ai/deepseek-harness 🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one… 32 r/LocalLLaMA community 1d ago How do you plan to run Qwen3.8-2.4T-A95B locally? To my fellow crazies, the few. Those who dared wrestle with llama-70b, mistral-large, goliath, mistral8x22B, DeepSeekV2/3, wept when llama4 behemoth was announced, picked yourself up and are now wrestling with DeepSeekV4Pro, GLM5.2, MiMoV2.5Pro and sometimes dare dream of… 15 Ollama releases dev-tools 1d ago v0.32.10 agent: allow multiple edits per edit tool call ( #17711 ) 15 Ollama releases dev-tools 1d ago v0.32.10 What's Changed Models that don't set a repeat_penalty now default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding; set a per-model parameter if an older model repeats itself. Faster prefill on NVFP4 MLX models with a global scale, about… 5 Ollama releases dev-tools 1d ago v0.32.10-rc0: nn: speed up prefill on double-scale nvfp4 models ModelOpt checkpoints apply a float32 global scale to every projection output on top of the per-group quantization scales. Running the multiply and the cast back to the activation dtype as separate eager ops costs an extra kernel launch and a materialized intermediate per… 14 vLLM releases dev-tools 1d ago v0.27.2rc0: [Spec Decode] DSpark confidence-scheduled verification (#47808) Signed-off-by: Lucas Wilkinson [email protected] Signed-off-by: Lucas Wilkinson [email protected] Signed-off-by: Benjamin Chislett [email protected] Signed-off-by: Lucas Wilkinson [email protected] Signed-off-by: Nick Hill… 26 r/LocalLLaMA community 1d ago Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show Link to the article: KV Cache Quantization on Gemma 4 31B: Non-QAT vs QAT KLD benchmarks with BeeLlama.cpp v0.4.3 , fork of llama.cpp with more KV cache quantization options, comparing Gemma Q4_0 non-QAT vs Gemma Q4_0 QAT. Long story short: QAT is much more friendly to KV cache… 4 arXiv — NLP / Computation & Language research 2d ago Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation arXiv:2608.10812v1 Announce Type: new Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a… 30 OpenAI Python SDK releases dev-tools 2d ago v3.0.0 3.0.0 (2026-08-12) ⚠ BREAKING CHANGES api: HTTPX2 is now the default HTTP client, and httpx is no longer installed automatically. Applications using custom HTTPX clients, transports, or configuration objects must migrate to their HTTPX2 equivalents or use the temporary,… 5 ComfyUI releases dev-tools 2d ago v0.32.0 ComfyUI v0.32.0 13 r/LocalLLaMA community 2d ago [llama.cpp PR #26608] Ling-3.0 support (unmerged) aetherbird has done some great work getting Ling-3.0 to work in llama.cpp. The architecture is generally identical to deepseekv2. I recently added a microscopic 40 line PR to his that adds support for the Tiny model, works great. Using it for home assistant voice with decent… 23 OpenAI Python SDK releases dev-tools 2d ago v2.54.0 2.54.0 (2026-08-11) Features api: Add new Responses model identifiers ( #3595 ) ( 0652787 ) Bug Fixes api: clarify audio upload metadata requirements ( #3596 ) ( 28888f9 ) Chores api: Update generated-file header attribution to Castiron ( #3583 ) ( ea17fda ) 23 vLLM releases dev-tools 3d ago v0.27.1: [CI] Limit Arctic import check to x86 test images The arm64 test lockfile intentionally omits arctic-inference, so only validate its native extension on platforms where the package is installed. Co-authored-by: OpenAI Codex [email protected] Signed-off-by: khluu [email protected] 19 Ollama releases dev-tools 3d ago v0.32.9: nemotron_h: support the Nemotron 3.5 prompt layout Select the 3.5 parser and renderer from its checkpoint template, preserve its prompt semantics, and map medium reasoning effort to the final-user annotation expected by the reference template. Exercise parser and renderer registration, create-time metadata inference, and exact… 12 Ollama releases dev-tools 3d ago v0.32.8-rc0 llama.cpp update ( #17659 ) 17 Ollama releases dev-tools 3d ago v0.32.8 llama.cpp update ( #17659 ) 29 Ollama releases dev-tools 3d ago v0.32.7 Muse Glimmer Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Support for NVIDIA, AMD, and other platforms will be available in the coming days. Muse Glimmer , Meta's newest open model and the first released by Meta… 33 vLLM releases dev-tools 4d ago v0.27.0: [Kimi][MM] disable kimi_vit's dynamic torch.compile for TPU (#51196) Signed-off-by: Linkun Chen [email protected] (cherry picked from commit 7f58e82 ) 11 vLLM releases dev-tools 4d ago v0.27.0rc2 v0.27.0rc2 31 ComfyUI releases dev-tools 6d ago v0.31.1 ComfyUI v0.31.1 17 ComfyUI releases dev-tools 6d ago v0.31.0 ComfyUI v0.31.0 32 Anthropic SDK (Python) releases dev-tools 6d ago v0.121.0 0.121.0 (2026-08-07) Full Changelog: v0.120.2...v0.121.0 Features api: add mid-conversation-tool-changes-2026-07-01 beta ( c7d1531 ) api: add support for session budgets, advisor tool, pinned inference location and skills auto-loading from GitHub ( 193bae0 ) Chores api: remove… 16 r/LocalLLaMA community 7d ago My issue with Artificial Analysis's 'intelligence index' I swear AA is not the bipartisan they so claim. An open source mode (Qwen 3.8 max) was number 1 on the agentic index, then they just so happen to launch "v4.1.1" of their index in which they just adjusted the weights of the gdpval and t3 banking so that it would be lower than… 13 vLLM releases dev-tools 7d ago v0.27.0rc1 v0.27.0rc1 10 r/LocalLLaMA community 7d ago KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0 , fork of llama.cpp with more KV cache quantization options. Models: Qwen 3.6 27B Q5_K_S 64k context, Gemma 4 31B Q5_K_S 16k context Standard quants, extended:… 9 ComfyUI releases dev-tools 9d ago v0.30.2 ComfyUI v0.30.2 33 Ollama releases dev-tools 9d ago v0.32.6-rc0 llama.cpp update ( #17545 ) 28 Ollama releases dev-tools 9d ago v0.32.6 llama.cpp update ( #17545 ) 33 r/LocalLLaMA community 9d ago I built a DwarfStar-inspired Vulkan/Metal inference engine for Qwen3.6-35B-A3B on 16 GB machines Disclosure: I’m the author and maintainer of QuarkStar. I built QuarkStar , a small native inference engine inspired by Antirez’s DwarfStar. QuarkStar currently supports: Qwen3.6-35B-A3B , using the same Antirez-inspired Q2 and Q2/Q4 quantization recipes KAT-Coder-V2.5-Dev , the… 6 r/LocalLLaMA community 9d ago Optimised DSv4-Flash for 2x GH200: 10,000 tok/s PP, >300 tok/s TG on SGLang There are some PRs to use and a nice trick to speed up PP on really longs contexts in my write up. Hope it helps! TL;DR: On this dual GH200 box, you build vLLM v0.26.0 from source, add the merged DSV4 cache-layout patch (PR #48993), disable async scheduling, and run DSpark at 6… 25 OpenAI Python SDK releases dev-tools 10d ago v2.53.0 2.53.0 (2026-08-03) Features api: Add gpt-5.5 and tool name/namespace to Responses types ( #3569 ) ( dd1202d ) Bug Fixes ci: avoid NumPy source builds and duplicate HTTPX coverage ( #3573 ) ( b58332f ) 8 OpenAI Python SDK releases dev-tools 10d ago v2.52.1 2.52.1 (2026-07-31) Full Changelog: v2.52.0...v2.52.1 Chores ci: pin setup-uv v5 to its underlying commit ( #3560 ) ( cbdc98b ) 37 r/LocalLLaMA community 10d ago an espresso Q/A model running fully offline on an ESP32S3 i already had an esp32 generating stories, but generating text is not the same as receiving a question and giving a useful answer. barista v0.1, a small model trained for espresso troubleshooting and running on an esp32s3 n16r8 witout cloud. you type a question(over usb for… 24 r/LocalLLaMA community 11d ago [RELEASE] SupraBrain-50M-v0.1 Hey there! So today we're releasing SupraBrain-50M, a hybrid language model that combines Gated DeltaNet linear recurrence with Sliding-Window Attention and Surprise-Gated update mechanisms to deliver very strong performance. Here are the benchmarks:… 19 llama.cpp releases dev-tools 11d ago b10237 llama : MTP support for DeepSeek V3.2 ( #26457 ) llama : MTP support for DeepSeek V3.2 model : no need to include MTP layers during DeepSeek V3.2 model type discovery Co-authored-by: Stanisław Szymczyk [email protected] Website: https://llama.app macOS/iOS: macOS Apple Silicon… 10 ComfyUI releases dev-tools 11d ago v0.30.1 ComfyUI v0.30.1 34 ComfyUI releases dev-tools 11d ago v0.30.0 ComfyUI v0.30.0 7 r/LocalLLaMA community 12d ago Xberg v1 is out Hi all, I'm happy to announce that Xberg v1 is out. Xberg is the successor to Kreuzberg, equivalent to what would have been Kreuzberg v5. It's a content intelligence framework that handles a very wide range of inputs: documents (currently 101 formats), code and data formats… 16 r/LocalLLaMA community 12d ago Koboldcpp v1.118 released   submitted by   /u/Fcking_Chuck [link]   [comments] 12 r/LocalLLaMA community 12d ago DSv4 Flash 0731 Running on Unoptimized Single 3090 System https://preview.redd.it/eqqebec92sgh1.png?width=873&format=png&auto=webp&s=4e9a120e439dc0a64ee4aef87d84b45c51f6bf02 - 3x32 DDR4 2400 Mhz ECC. CPU itself supports quad channel, and we already know why am i haven't filling those slot yet. - E5 2690v4. - 3090 Running on 250W. -… 15 OpenAI Python SDK releases dev-tools 13d ago v2.52.0 2.52.0 (2026-07-31) Full Changelog: v2.51.0...v2.52.0 Features api: content provenance checks ( 1d6c118 ) Bug Fixes client: honor Retry-After delays up to two minutes ( #3555 ) ( 7fa7946 ) Documentation add API-key mTLS HTTP client recipes ( #3552 ) ( 7a3d5e4 ) 30 ComfyUI releases dev-tools 14d ago v0.29.2 ComfyUI v0.29.2 37 ComfyUI releases dev-tools 14d ago v0.29.1 ComfyUI v0.29.1 9 r/MachineLearning community 14d ago MLVC: Multi-platform Learned Video Codec for Real-World Deployment [P] I've always found it a little strange that AI is everywhere, but the codecs we use in practice are the traditional hand-engineered systems like h.264, h.265, av1. Alexnet started the wave of neural networks replacing hand-engineered systems, but 14 years later traditional codecs… 34 OpenAI Python SDK releases dev-tools 14d ago v2.51.0 2.51.0 (2026-07-30) Full Changelog: v2.50.0...v2.51.0 Features api: fast tier ( 8808ed2 ) Bug Fixes api: add fast tier to helper methods ( 6064126 ) 28 Page 1 of 8 · 381 articles Older →