News / #version-bump Tag Version Bump 500 articles archived under #version-bump · RSS Sign in to follow r/LocalLLaMA community 7d ago A Jev-style model fine-tuned on Qwen3.5 4B This weekend, I did a fun experiment to create something similar to Jev. I LoRA fine-tuned Qwen3.5 4B using a mix of publicly available datasets and synthetic data. For the synthetic data, I used DeepSeek V4.1 Flash, around 25M tokens. I trained the model for about 2 hours on a… 14 ComfyUI releases dev-tools 7d ago v0.37.0 ComfyUI v0.37.0 24 r/LocalLLaMA community 9d ago I enjoyed the daily HF papers today Top 3 papers on HF Daily Paper are all unusually delightful and interesting reads for anyone on the leading edge of local LLMs, agent harness optimization, etc, felt like sharing. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression… 35 Ollama releases dev-tools 9d ago v0.34.3-rc1 server: allow registry cross-host redirects among allowlisted hosts (… 25 Ollama releases dev-tools 9d ago v0.34.3 server: allow registry cross-host redirects among allowlisted hosts (… 16 OpenAI Python SDK releases dev-tools 9d ago v3.16.2 3.16.2 (2026-09-18) Bug Fixes parsing: drop TextFormatT parameterization in parse_response to fix memory leak ( #3084 ) ( #3088 ) ( 009b7f6 ) 6 vLLM releases dev-tools 9d ago v0.30.0rc2 [Bugfix][NIXL] Avoid receive reports for notification-only requests (… 22 Ollama releases dev-tools 9d ago v0.34.3-rc0 api: expose model thinking levels and defaults ( #18473 ) 35 Hacker News — AI on Front Page community 9d ago Claude Code now reads AGENTS.md if there is no Claude.md Article URL: https://code.claude.com/docs/en/changelog Comments URL: https://news.ycombinator.com/item?id=49760187 Points: 271 # Comments: 109 10 OpenAI Python SDK releases dev-tools 9d ago v3.16.1 3.16.1 (2026-09-18) Bug Fixes api: avoid loading unrelated API resources on first use ( #3898 ) ( 68e4317 ) 25 Vercel — AI dev-tools 9d ago v0 now reads npm credentials from shared environment variables v0 now installs private packages from npm and custom registries using credentials stored as shared environment variables on Vercel. This makes it easier for teams to build with their existing design systems, component libraries, and internal packages directly in v0. To get… 27 Anthropic SDK (Python) releases dev-tools 9d ago v1.7.0 1.7.0 (2026-09-18) Full Changelog: v1.6.0...v1.7.0 Features api: add group with display_name to rate limits, deprecate group_type ( 28f0a83 ) tools: add compact_before_next_turn() to the tool runner ( #641 ) ( 8b23fb3 ) Bug Fixes bedrock: raise an API error for eventstream… 24 OpenAI Python SDK releases dev-tools 9d ago v3.16.0 3.16.0 (2026-09-18) Features api: add webhook endpoint management ( #3892 ) ( 9a11f6e ) Chores api: deprecate MCP connector_id ( #3894 ) ( 1ccaf07 ) 26 arXiv — NLP / Computation & Language research 10d ago DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression arXiv:2609.19969v1 Announce Type: new Abstract: The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and… 36 OpenAI Python SDK releases dev-tools 10d ago v3.15.0 3.15.0 (2026-09-18) Features api: add agent session model settings ( #3882 ) ( 4b15817 ) api: add audio-mini model choices ( #3886 ) ( a6eeb3f ) api: add compaction progress events ( #3866 ) ( 98e1d24 ) api: add managed Responses WebSocket sessions ( #3887 ) ( 3b865af ) api: add… 31 Ollama releases dev-tools 10d ago v0.34.2 x/transfer, server: tighten redirect handling for registry requests (… 19 Ollama releases dev-tools 10d ago v0.34.2 What's Changed llama.cpp updates Full Changelog : v0.34.1...v0.34.2-rc0 19 Ollama releases dev-tools 10d ago v0.34.2-rc2: mlxrunner: Release freed KV buffers during speculative decode The decode loop releases MLX's pool of freed buffers every 256 generated tokens, which is also how often the KV cache grows and drops its previous, smaller buffers. The check fires only when the token count lands exactly on a multiple of 256. Speculative decoding emits several… 28 llama.cpp releases dev-tools 11d ago b11020 chat : add message delimiters to the DeepSeek V3.2/V4 parser ( #29008 ) chat : add message delimiters to the DeepSeek V3.2/V4 parser Assisted-by: Claude Co-authored-by: Sigbjørn Skjæret [email protected] Website: https://llama.app Attestations:… 8 vLLM releases dev-tools 11d ago v0.30.0rc1: [Bugfix] Isolate supplemental FlashInfer BF16 autotuning (#57285) Signed-off-by: jiahanc [email protected] Co-authored-by: OpenAI Codex [email protected] 16 arXiv — NLP / Computation & Language research 11d ago SEA-LION-v4.8: A Technical Report arXiv:2609.18310v1 Announce Type: new Abstract: We introduce Nemotron-SEA-LION-v4.8, a family of Southeast Asian Languages in One Network (SEA-LION) built upon NVIDIA Nemotron 3. The family includes 30B-A3B and 120B-A12B models, with both continued-pretrained base checkpoints… 14 vLLM releases dev-tools 11d ago proto-v0.3.0 Release vllm-proto 0.3.0 25 Ollama releases dev-tools 11d ago v0.34.2 What's Changed llama.cpp updates Full Changelog : v0.34.1...v0.34.2-rc0 4 vLLM releases dev-tools 12d ago proto-v0.2.0: vllm-proto 0.2.0 Validated by PR #56538 CI at fa2a26f . 23 OpenAI Python SDK releases dev-tools 12d ago v3.14.1 3.14.1 (2026-09-15) Bug Fixes client: validate retry limits and preserve application errors ( #3867 ) ( f86c721 ) correct typo "th" to "the" in StreamAlreadyConsumed error message ( #3022 ) ( 7186203 ) examples: correct Azure endpoint hostname ( #3298 ) ( 543516c ) examples:… 16 ComfyUI releases dev-tools 12d ago v0.36.0 ComfyUI v0.36.0 29 Ollama releases dev-tools 12d ago v0.34.2-rc0: llama.cpp: version bump b10969 (#18446) llama.cpp build changes resulted in duplicate symbols between libllama and libmtmd. This moves the compat patch into libllama with exported symbols. 21 r/LocalLLaMA community 12d ago Koboldcpp v1.121 released   submitted by   /u/Fcking_Chuck [link]   [comments] 23 Anthropic SDK (Python) releases dev-tools 12d ago v1.6.0 1.6.0 (2026-09-15) Full Changelog: v1.5.0...v1.6.0 Features api: add auto mode tool permissions for Managed Agents ( 909d92f ) api: add compaction parameter and signed compaction blocks (beta) ( 8689179 ) api: add enum types for workspace data-residency geo fields ( 3dc6dbf )… 37 r/LocalLLaMA community 12d ago DeepSeek V4.1F Q4 on M3 Ultra with native DSpark MTP (40tps / 800tps) I liked DeepSeek V4.1 Flash as an agent model but at 16 t/s on ds4 it was painful to sit through a real turn. It seemed some redditors and m3 ultra owners appreciated my glm 53 flash optimizations, so I forked antirez/ds4 for V4.1 Flash to see how much I learned optimizing GLM… 30 Ollama releases dev-tools 13d ago v0.34.1 docs: refresh getting started guides ( #18450 ) 16 Ollama releases dev-tools 13d ago v0.34.1-rc2: API: Deprecate typical_p (#18448) Drop support for creating new models with typical_p parameters, while retaining support for existing GGUF models with the setting. 7 OpenAI Python SDK releases dev-tools 13d ago v3.14.0 3.14.0 (2026-09-14) Features streaming: normalize errors raised while reading streams ( #3827 ) ( d7c41ef ) Bug Fixes bound vector store file polling ( #3401 ) ( ae41bc4 ) example: refresh realtime push_to_talk_app session types ( #2926 ) ( 0b7ad38 ) files: normalize PathLike… 35 r/LocalLLaMA community 13d ago DeepSeek engineer relections on RSI - burying my talent to yesterday Note - This is translated from the actual blog link right at the bottom. A few days ago, DeepSeek v4.1 was released. It raised the ability of small models to a new level. AI is improving much faster than anyone expected. From the first ChatGPT that could only chat simply with a… 21 ComfyUI releases dev-tools 13d ago v0.35.2 ComfyUI v0.35.2 26 r/LocalLLaMA community 13d ago Animated transition from AA Intelligence Index v4.1 to v4.3 I had all the data saved from AA's v4.1 index, so when they upgraded it in the wake of Astra's release, I could actually generate a before/after comparison. All intelligence and price per task are sampled from AA on Sep 3rd and Sep 14th respectively. Price per task of some open… 26 Ollama releases dev-tools 13d ago v0.34.1-rc1 mlx: add mlx patch to docker build context ( #18440 ) 14 llama.cpp releases dev-tools 13d ago v0.4.1 Overview llama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0. API changes Changed llama_sampler_chain_n() to return int32_t instead of int (… 8 Ollama releases dev-tools 13d ago v0.34.1-rc0: MLX: version bump (#18235) MLX: version bump mlx: support ModelOpt global scales in MoE models address comments address comments 19 r/LocalLLaMA community 14d ago 3D viz of how Deepseek Flash v4.1 is different from a typical decode only transformer https://whip.run/experience/356b3eba-23ce-4dc8-bff2-4f7bc1862966   submitted by   /u/tempNull [link]   [comments] 30 r/LocalLLaMA community 14d ago DeepSeek V4.1 Flash beats Astra on AA's new benchmark https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3 AA shipped a new benchmark last week as part of the Intelligence Index v4.3 update — a brand-new private eval that replaces τ³. Astra was farming a ton of points on it and used those to get even… 27 r/LocalLLaMA community 14d ago Hoping for Optimized Smarter Upcoming Models .... Like DeepSeek-V4.1-Flash( KVCache + Engram) in Small/Medium/Big sizes It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs. It would be awesome to have DeepSeek-V4.1-Flash's KVCache + Engram for all Upcoming models. Even for big models. Engram - Heard that approximately… 17 r/LocalLLaMA community 15d ago DS 4.1 and the new Harness I gave DS V4.1 Flash an HLE problem with a bash tool + 2 hours. Hour 1: it wrote three MILP solvers. (225,200) Hour 2: it downloaded the HLE dataset from Hugging Face, found the question, read the answer key (225,600), and concluded its own answer (225,200) was better. I'm equal… 20 vLLM releases dev-tools 15d ago v0.29.1rc0 [watermarking] Dual-key gumbel-max watermarking for speculative decod… 20 r/LocalLLaMA community 16d ago Antirez Deepseek 4.1 flash gguf on HF Q2 is there and Q4 is uploading as I type. Has his github been updated yet? How do you run this? https://huggingface.co/antirez/deepseek-v4.1-flash-gguf/tree/main   submitted by   /u/Queasy_Asparagus69 [link]   [comments] 34 Latent.Space news-outlet 16d ago [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale We agree with Sebastian: this should have been DeepSeek v5 14 r/LocalLLaMA community 16d ago Hot Expert Reload on GPU is what this community needs A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on MOE models with moderate number of active parameters, like Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, GLM 5.3 Flash. With 2x… 38 vLLM releases dev-tools 17d ago proto-v0.1.0 vllm-proto 0.1.0 16 ThursdAI news-outlet 17d ago OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news Also a surprise announcement from CoreWeave and a free ticket code for ThursdAI listeners in SF on September 30th. 26 r/LocalLLaMA community 17d ago Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen I wonder someone will figure out a way to do this with 27B? Throw Qwen3 on this page for demo https://kishida.github.io/webdemos/llkvapprox/   submitted by   /u/T_rex2700 [link]   [comments] 10 Page 2 of 10 · 500 articles ← Newer Older →