News / #version-bump Tag Version Bump 500 articles archived under #version-bump · RSS Sign in to follow llama.cpp releases dev-tools 17h ago b11222 common : avoid side effects around params parsing ( #29537 ) register --rpc unconditionally and call llama_supports_rpc() only from its handler print server "initialization ..." log after args are parsed Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL Website: https://llama.app… 32 r/LocalLLaMA community 17h ago Mimo v2.6 flash MOPD   submitted by   /u/Automatic-Arm8153 [link]   [comments] 16 r/LocalLLaMA community 1d ago Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_nonogram_logic_puzzle/ . That thread shaped v1.2: Reasoning effort is explicit per run Every prompt and output is public. All current top ranking private and open weight… 7 r/LocalLLaMA community 1d ago 5 items in MiMo-V2.6 3 visible models are 1T, 311B and 9B, let's dream about two more... ;)   submitted by   /u/jacek2023 [link]   [comments] 32 r/LocalLLaMA community 1d ago Koboldcpp v1.122 released   submitted by   /u/Fcking_Chuck [link]   [comments] 18 r/LocalLLaMA community 1d ago Where is MiMo V2.6 9B Distill RL? Was gonna test it out but I can't find the RL checkpoint.   submitted by   /u/Aggravating-Push-207 [link]   [comments] 13 r/LocalLLaMA community 2d ago Kev 4B topped out in every Tetris game I ran. Mica v0.1 4B cleared about 4x more lines and survived two of them to the end I had Mica (my 4B decision model) and Kev 4B play the same Tetris games, same seed and same piece order, one RTX 3090. For context, this is a side project. Training and all the experiments ran on rented 3090s, about $30 in total. Every turn both get the board and 4 possible… 19 r/LocalLLaMA community 2d ago Mica v0.1 4B got an iron pickaxe in real Minecraft without generating a single token Mica v0.1 4B playing a real Minecraft 1.20.4 server. Video attached. How it works - Each step the bot's live game state (inventory, nearby blocks, entities, last result) is written out as text. - Mica scores the candidate commands and picks the next one. It never generates text.… 31 r/LocalLLaMA community 2d ago Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time I've been building a small decision model for agent loops: gates, routers, "should I ask the user or just act" checks. It's out now as Mica v0.1 4B (Apache-2.0). What it does You give it a state, a question and the allowed answers, and it returns a calibrated probability for… 18 ComfyUI releases dev-tools 2d ago v0.37.4 ComfyUI v0.37.4 36 ComfyUI releases dev-tools 2d ago v0.37.3 ComfyUI v0.37.3 19 r/LocalLLaMA community 2d ago Ion v0.2.0 — No install. No backend. Just one HTML file. The new Ion is available, a harness that run directly from a single HTML file, no install or backend required: is now more capable, more customizable, and has better tools!   submitted by   /u/fredconex [link]   [comments] 26 Ollama releases dev-tools 3d ago v0.40.0-rc0: llama-server: prepare to remove compatibility patch Add manifest-list storage so runner-specific manifests can coexist under one tag while preserving existing v1 tags as best-effort downgrade anchors. Show/list/copy/remove/pull/push now understand runner and digest selection and transfer referenced child manifests and layers. Add… 38 r/LocalLLaMA community 3d ago R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB… Here's my *first* implementation of KVA projectors on QFN (just the uncensored model for now) the highlights are basically as follows for using the projectors at each different layer: Starting at layer 12, prompt processing speeds up 1.85x [1700 t/s -> 3150 t/s] at the tradeoff… 24 OpenAI Python SDK releases dev-tools 4d ago v3.19.2 3.19.2 (2026-09-23) Bug Fixes preserve single files for fallback extraction paths ( #3875 ) ( bfd3680 ) Chores api: clarify approximate web search location defaults ( #3953 ) ( a95c95e ) api: clarify Realtime modality array definitions ( #3954 ) ( e79cf53 ) api: correct… 9 Ollama releases dev-tools 4d ago v0.34.4-rc1: mlxrunner: Update XGrammar to 0.2.7 for structured outputs We pick up schema fixes for typed dictionary values and short arrays. 12 Ollama releases dev-tools 4d ago v0.34.4: mlxrunner: Update XGrammar to 0.2.7 for structured outputs We pick up schema fixes for typed dictionary values and short arrays. 18 llama.cpp releases dev-tools 4d ago v0.5.0 Overview This release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. It adds HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, ggml 0.25.0 backend improvements, multi-address HTTP… 26 OpenAI Python SDK releases dev-tools 4d ago v3.19.1 3.19.1 (2026-09-23) Bug Fixes chat: preserve single-pass tool iterables ( #3770 ) ( 33ffa1f ) client: merge HTTP headers case-insensitively ( #3486 ) ( 5e39766 ) Chores api: clarify Chat Completions seed limits ( #3945 ) ( be9d666 ) Documentation clarify collaborator-only pull… 19 ComfyUI releases dev-tools 4d ago v0.37.2 ComfyUI v0.37.2 10 r/LocalLLaMA community 4d ago MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam MiMo-V2.6-Pro has an insanely high score of 46 on AA, putting it at the head of the opensource models available. It also costs pennies. Flash is not out on AA yet, but it costs less than half on datacenter and is slightly below on Xiaomi's own benchmarks. It also fits in 192GB,… 12 vLLM releases dev-tools 5d ago v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281) Signed-off-by: Andreas Karatzas [email protected] Co-authored-by: OpenAI Codex [email protected] 21 Ollama releases dev-tools 5d ago v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550) mlx: speed up Qwen 3.8 prompt processing Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU. address comments 31 OpenAI Python SDK releases dev-tools 5d ago v3.19.0 3.19.0 (2026-09-22) Features api: add GCP external storage support ( #3943 ) ( d12d60f ) api: add GPT-Rosalind research model ( #3940 ) ( d041d73 ) Bug Fixes _utils/_transform: propagate api_exclude in _async_transform_recursive ( #3324 ) ( 161ae65 ) api: handle omission markers… 7 Simon Willison community 5d ago Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war Yesterday was Grok 4.7 ( pelicans ) and MiMo v2.6 Flash/Pro ( more pelicans ). Today Anthropic released Claude Opus 5.5 , and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna . It's going to take a while to get a good read on all of these new models, but here are my… 10 r/LocalLLaMA community 5d ago MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap Some people here say MiMo-V2.6 is bad with tools and are going back to GLM-5.3-Flash. I spent today running MiMo-V2.6-Flash-RL as the backend for an agent harness, on 2× DGX Spark with vLLM, using the tonyd2wild recipe. Most of the "tool problems" I hit turned out to be serving… 8 ComfyUI releases dev-tools 5d ago v0.37.1 ComfyUI v0.37.1 38 OpenAI Python SDK releases dev-tools 5d ago v3.18.0 3.18.0 (2026-09-22) Features api: add GPT-6 Sol and Luna model identifiers ( #3935 ) ( 455ce1b ) 34 Anthropic SDK (Python) releases dev-tools 5d ago v1.8.0 1.8.0 (2026-09-22) Full Changelog: v1.7.0...v1.8.0 Features api: add support for claude-opus-5-5, inline tool definitions and MCP tool-list pinning (beta) ( b5cc700 ) Bug Fixes api: share one evaluated_permission enum across Managed Agents events ( f4f51c8 ) streaming: avoid a… 28 llama.cpp releases dev-tools 5d ago b11102 convert: add MiMo-V2.6 support ( #29257 ) convert: add MiMo-V2.6 support Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert Update conversion/mimo.py fix: use autoparser Co-authored-by: Sigbjørn Skjæret… 5 r/LocalLLaMA community 5d ago Fork of FreeToken with DeepSeek-V4.1, vision and speculative decoding (2x3090 numbers inside) I've been running FreeToken on my 2x3090 box for a while and ended up maintaining a fork of it. Posting it in case it's useful to anyone else here. Quick context if you haven't used it: FreeToken is an edge-native MoE serving engine. It offloads experts to host RAM/NVMe and… 31 Hacker News — AI on Front Page community 5d ago Fearless SIMD v1.0 Article URL: https://linebender.org/blog/fearless-simd-1-0/ Comments URL: https://news.ycombinator.com/item?id=49800085 Points: 227 # Comments: 34 26 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B · Hugging Face   submitted by   /u/Aggravating-Push-207 [link]   [comments] 36 r/MachineLearning community 6d ago Xiaomi releases MiMo-V2.6: "Frontier intelligence, all the modalities, built in public." [N] The total training cost was just $3.5M. The model comes with a live benchmaxxing dashboard. https://preview.redd.it/89uurv5r21rh1.png?width=1518&format=png&auto=webp&s=e5d641ef941882e6dc5c023a78695f40173c5c80 https://mimo.xiaomi.com/mimo-v2-6   submitted by  … 35 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B https://preview.redd.it/ybwqqbrst0rh1.png?width=811&format=png&auto=webp&s=1e40c5b53e304caf2a10efa2baa6f998b9a9a0bb MiMo-V2.6-Distill-Qwen-9B is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data. Made by Xiaomi… 38 Latent.Space news-outlet 6d ago [AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M crowning a new Chinese frontier lab 38 llama.cpp releases dev-tools 6d ago b11096 ci : update Level Zero SDK to v1.33.1 and enable the L0/oneDNN CMake … 11 OpenAI Python SDK releases dev-tools 6d ago v3.17.0 3.17.0 (2026-09-22) Features api: add external storage configuration management ( #3909 ) ( 6332577 ) api: add safety case retrieval ( #3911 ) ( a87b938 ) api: add safety warning and deactivation webhook events ( #3908 ) ( 19f1f37 ) api: add session environment reset events (… 14 r/LocalLLaMA community 6d ago Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact Recently, the new DeepSeek-V4.1-Flash architecture showed how a causal encoder-decoder can work, but it was trained from scratch. Model Grafting does it to an existing model: cut at some depth, let the lower layers read the prompt, and use the upper layers get for encoder's… 38 r/LocalLLaMA community 6d ago MiMo-V2.6 distilled themselves into Qwen 9B! https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B That's serious.   submitted by   /u/Beamsters [link]   [comments] 6 r/MachineLearning community 6d ago Jev's calibration was measured. The LLMs won [D] Source: Jev Benchmarks Its training method is literally called "Reinforcement Learning for Calibrated Decisions." Calibration gap vs human labels (lower = better): Yes/no: Jev 5.0, Gemini 3.8 Flash 2.0 Pick-one: Jev 9.8, DeepSeek V4.1 Flash 2.8 Rubric: Jev 19.7, GLM-5.3 12.9 It… 29 r/LocalLLaMA community 6d ago Wow, Mimo 2.6 pro seems to be pretty good, but requires more prompting than Sol Although it makes mistakes and use more tokens, but after some reprompting , It(max) gives really good outputs like on par with 5.6 sol and close to astra High in one task.. IT will likely vary on tasks. I hope ds v4.1pro is gonna be really good or at least as good as Kimi K3.… 27 r/LocalLLaMA community 6d ago Mimo v2.6-Flash-RL vs open-weight models Since there’s no comparison chart on the model page, I asked Perplexity to compare it against some relatively small open-weight models in a similar size range. Here are the results. Upd. Terminal-Bench 4.0 results: MiMo‑V2.6‑Flash‑RL — 28.8% DeepSeek‑V4‑Flash‑0731 — 12.0%… 31 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B we're so back?!?   submitted by   /u/VoiceApprehensive893 [link]   [comments] 24 Hacker News — AI on Front Page community 6d ago Xiaomi MiMo v2.6 Article URL: https://mimo.xiaomi.com/mimo-v2-6 Comments URL: https://news.ycombinator.com/item?id=49792730 Points: 238 # Comments: 104 28 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face   submitted by   /u/Bestlife73 [link]   [comments] 25 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face   submitted by   /u/Bestlife73 [link]   [comments] 9 vLLM releases dev-tools 7d ago v0.30.0: [Build] Fix DeepGEMM CUDA 12.9 release builds (#57554) Signed-off-by: khluu [email protected] Co-authored-by: OpenAI Codex [email protected] Co-authored-by: Jee Jee Li [email protected] (cherry picked from commit eb87980 ) 15 arXiv — NLP / Computation & Language research 7d ago Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5 arXiv:2510.07024v3 Announce Type: replace Abstract: LLMs are remarkable artifacts that have revolutionized a range of knowledge-intensive tasks. A significant contributor is their factual knowledge, which, to date, remains poorly understood, and is usually analyzed from biased… 33 Vercel — AI dev-tools 7d ago MiMo V2.6 models now available on AI Gateway MiMo V2.6 Pro , MiMo V2.6 Flash , and MiMo V2.6 Pro UltraSpeed from Xiaomi are now available on AI Gateway . MiMo V2.6 combines coding, reasoning, and tool use with native text, image, audio, and video understanding. Its 1M token context supports long repositories, tool traces,… 14 Page 1 of 10 · 500 articles Older →