News / #version-bump Tag Version Bump 381 articles archived under #version-bump · RSS Sign in to follow r/LocalLLaMA community 2mo ago Xiaomi just claimed 1,000+ tps on a 1T model using a standard 8-GPU server Just saw Xiaomi MiMo announce MiMo-V2.5-Pro UltraSpeed , claiming they broke the 1,000 tokens/sec output barrier on a 1 trillion parameter MoE model . According to them, they’re doing it on a single standard 8-GPU node , not custom wafer-scale hardware like Cerebras and not… 34 Hacker News — AI on Front Page community 2mo ago MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second Article URL: https://mimo.xiaomi.com/blog/mimo-tilert-1000tps Comments URL: https://news.ycombinator.com/item?id=48446639 Points: 252 # Comments: 175 30 Ollama releases dev-tools 2mo ago v0.30.7 docs: update docs examples to use Gemma 4 instead of Gemma 3 ( #16607 ) 7 Anthropic SDK (Python) releases dev-tools 2mo ago v0.107.1 0.107.1 (2026-06-07) Full Changelog: v0.107.0...v0.107.1 Bug Fixes foundry: send x-api-key header for API-key auth ( #62 ) ( 1338141 ), closes #1661 31 Anthropic SDK (Python) releases dev-tools 2mo ago v0.107.0 0.107.0 (2026-06-06) Full Changelog: v0.106.0...v0.107.0 Features api: small updates to Managed Agents types ( 72923f9 ) 35 Ollama releases dev-tools 2mo ago v0.30.7-rc1 openai: align models list with tags ( #16556 ) 13 Ollama releases dev-tools 2mo ago v0.30.7-rc0 launch: use native Windows Hermes config path ( #16558 ) 5 Anthropic SDK (Python) releases dev-tools 2mo ago v0.106.0 0.106.0 (2026-06-05) Full Changelog: v0.105.2...v0.106.0 Features api: mark Claude Opus 4.1 as deprecated ( 85068cc ) Bug Fixes client: make Foundry client copy() and with_options() work ( 94146ac ) transform schema: preserve $defs when schema root is a $ref ( #1642 ) ( fc58e06… 19 Ollama releases dev-tools 2mo ago v0.30.6-rc0 launch: oh-my-pi ( #16410 ) 34 Ollama releases dev-tools 2mo ago v0.30.6 launch: oh-my-pi ( #16410 ) 21 r/LocalLLaMA community 2mo ago BeeLlama v0.3.1 – latest llama.cpp with extras! DFlash, MTP, q6_0 cache, TurboQuant. Single RTX 3090: Qwen 3.6 27B & Gemma 4 31B up to 177.8 tps (4.93x over baseline) BeeLlama v0.3.0 and v0.3.1 are here! Big architectural update to align the fork with upstream llama.cpp and integrate all its additions like MTP and Gemma 4 12B support, while also updating DFlash to handle complex configurations like multi-slot and multi-GPU. Now also… 5 Ollama releases dev-tools 2mo ago v0.30.5: launch: hermes-desktop app (#16516) Add support to launch the hermes-desktop app alongside the hermes agent from ollama launch. It will go through the install on first run if hermes-desktop is not already installed. 9 ComfyUI releases dev-tools 2mo ago v0.24.1 ComfyUI v0.24.1 8 Ollama releases dev-tools 2mo ago v0.30.5-rc0: llama.cpp version update (#16511) Bump llama.cpp to b9509, which includes the upstream Gemma 4 12B multimodal projector fixes for the n_head=0 divide-by-zero crash seen on x86/CUDA/Linux/Windows. Fixes #16479 Fixes #16489 Fixes #16491 Fixes #16492 Fixes #16495 11 r/LocalLLaMA community 2mo ago The first Gemma 4 12B finetunes are ready Now you can start building your Gemma 4 12B collection :) https://huggingface.co/igorls/gemma-4-12B-it-heretic-GGUF https://huggingface.co/ReadyArt/Melody1437-12B-v0.4-GGUF https://huggingface.co/DuoNeural/Gemma4-12B-IT-Abliterated-GGUF… 26 vLLM releases dev-tools 2mo ago v0.22.1rc2: fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init Signed-off-by: khluu [email protected] 9 vLLM releases dev-tools 2mo ago v0.22.1: fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init Signed-off-by: khluu [email protected] 28 OpenAI Python SDK releases dev-tools 2mo ago v2.41.0 2.41.0 (2026-06-03) Full Changelog: v2.40.0...v2.41.0 Features api: responses.moderation and chat_completions.moderation ( 87e46c2 ) 33 Ollama releases dev-tools 2mo ago v0.30.4-rc1: llama-server: fix gemma4 patch wiring (#16477) This will fix the "clip.cpp:4399: Unknown projector type" crash. 4 Ollama releases dev-tools 2mo ago v0.30.4: llama-server: fix gemma4 patch wiring (#16477) This will fix the "clip.cpp:4399: Unknown projector type" crash. 38 r/LocalLLaMA community 2mo ago Big Model Value Wars - DeepSeek V4 Pro vs MiMo-V2.5-Pro vs MiniMax M3 For those who sometimes boost their local model use with openrouter options, or the madlads who have the infrastructure to actually run those locally, it feels like those three model have the edge in best bang for your buck. How then do you decide which one to use? Do you have a… 19 Hacker News — AI on Front Page community 2mo ago Elixir v1.20: Now a gradually typed language Article URL: https://elixir-lang.org/blog/2026/06/03/elixir-v1-20-0-released/ Comments URL: https://news.ycombinator.com/item?id=48388324 Points: 252 # Comments: 71 34 Ollama releases dev-tools 2mo ago v0.30.4-rc0: Kill llama-server during Windows cleanup (#16458) Windows installer and app cleanup could leave llama-server.exe running when ollama.exe was killed directly, so cleanup now includes llama-server.exe and taskkill /T. 28 ComfyUI releases dev-tools 2mo ago v0.24.0 ComfyUI v0.24.0 32 Ollama releases dev-tools 2mo ago v0.30.3 models: add support for gemma4-12b ( #16457 ) 30 r/LocalLLaMA community 2mo ago How does the new abliteration tool Apostate compare with others? - Abliterlitics Why Qwen 2.5 7B? Apostate is a new abliteration tool by heterodoxin. He asked me to benchmark it. Qwen 2.5 7B was recommended by heterodoxin as it's the most tested model for Apostate. I abliterated the model with Heretic v1.3.0 and Apostate. The models are available on… 33 Hugging Face Daily Papers research 2mo ago PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training Abstract PaddleOCR-VL-1.6 enhances document parsing performance through targeted data optimization and progressive post-training techniques, achieving state-of-the-art results on OmniDocBench v1.6. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We introduce PaddleOCR-VL-1.6, an… 9 vLLM releases dev-tools 2mo ago v0.22.1rc1: [docker] Stop using extra-index-url for flashinfer-jit-cache (#44366) Signed-off-by: Kevin H. Luu [email protected] 34 Ollama releases dev-tools 2mo ago v0.30.2-rc0: fix laguna patch build breakage (#16445) Follow up to #16396 Fix kernel template instantiation so the symbols are exported in the library. 29 Ollama releases dev-tools 2mo ago v0.30.2: fix laguna patch build breakage (#16445) Follow up to #16396 Fix kernel template instantiation so the symbols are exported in the library. 38 Ollama releases dev-tools 2mo ago v0.30.1: llm: ignore llama-server SSE ping comments (#16443) llama.cpp b9478 added a default 30s SSE ping that emits colon-only comment frames (":\n\n") while streamed requests are idle; Ollama treated non-data SSE lines as JSON, so skip SSE comments in completion and chat streams. 36 Ollama releases dev-tools 2mo ago v0.30.1-rc0 launch: isolate Codex launch configuration ( #16437 ) 7 r/MachineLearning community 2mo ago Backpropagation destroys V1 brain alignment in one epoch, tracking RSA alignment to fMRI across training for BP, FA, predictive coding, and STDP [R] Third in a series of papers tracking learning rules vs. human fMRI (THINGS dataset, V1–IT, N=3 subjects). Previous finding: untrained CNNs match backprop at V1. This paper asks: when does training break that, and does the learning rule matter? Setup: RSA alignment measured at 8… 30 OpenAI Python SDK releases dev-tools 2mo ago v2.40.0 2.40.0 (2026-06-01) Full Changelog: v2.39.0...v2.40.0 Features api: Add Amazon Bedrock Responses support Bug Fixes api: allow setting bedrock api keys on the client directly ( 4d5bfde ) 19 Ollama releases dev-tools 2mo ago v0.30.0: launch: migrate Codex config (#16397) launch: migrate Codex config 30 OpenAI Python SDK releases dev-tools 2mo ago v2.39.0 2.39.0 (2026-06-01) Full Changelog: v2.38.0...v2.39.0 Features api: workload identity in audit logs, additional_tools item in responses, fix ActionSearch.query to be optional. ( ab60d7a ) 10 ComfyUI releases dev-tools 2mo ago v0.23.0 What's Changed feat: MediaPipe face detection (CORE-235) by @kijai in #14009 Multi-threaded load of models from disk (big load time speedups & Offload to disk) (CORE-43,CORE-152,CORE-164,CORE-165,CORE-117) by @rattus128 in #13802 Repo security stuff. by @comfyanonymous in #14019… 28 Ollama releases dev-tools 2mo ago v0.30.0-rc32: llama-server followups (#16353) llama-server followups Misc fixes for #16031 Add back dropped ROCm build flag for multi-GPU support on windows Fix amdhip64_*.dll version detection for "latest" selection Fix embeddings API for consistent normalize behavior with prior versions ci: set up for automated llama.cpp… 19 r/LocalLLaMA community 2mo ago mistral.rs v0.8.2: up to 2.8x faster CUDA inference than llama.cpp on GB10, B200, and H100 Hey all! I’ve been working on CUDA performance in mistral.rs, and v0.8.2 is focused on CUDA throughput. The result: on Gemma 4 (dense & MoE), mistral.rs is faster than llama.cpp at every point in my release sweep on GB10/H100/B200. See some results below on GB10 and B200:… 24 r/LocalLLaMA community 2mo ago Llama Studio v0.2.0 I have made an update to my llama-server WebUI based on some awesome feedback and interaction with the community. 1) JSON model config replaced by per-model shell scripts. Run from CLI, paste from unsloth, email to your buddy or post to reddit: Using real shell scripts to store… 17 Hacker News — AI on Front Page community 2mo ago The AV2 Video Standard Has Released (Final v1.0 Specification) Article URL: https://av2.aomedia.org Comments URL: https://news.ycombinator.com/item?id=48340910 Points: 203 # Comments: 80 34 r/LocalLLaMA community 2mo ago this new Moss tts 1.5 is damn good with voice cloning https://huggingface.co/spaces/OpenMOSS-Team/MOSS-TTS-v1.5 I prefer this over fish audio s2 pro because fish audio dont allow commercial use Long Cat DiT 3.5 is also a another good model.   submitted by   /u/9r4n4y [link]   [comments] 38 vLLM releases dev-tools 2mo ago v0.22.1rc0: [CI] Make Model Executor test hangs fail fast with a traceback (#43971) Signed-off-by: khluu [email protected] Co-authored-by: Claude [email protected] 10 llama.cpp releases dev-tools 2mo ago b9411 model : support for DeepseekV32ForCausalLM with generic DeepSeek Sparse Attention (DSA) implementation ( #23346 ) llama : support DeepSeek V3.2 model family (with DSA lightning indexer) convert : handle DeepseekV32ForCausalLM architecture ggml : support for f16 GGML_OP_FILL… 34 Ollama releases dev-tools 2mo ago v0.30.0-rc31 ci fix - non-shallow MLX checkout 29 Ollama releases dev-tools 2mo ago v0.30.0-rc30 version bump 18 Anthropic SDK (Python) releases dev-tools 2mo ago v0.105.2 0.105.2 (2026-05-29) Full Changelog: v0.105.1...v0.105.2 14 Anthropic SDK (Python) releases dev-tools 2mo ago v0.105.1 0.105.1 (2026-05-29) Full Changelog: v0.105.0...v0.105.1 Chores internal: use Trusted Publishing for PyPI releases ( 1d04fc5 ) 34 Ollama releases dev-tools 2mo ago v0.30.0-rc29 review comments 24 Anthropic SDK (Python) releases dev-tools 2mo ago v0.105.0 0.105.0 (2026-05-28) Full Changelog: v0.104.1...v0.105.0 Features api: Add support for claude-opus-4-8, mid-conversation system blocks, and usage.output_tokens_details ( f18b014 ) support custom file size caps ( #1825 ) ( 7e5f944 ) Chores examples: rename managed-agents… 12 Page 5 of 8 · 381 articles ← Newer Older →