News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow Hacker News — AI on Front Page community 10d ago FFmpeg 9.0 Article URL: https://github.com/FFmpeg/FFmpeg/blob/n9.0/RELEASE_NOTES Comments URL: https://news.ycombinator.com/item?id=49166202 Points: 219 # Comments: 41 21 llama.cpp releases dev-tools 10d ago b10258 llama : move n_vocab from llama_sampler_data to penalty_sampler ( #26520 ) This matches how it is done for logit_bias and mirostat samplers, see #25262 (comment) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled)… 13 Vercel — AI dev-tools 10d ago Vercel supports Next.js 16.3 Yesterday the Next.js team announced the release of Next.js 16.3 , with leaner prefetching, immutable static assets, and instant navigations. As part of this release, we worked with the Next.js team to fully support 16.3 on Vercel, including better performance and additional… 5 llama.cpp releases dev-tools 10d ago b10256 sycl: parallelize the non-contiguous concat kernel ( #25852 ) sycl: parallelize the non-contiguous concat kernel Launch geometry only: the non-contiguous concat kernel launched a single-lane work-group (1, 1, 1), now it will launch a (1, 1, SYCL_CONCAT_BLOCK_SIZE) one.… 30 Smol AI News news-outlet 10d ago not much happened today **Alibaba** launched **Qwen3.8-Max**, enhancing multimodal capabilities and agent ecosystem integration. **NVIDIA** introduced **Alpamayo 2 Super** for autonomous vehicle reasoning, while **Mistral AI** released **Shieldstral**, a 3B parameter open-weights safety model for… 17 llama.cpp releases dev-tools 10d ago b10254 chat : add new template for DeepSeek V4 Flash 0731 ( #26398 ) common/chat: update DeepSeek V4 templates Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. Default drop_thinking for DeepSeek V4 history so prior thinking is… 37 arXiv — NLP / Computation & Language research 10d ago Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification arXiv:2608.00581v1 Announce Type: new Abstract: Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set whose… 32 Vercel — AI dev-tools 10d ago Skill packs are now available on skills.sh You can now bundle multiple agent skills into a shareable pack on skills.sh . Hand anyone a curated set via a single URL, or share it with your GitHub organization to standardize the skills your team's agents use across any project. Create a pack from community skills on… 25 Latent.Space news-outlet 10d ago [AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork Qwen is so back! 9 r/LocalLLaMA community 10d ago can someone create a website where people share specific hardware specs with specific llama cpp flags so we see what works? title   submitted by   /u/Rank201AltAccount [link]   [comments] 20 r/LocalLLaMA community 10d ago More Qwen 3.8 sizes coming   submitted by   /u/appakaradi [link]   [comments] 14 Simon Willison community 10d ago Quoting Steve Yegge Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever… 19 Simon Willison community 10d ago Quoting Steve Yegge Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever… 33 OpenAI Python SDK releases dev-tools 10d ago v2.53.0 2.53.0 (2026-08-03) Features api: Add gpt-5.5 and tool name/namespace to Responses types ( #3569 ) ( dd1202d ) Bug Fixes ci: avoid NumPy source builds and duplicate HTTPX coverage ( #3573 ) ( b58332f ) 8 r/LocalLLaMA community 10d ago G9v3-39A5B: Agentic heavy MOE with low hallucination Hugging Face Artificial Analysis Should be a sweet spot for general work. Seems like coding is the only part that is inferior to Qwen.   submitted by   /u/axseem [link]   [comments] 16 r/LocalLLaMA community 10d ago Ling-3.0-flash is another potential model to test before qwen3.8 27b I tested Ling-3.0-flash with hard bugs and it fixed bugs that qwen3.6-27b could not. This models speed faster than deepseek v4 flash but almost the same level as (old) deepseek v4 flash. Note: hard bugs mean they don't have "error messages" but they are unexpected behaviors of a… 13 r/LocalLLaMA community 10d ago DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config Ran DeepSeek V4-Flash-0731 — the full official checkpoint, not a re-quant — on commodity used hardware. Sharing because I couldn't find anyone else publishing Ampere results for this engine. Why bother with a 2018 server The model is 156 GB . That number decides everything… 34 llama.cpp releases dev-tools 10d ago b10243 llama : allocate indexer cache only in "full" indexer layers ( #26474 ) Co-authored-by: Stanisław Szymczyk [email protected] Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS… 4 Don't Worry About the Vase community 10d ago OpenAI's Unreleased Model Astra Solves Ten Major Open Mathematics Problems Math is hard. 10 r/LocalLLaMA community 10d ago Qwen3.8-Max matches Kimi K3 and DeepSeek V4 Flash Qwen3.8-Max (2.4T) is another massive contribution to the open weight community. On benchmarks, it performs closely to Kimi K3 and DeepSeek V4 flash across all categories and is better at coding and software tasks. Qwen3.8-27B will also be open weight soon too. Weights are being… 11 Hacker News — AI on Front Page community 10d ago Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone Article URL: https://github.com/leonickson1/Swiftlet Comments URL: https://news.ycombinator.com/item?id=49158333 Points: 208 # Comments: 95 34 r/LocalLLaMA community 10d ago I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane! So this is the stuff of absolute insanity. In less than 20 months we've gone from super expensive cloud models only, to being able to run a Q3 quant of DeepSeek on an Intel Windows PC with a very average 24GB of VRAM. No wonder the big boys are panicking (and yes it's slow as… 10 r/LocalLLaMA community 10d ago All DeepSeek model oneshots: 242 outputs to look at and compare! Continuing my weekend of oneshotting the cheap OpenRouter models, here are all 10 DeepSeek models across the same 35 prompts. DeepSeek had a rougher time (more provider errors / empty completions), so only 242 made it out of the 10*35 matrix. Here they are… 31 r/LocalLLaMA community 11d ago Quantization hurts knowledge nonlinearly - Qwen3.6 27B case study   submitted by   /u/pmigdal [link]   [comments] 35 Interconnects (Nathan Lambert) research 11d ago Introducing our Artifacts Hub and Adoption Dashboard Scaling our curation and measurement of the open ecosystem. 14 r/LocalLLaMA community 11d ago model: MTP support for Qwen3-Next by yomaytk · Pull Request #25589 · ggml-org/llama.cpp Now we can run Qwen3-Next at “full speed” :) Do you still remember this model?   submitted by   /u/jacek2023 [link]   [comments] 26 r/MachineLearning community 11d ago Deep Dive on RL and OPD for Training LLMs [D] Hi everyone, if you have been reading the tech reports of Kimi, DS, Qwen and GLM, you will realize how much on policy distillation and GRPO style algorithms power the frontier. I thought it will be quite beneficial to do a deep dive explaining the maths and code behind this… 4 r/LocalLLaMA community 11d ago KAT Coder 2.5 dev: Do yourself a favor and try it! It is so good! I don't know why there aren't more people talking about it. Fewer tokens, faster and more accurate than Qwen 3.6 35b a3b. On my setup it's nearly as good as 27b, but 5x faster. And it completely trashes the Gemma 4 models. At least for my use case, it feels… 31 r/LocalLLaMA community 11d ago AI9Stars released G9v3-39A5B AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under the Apache 2.0 license making it fully open for personal and commercial use It… 10 llama.cpp releases dev-tools 11d ago b10238 model: MTP support for Qwen3-Next ( #25589 ) mtp for qwen3nex fix for python type-check Fix to compute num_mtp from directly mtp layer define opt_num_mtp_layers in _QwenMtpMixin and fix some comments Fix for python type check Update gguf-py/gguf/constants.py Co-authored-by:… 26 r/LocalLLaMA community 11d ago [RELEASE] SupraBrain-50M-v0.1 Hey there! So today we're releasing SupraBrain-50M, a hybrid language model that combines Gated DeltaNet linear recurrence with Sliding-Window Attention and Surprise-Gated update mechanisms to deliver very strong performance. Here are the benchmarks:… 19 llama.cpp releases dev-tools 11d ago b10237 llama : MTP support for DeepSeek V3.2 ( #26457 ) llama : MTP support for DeepSeek V3.2 model : no need to include MTP layers during DeepSeek V3.2 model type discovery Co-authored-by: Stanisław Szymczyk [email protected] Website: https://llama.app macOS/iOS: macOS Apple Silicon… 10 r/LocalLLaMA community 11d ago Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090 https://preview.redd.it/3zcvpbds14hh1.png?width=1911&format=png&auto=webp&s=a79aafb71eeca97638da93d2591902631e897fd5 I tried a test similar to the recent model-quant comparisons, but this time I focused only on: DeepSeek-V4-Flash-0731-IQ2_XS-Experts-Q8_0 model link… 16 OpenAI official-blog 11d ago How we built a realtime system for responsive voice AI in six months GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations. 16 r/LocalLLaMA community 11d ago I benchmarked classic vector RAG vs Google's new OKF format vs both combined — same corpus, same 7 questions, all local (Ollama + ChromaDB) Google Cloud published OKF (Open Knowledge Format) on June 12th — a spec for storing curated knowledge as a directory of markdown files with YAML frontmatter. One concept per file, linked to each other, with an index.md for progressive disclosure. The only required field is… 7 r/LocalLLaMA community 11d ago Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍   submitted by   /u/quantier [link]   [comments] 7 Smol AI News news-outlet 11d ago Qwen 3.8 Max **Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with… 21 Simon Willison community 11d ago condense-json 1.1 Release: condense-json 1.1 After shipping condense-json 1.0 I started integrating it into LLM, and found there were some desirable new features already: Replacements object can now include values other than strings. These will be identified and used as structural replacements by… 14 r/LocalLLaMA community 11d ago Can't wait to see Qwen3.8-27B Qwen announced Qwen3.8 a few hours ago, and it looks like we’re getting a new 27B model! Really excited to try this one locally.   submitted by   /u/FormOne2615 [link]   [comments] 34 r/LocalLLaMA community 11d ago Open weight has made to frontier Am looking forward to this! Open weight has come near frontier for 5x less the cost per/M tokens on task completion Open weight ranking Kimi k3 Qwen 3.8 GLM 5.2 Deepseek v4 flash 07/31   submitted by   /u/Specialized-Trap404 [link]   [comments] 4 arXiv — Machine Learning research 11d ago The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting arXiv:2607.29503v1 Announce Type: new Abstract: While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in… 24 arXiv — Machine Learning research 11d ago Geographically Weighted Surrogate Models for Rapid Small-Area Chronic Disease Estimation arXiv:2607.28655v1 Announce Type: cross Abstract: Small-area estimation (SAE) enables researchers and policymakers to identify spatial disparities in health outcomes, but survey-based SAE products carry an inherent lag. Gold-standard estimates such as CDC PLACES are released… 10 arXiv — NLP / Computation & Language research 11d ago Tokenizer-Agnostic Engram Module arXiv:2607.29065v1 Announce Type: new Abstract: Deepseek's Engram, a conditional memory module, was introduced to trade-off storage versus reasoning in large language models. However, the module relies on token-level $N$-gram hashing for Engram embedding lookup, introducing a… 31 r/LocalLLaMA community 11d ago Qwen3.8-27B announced alongside Qwen3.8-Max https://preview.redd.it/gy0tgokdl2hh1.png?width=540&format=png&auto=webp&s=7db9e034613a915cb33d378b99ad72c31c7cc18f source: https://x.com/Alibaba_Qwen/status/2084100707423289643   submitted by   /u/TKGaming_11 [link]   [comments] 29 Hacker News — AI on Front Page community 11d ago Qwen3.8-Max: A New Bar for Coding and Cowork Article URL: https://qwen.ai/blog?id=qwen3.8 Comments URL: https://news.ycombinator.com/item?id=49150470 Points: 212 # Comments: 74 36 r/LocalLLaMA community 11d ago Qwen 3.8 is live now. Update: And yes, Qwen3.8-27B is coming too. Next week! 2.4-trillion-parameter MoE flagship delivering a comprehensive leap in coding and professional work. Autonomously codes and delivers complete projects spanning 10+ days. Open weights coming soon! It is live at… 19 Simon Willison community 11d ago condense-json 1.0 Release: condense-json 1.0 I'm trying to get braver at releasing 1.0 versions. This little library is a year and a half old now - I've applied some sensible and non-disruptive fixes and shipped the big 1.0 for it. Here's an example of what it can do, lifted from the README: {… 29 r/LocalLLaMA community 11d ago You really should not quantize KV Cache for DeepSeek V4 Flash I don't think anyone should quantize the KV with DS4F. I checked the the quality impact (PPL, KLD, Same TopP) for swhitching from BF16 KV to Q8 KV, and it appears significant. Very much in contrast to Qwen 397B. Here are the results for DS4F: ====== Perplexity statistics ======… 11 r/LocalLLaMA community 11d ago PSA: llama.app, Mac app and llama serve from llama.cpp https://llama.app/ Been using llama.cpp for years now and im on here all the time (im a mod..), but somehow I totally missed that llama.app exists and its official from the HF/llama.cpp team. So posting this as I'm quite sure I'm not the only one in this boat. The llama.cpp team… 5 r/LocalLLaMA community 11d ago https://huggingface.co/poolside/Laguna-S-2.1-NVFP4 Updated release (August 2026). This is a new checkpoint that supersedes the earlier version of this repository. The weights have changed, not only the config, so if you downloaded a previous copy please re-download to pick up the current checkpoint.   submitted by  … 23 Page 8 of 10 · 500 articles ← Newer Older →