News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 2d ago 4-5 days replacing Claude w Qwen 3.8 Next Hi, human here with rambling thoughts to share. Feel free to skip Overall, I don’t feel like I’m missing much; if anything. On my hardware(M1 ultra w 128Gb) it’s probably not as fast as Claude but I’ve been using opencode for research and other business related tasks and it’s… 37 r/LocalLLaMA community 2d ago Qwen3.8-27B: Using KV Cache Transplants to Boost Output Quality Since my last post , I've been thinking about different options for dynamic performance degradation, trying to squeeze as much high-quality inference out of my GPU as I can. Over the weekend I read this really interesting paper: Cache-to-Cache: Direct Semantic Communication… 18 r/LocalLLaMA community 2d ago Swift1.5-Qwen3.8-Flash-Next is phenomenal vs. base 3.8-Flash! TL;DR - Swift Flash is a killer model that massively reduces excess reasoning. Try it out! If you haven't seen from my previous comparison posts , I'm a huge fan of the Swift Qwen3.8 models. I've been using 27B since it dropped, and I'm really impressed with the performance and… 15 OpenAI official-blog 2d ago Proaction boosts sales 60% and saves 75+ hours with Codex With Codex, GPT-Live-1, and GPT-6 Astra, Proaction builds, operates, and sells modern fleet management faster. 14 r/LocalLLaMA community 2d ago What I learned letting a local 27B run overnight long-horizon coding on my own rig Hi reddit, i know you hate AI slop so i indeed write the intro myself! iam dev and curios about local inference and long hoirzon coding on my own box. last day-ish i let my local model (qwen 27b on llama.cpp, 2x 16GB cards) go on a long coding tour inside deepseek harness while… 35 Hacker News — AI on Front Page community 2d ago Ollaya – Ollama for open-source, Jev-style decision models Article URL: https://ollaya.dev/ Comments URL: https://news.ycombinator.com/item?id=49848269 Points: 255 # Comments: 81 30 r/LocalLLaMA community 2d ago Trained my first small language model I have a tool that uses Gemini Flash with the lowest thinking budget to do summarization work. It's very fast, 0.9-1.2s in most cases. But I have a user experience problem where people make the wrong choice when using an internal app for the team. Gemini Flash can figure out… 12 TechCrunch — AI news-outlet 2d ago Meta’s Muse just stole the AI spotlight from OpenAI and Anthropic When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s… 24 r/LocalLLaMA community 2d ago Jev vs. Kev: open-source Jev alternative tested side by side We hosted Kev 4B (Jared Palmer's Apache-2.0 fine-tune of Qwen3.5-4B) and ran it side by side with Jev on the same endpoint to see how it compares. We built a fresh set of 362 items published after both models shipped (new arXiv papers, Stack Exchange questions, GitHub issues),… 9 llama.cpp releases dev-tools 2d ago b11183 metal : split fa kernels into per-dtype libraries ( #29329 ) metal : split fa kernels into per-dtype libraries Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp cont : minor fix comment Website: https://llama.app Attestations:… 11 TechCrunch — AI news-outlet 2d ago Astra and Opus just passed Turing’s other test Frontier AI models are finishing Alan Turing's World War II codebreaking work. 37 r/LocalLLaMA community 2d ago Make Volta Fast Again For those who have V100 cards, I wanted to point you to 1Cat-vLLM, a vLLM fork that enables optimized serving for these cards. Showing stats for Qwen3.6-35b comparing a Strix Halo with a hughly optimized llama.cpp fork (pwilkin) and the V100 with 1Cat. It’s not apples to apples,… 22 llama.cpp releases dev-tools 2d ago b11182 llama : add llama_prec_policy + model-driven W4A4 path ( #24364 ) Rebase and update based on #26675 Signed-off-by: ynankani [email protected] CI failure fix(launh_bounds overload on HIP) and cleanup Signed-off-by: ynankani [email protected] Address review comments… 10 TechCrunch — AI news-outlet 2d ago Meta’s AI Tamagotchi bet is…working? When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s… 21 r/LocalLLaMA community 2d ago How do you use subagents & multiple agent with local models, and how many? Running qwen3.8 27b nvfp4 on vllm at max context only gives around 8 agents with 32k context each. That doesnt seem like much; what use cases do people use multi-agent frameworks and find it helpful for?   submitted by   /u/Ambitious_Fold_2874 [link]   [comments] 15 LangChain releases dev-tools 2d ago langchain-fireworks==1.6.3 Changes since langchain-fireworks==1.6.2 release(fireworks): 1.6.3 ( #40834 ) fix(fireworks): declare native PDF inputs unsupported ( #40814 ) fix(fireworks): preserve malformed tool arguments as diagnostic JSON ( #40818 ) chore(model-profiles): refresh model profile data (… 27 The Information — AI news-outlet 2d ago Microsoft Launches Revamped Copilot ‘Super App’ with Muse Competitor Microsoft on Friday launched a revamped version of its Copilot app that combines features that had previously been sold separately, including AI coding tools, features that automate tasks in Office 365, and an always-on “Autopilot” agent similar to Meta’s Muse. The overhaul is… 28 r/LocalLLaMA community 2d ago Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity I’ve been experimenting with whether Qwen3.8-Flash-Next’s pretrained PLE n-gram memory can improve a much smaller Qwen3.5-0.8B model. I trained the 0.8B setup with limited resources, mostly using free Kaggle notebook GPUs. The setup keeps both the Qwen3.5-0.8B backbone and the… 23 r/LocalLLaMA community 3d ago Did anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some I let Codex do some eval on Qwen 3.8 Flash Next IQ4_XS (served via vllm and r9v) and Qwen 3.8 FP8 (served via vllm and radiance). Here are the results: Benchmark Flash-Next IQ4_XS Qwen3.8 27B FP8 Result MMLU-Pro 83.8% 75.0% Flash-Next GPQA Diamond 42.5% 27.5% Flash-Next GSM8K… 8 llama.cpp releases dev-tools 3d ago b11177 CUDA: fuse RMS_NORM + SCALE into one kernel ( #29393 ) #28068 builds the GDN q/k l2norm as ggml_scale(ggml_rms_norm(x, eps/n), 1/sqrt(n)). This adds 2 SCALE nodes per GDN layer, 96 extra kernel launches per ubatch on Qwen3.8-27B (48 GDN layers). The extra kernels take no… 38 ThursdAI news-outlet 3d ago Opus 5.5 is your new workhorse! OpenAI ships GPT 6 Sol and Luna before DevDay and Meta goes all in on Muse! Your friday read is here From CoreWeave: Opus 5.5 is back and 40% cheaper, GPT-6 halves prices, Grok orders Starbucks from your Tesla, and Google clones your voice in 30 seconds 11 arXiv — Machine Learning research 3d ago When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection arXiv:2609.29580v1 Announce Type: new Abstract: A released table is often treated as an i.i.d. sample, although its repeated rows may encode business frequency, repeated entities, joins, resampling, or extraction errors. We show that this ambiguity creates a hidden measurement… 35 arXiv — NLP / Computation & Language research 3d ago YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech arXiv:2609.29448v1 Announce Type: new Abstract: We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date,… 30 arXiv — NLP / Computation & Language research 3d ago Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases arXiv:2609.30012v1 Announce Type: new Abstract: Language models advise people, keep them company, and write software while they sleep. Measuring what they do is hard: behavior has to be sampled repeatedly across models, prompts and releases, most of it lives in unstructured text… 23 llama.cpp releases dev-tools 3d ago b11176 llama : fix tensor split for fused qkv with uneven K/V head sizes ( #2 … 6 r/LocalLLaMA community 3d ago CachyOS Qwen 3.8 27b on 2x 5090 vLLM   submitted by   /u/piddlefaffle12 [link]   [comments] 7 The Information — AI news-outlet 3d ago Why Meta’s Moment in the AI Sun Won’t Last If Meta Platforms wants to raise money by selling equity, now is the time. Meta shares have gone on an incredible run lately, jumping 27% since Sept. 8, when the company released its Muse personal AI agent. Meta stock is now up 18% so far this year, a bigger gain than that for… 21 Vercel — AI dev-tools 3d ago Pixel Canary is now available in stealth for free on AI Gateway Pixel Canary is now available on AI Gateway as stealth/pixel-canary . It's free for a limited time while in stealth. Pixel Canary is strong at coding, including building applications and refactoring existing code. It is well suited to frontend development and mobile app design,… 8 llama.cpp releases dev-tools 3d ago b11172 metal : optimize sparse FA + clean-up ( #29377 ) metal : cache sparse FA indices in shared memory Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp metal : simplify shared memory size calculation Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp pi : update general… 7 r/LocalLLaMA community 3d ago Is Qwen Flash Next at like Q2 better than 27B at Q4? I know questions like this are asked often but I didn’t see this specific one   submitted by   /u/Borkato [link]   [comments] 19 r/LocalLLaMA community 3d ago 7900 XTX — two "low-thinking" Qwen 3.8 27B quants (Swift + ThinkingCap) vs the regular quant First, do they actually produce less tokens? Yes. Total tokens per benchmark run (4 scenarios): base quant ~66k, ThinkingCap ~49k (−26%), Swift ~45k (−33%). So the "less thinking" is real — and Swift cuts the most. Then the cost: and this is where it got interesting. The two… 29 r/LocalLLaMA community 3d ago Qwen3.8 27b practical modeling for 3d printing I spent the past day and a half trying to get qwen27b to complete some practical work for me. I have a Bambu h2c I have been wanting to get more use out of so thought this would be a fun experiment. I have 27b running on my 5090 and qwen image 2.1 running on a 3080 10gb with… 18 The Information — AI news-outlet 3d ago Databricks Acquires Microsoft Excel Competitor Row Zero Databricks is acquiring Row Zero, a five-year-old Microsoft Excel competitor that sells spreadsheet software that’s designed for analyzing massive amounts of data. The deal could help Databricks attract more non-technical users to Genie, a chatbot it launched last summer that… 17 r/LocalLLaMA community 3d ago Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram? I was actually pretty happy with my Qwen3.8-27b setup, and I'd been tinkering with Ninfer to have a version that was "fast but maybe a bit stupid" and the speed was nice to have as a backup. But I was curious how the Flash-Next version might work, after I learned it didn't need… 7 r/LocalLLaMA community 3d ago Qwen 3.8 27b be like... The user is frustrated — I rambled too much and didn't act. Let's just run the test suite and move on. No more forensics. One command, execute, then report. (Original memo is a casual internal monologue in English. Translating faithfully while preserving the informal,… 9 r/LocalLLaMA community 3d ago Dailychained PLX 88096 switches, Quad RTX 5070 Ti + Quad RTX 5060 Ti Left: Quad RTX 5060 Ti, Middle: Quad RTX 5070 Ti, both PLX 88096 switch, Right = rehomed host Edit: Benchmarsk were 1 line = fixed WHY: -->> DATA SOVEREIGNITY / PRIVACY<<-- hey, this is (localllama right?), this makes no less sense than my dropping the same $$$ on a motorbike I… 14 Hacker News — AI on Front Page community 3d ago Opus 5.5 is good at explainer videos Article URL: https://launchvideo.io Comments URL: https://news.ycombinator.com/item?id=49836374 Points: 277 # Comments: 140 8 r/LocalLLaMA community 3d ago I'm new and it's kinda overwhelming to get into Hi, sorry if this doesn't belong here. Getting to the point basically, I've been using online-only AI like GPT/Gemini since 2022, and have been interested in local models but am clueless overall. Yes I'm extremely late. I only use laptop (I'm a student), and I currently own:… 38 Simon Willison community 3d ago commit-rewriter 0.2 Release: commit-rewriter 0.2 Support for branches other than the default branch. Use uvx commit-rewriter --branch other to run against another branch. #3 Tags: git 35 r/LocalLLaMA community 3d ago FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face from FreedomIntelligence: HuatuoGPT-3-27B is a medical LLM built on Qwen3.8-27B with One-stage Policy Optimization (OnePO) . OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher… 12 Simon Willison community 3d ago datasette 1.0a41 Release: datasette 1.0a41 Alec Garcia added support for OpenTelemetry to Datasette in this release. I've also refactored all of Datasette's modal dialogs to a single Web Component, which is now documented for other plugins to use . Tags: javascript , datasette , web-components ,… 12 LangChain releases dev-tools 3d ago langchain-core==1.6.5 Changes since langchain-core==1.6.4 release(core): 1.6.5 ( #40816 ) fix(core): abbreviate long tool IDs in XML buffer strings ( #40792 ) 37 r/LocalLLaMA community 3d ago Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~430 tok/s… 38 TechCrunch — AI news-outlet 3d ago Google Photos ‘Clueless’-inspired virtual closet is now available on Android and iOS The AI-powered feature builds a virtual wardrobe from your photos, and is now broadly available after first rolling out to Android users in June. 11 r/LocalLLaMA community 3d ago UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy Hey everyone, Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens related to pathological overthinking patterns and restoring accuracy via RL (GSPO) and OPD . After amazing feedback and 350k+… 26 Google DeepMind official-blog 3d ago Introducing Gemini 3.8 Live with Live Avatar Introducing Gemini 3.8 Live with Live Avatar Sep 24, 2026 | x.com Facebook LinkedIn Mail Gemini 3.8 Live with Live Avatar brings real-time visual presence to Gemini’s conversational AI. By natively coupling our live dialogue capabilities with low-latency streaming video, Live… 16 Ars Technica — AI news-outlet 3d ago Google's first Suncatcher orbital data center test launches October 1 Google's experimental orbital data center will have four TPUs and only run for 15 minutes at a time. 21 r/LocalLLaMA community 3d ago ThinkingCap 3.8-27B vs. Swift 3.8-27B vs. Qwen 3.8-27B Benchmarks With the release of ThinkingCap-Qwen3.8-27B , I thought it would be worthwhile to do a comparison between the original Qwen3.8-27B, the new ThinkingCap, and Swift-Qwen3.8-27B . Both Swift which I already reviewed , and ThinkingCap do exactly the same thing: they reduce the… 8 TechCrunch — AI news-outlet 3d ago Google tests letting Gemini call businesses for you Google says the AI-calling feature will first be available to Pixel 11 owners in the U.S. who pay for a Gemini subscription. 38 Don't Worry About the Vase community 3d ago AI #187: Coming Into Play Opus 5.5 was released on Tuesday. 29 Page 2 of 10 · 500 articles ← Newer Older →