News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow Simon Willison community 7d ago datasette 0.65.3 Release: datasette 0.65.3 Back-ported the SQL Injection security fix from 1.0a38 . Tags: datasette 29 Simon Willison community 7d ago datasette 0.65.3 Release: datasette 0.65.3 Back-ported the SQL Injection security fix from 1.0a38 . Tags: datasette 16 r/LocalLLaMA community 7d ago KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0 , fork of llama.cpp with more KV cache quantization options. Models: Qwen 3.6 27B Q5_K_S 64k context, Gemma 4 31B Q5_K_S 16k context Standard quants, extended:… 9 Hacker News — AI on Front Page community 7d ago Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users Article URL: https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/ Comments URL: https://news.ycombinator.com/item?id=49199357 Points: 233 # Comments: 174 34 r/LocalLLaMA community 7d ago 2 x 5070ti Qwen 27B full config / stats Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-nightly. Which has the KV cache connector fixes and performance improvements.… 34 Hacker News — AI on Front Page community 7d ago Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks Hey HN, we’re Will & Johnny from ProvenMetal ( https://provenmetal.com ). You send us design files or specs and we give you assembled boards domestically in days. The US produced 30% of PCBs globally in 2000, now they produce 4%. Chinese manufacturers have completely dominated… 21 r/LocalLLaMA community 7d ago How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a benchmarking issue? EDIT: The contribution of this benchmark to the… 28 TechCrunch — AI news-outlet 8d ago Google Maps adds agentic features, including food ordering and hotel bookings The launch of these new features reflects Google’s ambitions to transform Google Maps from a navigation tool into an assistant that's capable of helping users complete real-world tasks. 19 r/LocalLLaMA community 8d ago They almost catched up on Frontier performance, so now catching up on prices Users also report that the free version was significantly downgraded after the release of the new models this is very important for us when considering local hosting. A lot of people decided not to buy expensive hardware because DeepSeek’s prices made it very difficult to break… 25 r/LocalLLaMA community 8d ago Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090) TL;DR: On a Qwen3.6-35B-A3B Q6 setup sized for 64K context on a 24GB RTX 3090, spilling eight MoE expert layers to CPU freed enough VRAM to increase -b from 512 to 1024 and -ub from 128 to 512. Prompt processing improved by 2.36×, while generation speed remained unchanged within… 9 r/LocalLLaMA community 8d ago Final optimization: from ~10 tok/s to ~15 tok/s on DeepSeek-V4-Flash-0731 at 128K ctx - 1 RTX 3090 J'ai consacré beaucoup de temps à l'optimisation de DeepSeek-V4-Flash-0731 GGUF sur une seule RTX 3090. Mon exigence absolue pour chaque configuration était la suivante : Le modèle doit rester utilisable avec une fenêtre de contexte de 128 000 jetons. J'ai testé les différentes… 4 OpenAI official-blog 8d ago Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna. 29 r/LocalLLaMA community 8d ago Best llama cpp flags to run Deepseek-flash 0731 Hi all. These are my system specs: dual xeon e5 2696 v2 , 160gb DDR3 ram ECC(1600mhz), 3 gpus: 3060 12gb, p100 16gb, 3050 6gb. And a 400gb nvme sdd RAID0, 3000 mb/s. The model is Deepseek-flash-0731 UD_8_X_XL, loseless, 161gb. Now, I'm not too knowledgeable about llama cpp… 12 r/LocalLLaMA community 8d ago Qwen3.8-2.4T-A95B (aka Qwen3.8-Max) open release time: next wednesday https://preview.redd.it/9zwlpcushphh1.png?width=972&format=png&auto=webp&s=18cb49c738caba9799d22d4916337c772baa830f https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B   submitted by   /u/HugeConsideration211 [link]   [comments] 17 r/LocalLLaMA community 8d ago I get that AI labs need to make money, but zero-warning price spikes are a nightmare for production builds Seen a ton of posts today about the DeepSeek API price hike. Half the feed is doom-posting, the other half is explaining basic GPU economics. Honestly, I get the cost side. Sub-cent tokens were never gonna last forever. But what actually sucks is the zero-day notice. Dropping a… 23 OpenAI official-blog 8d ago Working with the American Psychological Association on youth mental health and AI OpenAI and the APA are launching a three-year partnership to develop guidance, resources, and safeguards for responsible AI use supporting youth mental health. 12 Smol AI News news-outlet 8d ago not much happened today **Meta's Muse Spark 1.2** rapidly rose to frontier-tier with top 5 ranking on Vals Index at **$0.69/test**, being **3x cheaper than Kimi** and **10x+ cheaper than Fable, Opus, and 5.6 Sol**. It achieved **gold-medal-level performance in five STEM Olympiads** with perfect theory… 20 r/LocalLLaMA community 8d ago GLM/Qwen Appreciation Post https://preview.redd.it/o6ik6qboeohh1.png?width=1134&format=png&auto=webp&s=4016f26c50c1d93bd3d0c7e880e9b55a2d75310f I have been running Qwen3.6 27b for a little while (mostly coding tasks) and recently trying out V4 flash 0731 in it's place. It was very apparent the new v4… 29 arXiv — NLP / Computation & Language research 8d ago Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv:2608.04899v1 Announce Type: new Abstract: Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extremely sparse outputs. For instance, Qwen3-32B… 27 arXiv — NLP / Computation & Language research 8d ago AI Literacy for Legal Translation: Developing Digital Resilience arXiv:2608.04641v1 Announce Type: cross Abstract: Generative AI is transforming legal translation by introducing opportunities alongside linguistic, technical, legal, ethical and cognitive risks. This chapter examines the implications of AI for professional legal translation and… 4 r/LocalLLaMA community 8d ago How many people in this sub try to train their own AI from scratch on their systems just for fun and to test out techniques from research papers? As for me, I own a system with an RTX 5090, Ryzen 9 9950X3D2, and 64 GB of DDR5. Every time I see research come out with a new way to train AI, I immediately think to try it on my system to see the results I get. Applying things like Titans, that one Deepseek paper on engrams,… 26 r/LocalLLaMA community 8d ago Introducing BetterBench - more accurate PP and TPS measurement I built this because the existing benchmarks were using random data and with MTP content types can vary a lot on what performance you see. 5% or more with content types. BetterBench is designed to have content consistency within 1% and also measures across different content… 8 Vercel — AI dev-tools 8d ago Introducing Agent Plugins Today, Agent Plugins 1.0.0 is publicly available. Agent Plugins is an open, vendor-neutral standard for plugins that extend AI agents. Agent Skills provide reusable instructions and resources for AI agents. MCP servers connect agents to tools and services. Both can be reused… 33 Vercel — AI dev-tools 8d ago Introducing Agent Plugins 1.0.0 Agent Plugins 1.0.0 is now available. It is an open, vendor-neutral standard for packaging Agent Skills and MCP servers into portable plugins. Compatible agent clients can discover and load them. Agent Plugins defines a common format: a root plugin.json manifest, plus fixed… 7 Vercel — AI dev-tools 8d ago Ling 3.0 Tiny is now available on AI Gateway Ling 3.0 Tiny from ANT Group is now on AI Gateway, free to use till 8:00am PT on 8/14. Ling 3.0 Tiny takes the free slot from Ling 3.0 Flash . Ling 3.0 Tiny is a MOE model with 7.9B total parameters and about 1.3B active per token, a 256K token context window, and up to 32K… 6 Vercel — AI dev-tools 8d ago Seedance 2.5 now available on Vercel AI Gateway Seedance 2.5 from ByteDance is now available on AI Gateway. It generates up to 30 seconds in a single clip, holding camera movement and continuity without stitching shots together in post. Short clips can also be extended with character, scene, and camera movement carried over.… 33 Simon Willison community 8d ago Introducing Muse Code and Muse Spark 1.2 Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as part of getting that to work! Muse Spark 1.2 is a coding-focused update to Muse… 11 Simon Willison community 8d ago Introducing Muse Code and Muse Spark 1.2 Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as part of getting that to work! Muse Spark 1.2 is a coding-focused update to Muse… 27 r/LocalLLaMA community 8d ago I remember a time when 'flash' meant 32B I mean, Deepseek V4 Flash is an absolutely fantastic model, even though I can't run it on my machine it's so fascinating to see how it performs. Knowing that potentially it could be run at home is really motivating and makes me hopeful that those capabilities will trickle down… 6 TechCrunch — AI news-outlet 8d ago Meta launches Muse Code, an AI agent for large code bases Meta expanded its AI coding offerings with a new agent that, it promises, can handle complex tasks with complex software. 23 r/LocalLLaMA community 8d ago Deepseek V4 Flash just hit Colibri, does anyone have numbers? I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older hardware will likely be slow - no Unsloth GGUF supported either. I'd be curious what people are getting with V100s, R9700s,… 8 r/LocalLLaMA community 8d ago Xiaomi-Robotics-1: New robotics model released Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipulation in unseen environments and efficient adaptation to new tasks. XR-1… 14 Simon Willison community 8d ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 29 Simon Willison community 8d ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 30 TechCrunch — AI news-outlet 8d ago Jeff Dean and other top AI researchers are leaving Google to launch their own startup The legendary Google executive is joined by other outgoing Google execs in a joint mission to use AI to push forward the process of scientific discovery. 36 Hacker News — AI on Front Page community 8d ago Muse Code and Muse Spark 1.2 Article URL: https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2 Comments URL: https://news.ycombinator.com/item?id=49187575 Points: 203 # Comments: 120 33 LangChain releases dev-tools 8d ago langchain-anthropic==1.5.4 Changes since langchain-anthropic==1.5.3 release(anthropic): 1.5.4 ( #39277 ) fix(anthropic): handle tool schemas with unsupported top-level composition ( #39273 ) chore: bump the minor-and-patch group across 3 directories with 7 updates ( #39187 ) fix(anthropic): preserve… 15 r/LocalLLaMA community 8d ago Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory A follow up to the launch of Mference , it now supports and runs Inkling-Small 276B-A12B . Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B total, ~12B active, 3.4 GB resident set , ~148 GB on disk. Measured on my M5,… 6 r/LocalLLaMA community 8d ago Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM Hey everyone! Scenema Audio is now a native ComfyUI custom node. Same model that powers scenema.ai now quantized so it fits on 8GB VRAM. When we first released it a few months ago as an API and Docker stack, the full precision transformers were too heavy for most people to… 6 Hacker News — AI on Front Page community 8d ago Beating GPT-5.6 Sol on retrieval with 100x cheaper open models Article URL: https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency Comments URL: https://news.ycombinator.com/item?id=49186762 Points: 226 # Comments: 47 29 r/LocalLLaMA community 8d ago Gemma 4 31b AttnRes Project I had Claude re-draft this for me, thus it has Em Dashes. It's correct with lots of "Claude" simplifications. --- Hey all. It's been a while since I posted about the AttnRes architecture so I figured I'd give an update on where things are. Short version: it's alive. Longer… 11 r/LocalLLaMA community 8d ago Mistral Releases premier Not-Hotdog model https://huggingface.co/mistralai/Shieldstral-1.0-3B   submitted by   /u/rpiguy9907 [link]   [comments] 35 r/LocalLLaMA community 8d ago DeepSeek V4 Flash 0731 at 10–17 t/s (nothink) on MacBook M5 Pro **64GB***, partly via SSD streaming Inspired by a post from u/giveen I motivated claude (no patinence on my side to work through everything myself) to help me get DS running on my MacBook M5 Pro 64GB and it exceeded my expectations.. because it worked, and at a quite usable generation speed! background: antirez… 12 r/MachineLearning community 8d ago Running Whisper, Qwen3-ASR, Nemotron & MOSS completely offline on iPhone [P] Over the past month, I've been building LiveTranscriber, an open-source iOS app for running modern speech and language models entirely on-device. The goal was to see whether recent open-source models could be turned into a practical mobile product—not just technical demos.… 8 r/LocalLLaMA community 8d ago Ling-3.0-flash MXFP4 released and running locally on one DGX Spark. In tests: ~80 tok/s decoding 2,500–3,500 tok/s long-input prefilling Smooth use by 3–4 concurrent users Private, on-device inference for coding, agents, and offline batch jobs   submitted by   /u/niacolhealth [link]   [comments] 30 Ars Technica — AI news-outlet 8d ago Google plans to kill Assistant on your phone on September 4 Assistant will disappear, leaving only Gemini for voice control in the coming weeks. 15 llama.cpp releases dev-tools 8d ago b10285 mtmd: support multi-row batching for deepseek-ocr ( #26154 ) mtmd: support multi-row batching for deepseek-ocr mtmd: weave deepseek-ocr rows in one shot instead of per row ( #26615 ) Co-authored-by: Saba Fallah [email protected] Website: https://llama.app macOS/iOS: macOS… 15 r/LocalLLaMA community 8d ago jabbatheduck/DeepSeek-v4-flash-mini · Hugging Face Because why not? How far can we go and make DeepSeek work?   submitted by   /u/giveen [link]   [comments] 6 r/LocalLLaMA community 8d ago MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE expert weights of the first N layers in system RAM and multiply them on the CPU;… 29 Hugging Face Daily Papers research 9d ago Multi-Task Multi-Frame Visual Piano Transcription Abstract Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT)… 38 Page 6 of 10 · 500 articles ← Newer Older →