News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow llama.cpp releases dev-tools 1h ago b10425 sycl: fuse the gated-delta-net state writeback cpy ( #26643 ) Port of #23940 . Arc Pro B70, Qwen 3.6 27B Q4_K - Medium (48 of its 64 blocks run gated_delta_net), -ngl 99 -fa 1 -ctk f16 -ctv f16 -b 2048 -ub 2048, interleaved A/B passes of r=3: tg128 23.91 / 23.90 / 23.90 -> 24.19… 26 Latent.Space news-outlet 2h ago [AINews] Gemini 3.7 Flash brings GDM back to the forefront Down, but not out! 6 r/LocalLLaMA community 2h ago GLM 5.3 Released Official Announcement https://z.ai/blog/glm-5.3   submitted by   /u/jmorant555 [link]   [comments] 37 Simon Willison community 8h ago sqlite-utils 4.2.1 Release: sqlite-utils 4.2.1 Fixes a crashing bug in sqlite-utils 4.2 . I'd introduced code that looks like this: from typing_extensions import Self It turned out the typing-extensions package was not listed as a dependency for sqlite-utils - it was installed by one of the other… 17 Ollama releases dev-tools 9h ago v0.32.11 launch: add DeepSeek Harness integration ( #17733 ) 23 r/LocalLLaMA community 9h ago Is waiting for Qwen 3.8 27B like waiting for Star War Episode one? Is waiting for Qwen 3.8 27B like waiting for Star War Episode one? I'm sweating waiting to get my hand on this to try it tomorrow morning. But it takes me back to Star Wars 1 and the disappointment after being so hyped to see it. Only 16 hours and 46 minutes to go... 45, ...… 22 llama.cpp releases dev-tools 9h ago b10419 OpenVINO: Qwen3.5, memory optimization, and test-recurrent-state-rollback ( #26952 ) OpenVINO backend: 1) enable gpt-oss moe on OV bk; 2) enable mxfp4 support OpenVINO backend: disable TOPK_MOE op test OpenVINO Backend: Add op FILL support OpenVINO backend: enable set rows with… 36 llama.cpp releases dev-tools 11h ago b10417 chat : fix LFM2 tool call arg name prefix ambiguity ( #26960 ) Assisted-by: Claude Opus 5 Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu… 6 r/LocalLLaMA community 11h ago Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release Qwen just released their first 3.8 model. The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting reasoning_effort to xhigh , medium , or low . However, the official template still has some serious problems: You cannot… 14 Simon Willison community 11h ago sqlite-utils 4.2 Release: sqlite-utils 4.2 Lots of improvements in this one relating to the table.transform() feature , which adds support for complex alter table operations by creating a fresh table, copying across the data and then dropping and replacing the old one. transform() now preserves… 13 r/LocalLLaMA community 12h ago Trained a 1.5B to write shell commands so I'd stop googling tar flags. Runs on a laptop CPU in ~1 sec. I've been googling "tar extract gz" for about ten years. and I finally did something about it. It started out as a research project and I ended up with a Fine-tuned Qwen2.5-Coder-1.5B on 125k natural-language/command pairs, merged and quantized to Q4_K_M. 941MB which runs… 16 TechCrunch — AI news-outlet 12h ago OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed OpenAI is launching a preview of a sped up version of its latest, most powerful model, in an effort to court enterprise users. 13 r/LocalLLaMA community 13h ago EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s I managed to get Qwen3.8-2.4T-A95B running locally with llama.cpp on mu PC just for fun, cause why not. I was using the Unsloth Qwen3.8-2.4T-A95B-UD-Q1_0 GGUF quantization. The full GGUF is about 397 GiB . The model uses 512 routed experts, with 10 active per token. My hardware:… 31 llama.cpp releases dev-tools 13h ago b10414 metal : add TQ2_0 support ( #26980 ) metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 cont : optimize mul_mv kernel float ops over integer ops precalculate sums… 29 Hacker News — AI on Front Page community 13h ago Accelerating GPT-5.6 Sol Ultrafast Article URL: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai Comments URL: https://news.ycombinator.com/item?id=49289844 Points: 250 # Comments: 87 6 Hacker News — AI on Front Page community 14h ago Gemini 3.7 Flash https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flas... Comments URL: https://news.ycombinator.com/item?id=49289112 Points: 250 # Comments: 169 21 r/LocalLLaMA community 14h ago MiniMax-Music3 released!   submitted by   /u/Acceptable-Cycle4645 [link]   [comments] 20 Hacker News — AI on Front Page community 14h ago Mistral OCR 4.1 Article URL: https://docs.mistral.ai/models/ocr-4-1 Comments URL: https://news.ycombinator.com/item?id=49288889 Points: 261 # Comments: 103 38 Google DeepMind official-blog 14h ago Introducing Gemini 3.7 Flash Introducing Gemini 3.7 Flash Aug 13, 2026 | x.com Facebook LinkedIn Mail Our most intelligent workhorse model yet for coding and agents. Tulsee Doshi Senior Director, Product Management, on behalf of the Gemini team Share x.com Facebook LinkedIn Mail Listen to article… 24 Ars Technica — AI news-outlet 14h ago Google announces Gemini 3.7 Flash just three weeks after previous release Gemini 3.6 Flash debuted just 3 weeks ago, but Google says 3.7 has "substantial improvements." 21 r/LocalLLaMA community 15h ago deepseek-ai/DeepSeek-V4-Pro-0813 (Available again) · Hugging Face   submitted by   /u/panchovix [link]   [comments] 30 r/LocalLLaMA community 16h ago unsloth/DeepSeek-V4-Pro-0813-GGUF · Hugging Face uploading...I think   submitted by   /u/mossy_troll_84 [link]   [comments] 14 Ars Technica — AI news-outlet 17h ago Anthropic could be worth $2 trillion when it goes public Rapid revenue growth fuels hope Claude maker's IPO is the biggest listing in history 22 r/LocalLLaMA community 18h ago Deepseek Harness is Up! DeepSeek Harness (dsh) is an open-source agent harness developed by DeepSeek AI. It uses an architecture where everything is a plugin, and is powered by Cordis, whose design is described in A Programming Paradigm for Spatiotemporal Composability. DeepSeek Harness is currently in… 13 r/LocalLLaMA community 18h ago Deepseek new pricing https://api-docs.deepseek.com/quick_start/pricing/   submitted by   /u/Comfortable-Rock-498 [link]   [comments] 34 MIT Technology Review — AI news-outlet 18h ago Flock is tightening its rules in response to a growing surveillance backlash The police-tech giant Flock is announcing today that it will change officers’ access to its nationwide network of license plate readers, in an apparent effort to quell a growing backlash and win back contracts lost amid concerns about mass surveillance and police abuse. Several… 24 r/LocalLLaMA community 18h ago GitHub - deepseek-ai/deepseek-harness 🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one… 32 LangChain releases dev-tools 18h ago langchain-openai==1.5.0 Changes since langchain-openai==1.4.3 release(openai): 1.5.0 ( #39629 ) feat(openai): support openai 3.0 SDK ( #39613 ) chore(partners): bump langgraph floor in openai and huggingface lockfiles ( #39617 ) 25 Hacker News — AI on Front Page community 18h ago DeepSeek Harness Article URL: https://github.com/deepseek-ai/deepseek-harness Comments URL: https://news.ycombinator.com/item?id=49285244 Points: 311 # Comments: 135 24 Hacker News — AI on Front Page community 18h ago DeepSeek Harness developer preview https://github.com/deepseek-ai/deepseek-harness https://deepseek-harness.github.io/deepseek-harness/en/guide... Comments URL: https://news.ycombinator.com/item?id=49285244 Points: 369 # Comments: 173 28 r/LocalLLaMA community 19h ago deepseek-ai/DeepSeek-V4-Pro-0813 · Hugging Face   submitted by   /u/mossy_troll_84 [link]   [comments] 8 r/LocalLLaMA community 19h ago DeepSeek: We’re launching DeepSeek-V4-Pro today! From DeepSeek on 𝕏: https://x.com/deepseek_ai/status/2087864585504305397   submitted by   /u/Nunki08 [link]   [comments] 28 r/LocalLLaMA community 20h ago The Qwen team is going live!   submitted by   /u/-MaskNinja- [link]   [comments] 16 Ars Technica — AI news-outlet 20h ago Claude's new Scarlet Letter watermark is invisible — for now The mark flags anything Claude processed, even human writing it only edited. 16 OpenAI official-blog 20h ago The builder’s guide to GPT‑5.6 Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabilities. 22 r/LocalLLaMA community 21h ago Qwen/Qwen3.8-27B · Official Countdown · Hugging Face   submitted by   /u/paf1138 [link]   [comments] 6 OpenAI official-blog 21h ago Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output tokens per second. 24 r/LocalLLaMA community 23h ago Minimax Music 3 open weight release soon? Diffusers has a PR with deets: https://github.com/huggingface/diffusers/pull/14456 Minimax is working on this repository right now and put up a bunch of samples: https://github.com/MiniMax-AI/music3-demo/tree/main/assets/audio/tracks Comfy-Org is teasing about a big release in… 30 r/LocalLLaMA community 1d ago The countdown to Qwen3.8-27B starts now!   submitted by   /u/Ok-Shower7286 [link]   [comments] 38 r/LocalLLaMA community 1d ago I asked DeepSeek-V4-Flash to work with Muse-Glimmer for Vision ability in PI agent and it produced this Same old prompt, just appended a TIP in the end: "Write a single HTML file with a full-page canvas and no libraries. Simulate a realistic side-view of a moving car as the main subject. Keep the car visible in the foreground while the background landscape scrolls continuously to… 33 Smol AI News news-outlet 1d ago not much happened today **Google** rapidly released **Gemini 3.7 Flash** just three weeks after 3.6 Flash, targeting coding, web development, knowledge work, and agentic workflows with a 50% introductory price cut and improved benchmark scores like **DeepSWE 65.3%** and **Code Arena Elo 1588**. The… 17 Hacker News — AI on Front Page community 1d ago Codex in ChatGPT desktop app for Linux is now in preview Article URL: https://community.openai.com/t/codex-in-chatgpt-desktop-app-for-linux-is-now-in-preview/1390027 Comments URL: https://news.ycombinator.com/item?id=49281916 Points: 374 # Comments: 263 24 arXiv — Machine Learning research 1d ago FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting arXiv:2608.11623v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational… 23 arXiv — Machine Learning research 1d ago DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks arXiv:2608.11873v1 Announce Type: new Abstract: In this work, we extend the Dependent Click Model (DCM) Bandits to a multiplayer information-asymmetric setting, where multiple agents interact with a shared ranked list and may observe multiple clicks per session, introducing new… 30 arXiv — NLP / Computation & Language research 1d ago Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing arXiv:2608.08514v1 Announce Type: cross Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B… 35 arXiv — NLP / Computation & Language research 1d ago Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets arXiv:2608.11233v1 Announce Type: new Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent… 19 arXiv — NLP / Computation & Language research 1d ago Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release arXiv:2608.11822v1 Announce Type: new Abstract: A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested… 5 arXiv — NLP / Computation & Language research 1d ago Explainability in Practice: A Survey of Explainable NLP Across Various Domains arXiv:2502.00837v3 Announce Type: replace Abstract: Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o, Gemini, and BERT increasingly inform decisions. The… 26 Simon Willison community 1d ago alchemy-utils 0.1a1 Release: alchemy-utils 0.1a1 Performance boost for DuckDB exports and CSV imports, see here . 8 LangChain releases dev-tools 1d ago langchain-anthropic==1.5.6 Changes since langchain-anthropic==1.5.5 release(anthropic): 1.5.6 ( #39622 ) fix(anthropic): normalize tool_search_tool_result blocks ( #39621 ) fix(anthropic): correct model profile data for Fable 5, Sonnet 5, Opus 4.1 ( #39604 ) 13 Page 1 of 10 · 500 articles Older →