News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow Latent.Space news-outlet 5d ago [AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50% overshadowing more efficient GPT6 models from OpenAI 32 r/LocalLLaMA community 5d ago GGUFs in transformers natively! Hey there folks! Aritra here from Hugging Face. I wanted to update you all about the latest changes in `transformers`. We now natively support GGUFs (llama cpp quants). You can use it like so: from transformers import AutoModelForCausalLM, AutoTokenizer model_id =… 12 arXiv — Machine Learning research 5d ago Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow arXiv:2609.25131v1 Announce Type: new Abstract: Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also exposes correct intermediate predictions to later errors. An experiment on Sudoku… 20 arXiv — Machine Learning research 5d ago Lightweight Ranking Heads: Accelerating Multi-Task Experimentation in Production Recommender Systems arXiv:2609.25433v1 Announce Type: new Abstract: Modern production-scale recommender systems rely on complex, multi-task ranking models. Introducing new prediction tasks into these massive systems often causes bottlenecks - it risks negative task conflicts with existing tasks,… 13 arXiv — Machine Learning research 5d ago An Exploratory Replica-Overlap Probe of the Grokking Transition arXiv:2609.25634v1 Announce Type: new Abstract: We trained 64 independently seeded networks in four configurations, continuing each to sustained convergence or a 40,000-epoch ceiling. We then asked whether an RSB-inspired distribution of pairwise weight overlaps changes across… 5 arXiv — NLP / Computation & Language research 5d ago Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum arXiv:2609.25028v1 Announce Type: new Abstract: QMSum provides no scorer, making query-focused meeting summarization results difficult to compare. We rescore or generate 15 systems under one implementation. Through a common inference port, a released 406M Fusion-in-Decoder… 38 arXiv — NLP / Computation & Language research 5d ago From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication arXiv:2609.25034v1 Announce Type: new Abstract: Central bank press conferences are not merely information releases --- they are structured narratives. We study whether the shape of sentiment within a statement, not just its average tone, carries policy-relevant signals.… 13 arXiv — NLP / Computation & Language research 5d ago LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay arXiv:2609.25053v1 Announce Type: new Abstract: Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our… 23 arXiv — NLP / Computation & Language research 5d ago TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks arXiv:2609.25356v1 Announce Type: new Abstract: Large language models (LLMs) offer great potential to automate a broad range of telecom engineering tasks by reasoning over standards, network configurations, mathematical models, source code, and operational logs. However,… 8 arXiv — NLP / Computation & Language research 5d ago Qwen3.8-Omni: Towards Native Omni-Modal Agents arXiv:2609.25611v1 Announce Type: new Abstract: We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash… 30 arXiv — NLP / Computation & Language research 5d ago Detecting GPT-Assisted Writing Using Interpretable Stylometric Features arXiv:2609.26687v1 Announce Type: new Abstract: Distinguishing GPT-assisted from independently authored student writing has become a critical challenge in academia. This paper evaluates the discriminative capability of interpretable stylometric features extracted solely from… 8 arXiv — NLP / Computation & Language research 5d ago Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction arXiv:2609.25176v1 Announce Type: cross Abstract: Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think… 9 r/LocalLLaMA community 5d ago DeepSeek and Moonshot AI face Beijing's probe over potential data leaks to Anthropic The rumor about Kimi execs getting arrested finally has some legs. I believe the reality is more like under investigation for potential arrests or fine.   submitted by   /u/Ok_Warning2146 [link]   [comments] 20 The Information — AI news-outlet 5d ago Anthropic in Talks to Cement Control Over More Data Centers Anthropic is in early talks to lease up to 1 gigawatt in compute capacity from a data center developer majority-owned by Apollo Global Management, part of a major effort by the AI company to reduce its reliance on cloud providers. The maker of Claude has discussed signing on as… 16 OpenAI official-blog 5d ago Airbnb widens access to GPT-6 Astra and OpenAI frontier models Learn how Airbnb is expanding access to GPT-6 Astra and OpenAI frontier models to help engineering teams solve bugs, design systems, and ship faster. 6 r/LocalLLaMA community 5d ago Unsloth Studio VS LM Studio... Which one do you prefer? So I've been experimenting with various platforms and even on day 1 Unsloth Studio was released, I knew that LM Studio had it's days numbered. LM studio will always be the OG but I wonder how much longer they have, especially with all these new platforms arising. It feels like… 28 Ollama releases dev-tools 5d ago v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550) mlx: speed up Qwen 3.8 prompt processing Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU. address comments 31 OpenAI Python SDK releases dev-tools 5d ago v3.19.0 3.19.0 (2026-09-22) Features api: add GCP external storage support ( #3943 ) ( d12d60f ) api: add GPT-Rosalind research model ( #3940 ) ( d041d73 ) Bug Fixes _utils/_transform: propagate api_exclude in _async_transform_recursive ( #3324 ) ( 161ae65 ) api: handle omission markers… 7 OpenAI official-blog 5d ago Grab and OpenAI bring practical AI skills to Southeast Asia OpenAI and Grab launch GO Forward with AI, a regional programme helping 30,000 partners build practical AI skills across Southeast Asia. 25 Vercel — AI dev-tools 5d ago Gemini 3.8 text-to-speech models now available on AI Gateway Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS from Google are now available on AI Gateway . Both models take text and generate speech in more than 100 languages. They support long-form narration, control over delivery, and two-speaker dialogue.… 20 Simon Willison community 5d ago Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war Yesterday was Grok 4.7 ( pelicans ) and MiMo v2.6 Flash/Pro ( more pelicans ). Today Anthropic released Claude Opus 5.5 , and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna . It's going to take a while to get a good read on all of these new models, but here are my… 10 The Information — AI news-outlet 5d ago Exclusive: Microsoft to Boost Copilot Discounts as It Launches AI ‘Super App’ Microsoft leaders told sales staff this week that it is authorizing steeper discounts of 30% to 50% on subscriptions of its Copilot AI software for corporate customers who commit to purchase a large number of seats and commit to additional payments based on their usage of… 30 The Information — AI news-outlet 5d ago Anthropic releases cheaper model in first launch since slowdown calls Anthropic announced the release of its newest model, Claude Opus 5.5 , saying it performs as well or better than Anthropic’s previous most powerful models, Fable 5.1 and Mythos 5.1, in domains including coding, reasoning, business workflows, and safety. At the same time, it will… 18 LangChain releases dev-tools 5d ago langchain-openai==1.6.4 Changes since langchain-openai==1.6.3 release(openai): 1.6.4 ( #40775 ) chore(model-profiles): refresh openai model profile data ( #40774 ) 27 LangChain releases dev-tools 5d ago langchain-anthropic==1.7.3 Changes since langchain-anthropic==1.7.2 release(anthropic): 1.7.3 ( #40773 ) chore(model-profiles): refresh anthropic model profile data ( #40772 ) fix(anthropic): auto-route with_structured_output to method="json_schema" for fable and opus 5.5 ( #40766 ) chore(anthropic):… 38 r/LocalLLaMA community 5d ago My local llm when I tell it to do any changes to my vLLM service better make no mistakes I've been running Qwen 3.8 Flash Next and it's a great driver for Hermes and Pi. I told it to add CUDA_DISABLE_PERF_BOOST=1 to reduce my server's idle power draw   submitted by   /u/ZaltyDog [link]   [comments] 38 OpenAI official-blog 5d ago Better prompt caching for GPT-6 Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs. 16 The Information — AI news-outlet 5d ago Wall Street’s GPU Futures Push Stalls at the CFTC The launch of futures for rental prices on Nvidia GPUs, which are seen as important to make AI compute an investable asset class, is hitting a snag. CME Group had hoped to launch contracts as soon as early October, but it won’t have regulatory approval by then. The Commodity… 15 r/LocalLLaMA community 5d ago quants for K2-Horizon are now available You can finally downloads quants from: https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B-GGUF https://preview.redd.it/cfhl43pps4rh1.png?width=2800&format=png&auto=webp&s=313eab309fb407a0a0e5da61e2a120c677d3d3cd https://huggingface.co/IFM/K2-Horizon-32B-GGUF… 12 r/LocalLLaMA community 5d ago Cloud-AI Cold-Turkey: Real Dev Work with Local AI (Ornith 1.5 35b-a3b and Qwen 3.8-27b; 8GB VRAM vs 32GB VRAM) Spent several weeks on an 'as-much-Local-AI-as-possible' regime and have been - mostly - impressed. Yes, Qwen 3.8 27b is the (rightful) star of the show (obligatory one-shot-mario-build-reddit-comment here! ). But don't underestimate what models with more modest hardware… 26 TechCrunch — AI news-outlet 5d ago Qualcomm launches two new smartphone chips with emphasis on AI Qualcomm said that its new top chip can run 30B mixture-of-expert model locally. 9 Simon Willison community 5d ago llm 0.36 Release: llm 0.36 New OpenAI models: gpt-6-sol for GPT-6 Sol and gpt-6-luna for GPT-6 Luna . #1702 Model plugins can now declare supports_conversation = False for models that only accept single-turn prompts. LLM raises llm.ConversationNotSupported when these models receive… 13 r/LocalLLaMA community 5d ago Qwen image 2.1 (Fast FP8) generates premium quality images Don't know how they did it, but for under 10GB model, the results are astonishing. I am running it on Unsloth Studio. They just released the update, so if you are not seeing the option, I recommend updating your Unsloth Studio. Cheers!   submitted by   /u/108er [link]… 18 r/LocalLLaMA community 5d ago Now Opus 5.5 is 58 on Artificial Analysis , how long do you have wait until an open model hits 58? It is 12 points higher than the best open model mimo 2.6 pro and a big jump from fable 5.1. Crazy, glm 5.5 and qwen 4 will be on par with gpt 6 sol or better since it has a score of 48 If it took 2 months for the best open model to go from 44 to 46, then at this rate, in 6… 28 OpenAI Python SDK releases dev-tools 5d ago v3.18.0 3.18.0 (2026-09-22) Features api: add GPT-6 Sol and Luna model identifiers ( #3935 ) ( 455ce1b ) 34 r/LocalLLaMA community 5d ago To my surprise I found gemma4 much better at tool-calling than Qwen I've had fairly good luck getting off the ground coding at home with both qwen3.6 35B a3b, and also qwen3.8 27B. However once I switched from a chat window (where the robot wrote code blocks that I could copy/paste into a text editor) to a simple agentic loop, things got funny.… 4 The Information — AI news-outlet 5d ago OpenAI Announces Cheaper GPT-6 Family Models OpenAI on Tuesday announced GPT-6 Sol and GPT-6 Luna, which it describes as cheaper and faster models in the GPT-6 Astra family. The company said that the new models show significant improvements compared to their predecessors, GPT-5.6 Sol and GPT-5.6 Luna, in areas like… 25 Hacker News — AI on Front Page community 5d ago GPT-6 Sol and Luna Article URL: https://openai.com/index/introducing-gpt-6-sol-and-luna/ Comments URL: https://news.ycombinator.com/item?id=49805509 Points: 408 # Comments: 207 33 TechCrunch — AI news-outlet 5d ago OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes OpenAI is launching two new models, which it says are cut from the same cloth as Astra. 32 OpenAI official-blog 5d ago Introducing GPT-6 Sol and Luna Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost. 13 The Information — AI news-outlet 5d ago Exclusive: Meta’s Muse Surpassed 500,000 Users After First Week Meta’s new personal AI agent, Muse, had more than 500,000 people try it out roughly a week after launch, including more than 250,000 daily active users, according to internal data viewed by The Information. Users have also submitted more than 2 million prompts. The data indicate… 34 llama.cpp releases dev-tools 5d ago b11109 metal : gate mul_mm_id src1 rescale behind ggml_prec ( #29029 ) metal : gate mul_mm_id src1 rescale behind ggml_prec Assisted-by: Claude Fable 5.1 ggml-webgpu: reject MUL_MAT_ID when src1 precision is F32 cuda/vulkan: reject MUL_MAT_ID in supports_op when src1 prec is F32 fix… 10 The Information — AI news-outlet 5d ago Why America’s AI Dream Is Failing to Launch In 2022, American dominance in AI looked almost unassailable. The U.S. led in advanced models, semiconductors, capital and talent. Its closest competitor, China, appeared to be reeling from export controls that cut it off from the world’s most advanced process nodes, chips and… 37 Simon Willison community 5d ago llm-anthropic 0.29 Release: llm-anthropic 0.29 Adds support for Claude Opus 5.5 : llm -m claude-opus-5.5 "prompt goes here" Tags: llm , anthropic 11 Hacker News — AI on Front Page community 5d ago Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max) Article URL: https://artificialanalysis.ai/models/claude-opus-5-5 Comments URL: https://news.ycombinator.com/item?id=49804316 Points: 246 # Comments: 77 5 r/LocalLLaMA community 5d ago Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW Someone recommended that I try the ByteShape Qwen 3.8 27B IQ3-XXS GGUF after seeing my previous testing of the GSQ quant. So I did. And the result was… surprisingly bad. For context, I'm running: RTX 3060 12GB 16GB DDR4 RAM, single channel CachyOS / Arch Linux llama.cpp Qwen 3.8… 6 TechCrunch — AI news-outlet 5d ago Anthropic releases Opus 5.5 with lower prices and Fable-level performance Anthropic called it "the strongest-performing model we've tested to date." 35 Hacker News — AI on Front Page community 5d ago Claude Opus 5.5 Article URL: https://www.anthropic.com/claude-opus-5-5 Comments URL: https://news.ycombinator.com/item?id=49803892 Points: 363 # Comments: 439 8 Anthropic SDK (Python) releases dev-tools 5d ago v1.8.0 1.8.0 (2026-09-22) Full Changelog: v1.7.0...v1.8.0 Features api: add support for claude-opus-5-5, inline tool definitions and MCP tool-list pinning (beta) ( b5cc700 ) Bug Fixes api: share one evaluated_permission enum across Managed Agents events ( f4f51c8 ) streaming: avoid a… 28 Simon Willison community 5d ago llm-typesafe 0.1a0 Release: llm-typesafe 0.1a0 I built this new plugin for LLM to add support for TypeSafe AI's new Jev model . Install it like this: llm install llm-typesafe Then set an API key ( get one here , the waitlist seems to move pretty fast): llm keys set typesafe # Paste key And now you… 34 Page 4 of 10 · 500 articles ← Newer Older →