News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 4h ago modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection link to original article TLDR: security researcher eddie zhang used a modified+uncensored local qwen 3.8 27b to create an executable capable of dumping LSASS memory for credential harvesting while evading 2 modern EDR security products. this makes me reflect on how cloud… 15 r/LocalLLaMA community 5h ago GPT-3 is discontinued today It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our memories. Arguably what's more infuriating is that they suggest using GPT-5.6 Terra as a replacement. Keep in mind that… 20 r/LocalLLaMA community 6h ago macOS 27 ships a free local LLM on Apple Silicon Macs. I made it easy to use from Node and Python Apple Silicon Macs on macOS 26+ come with a small LLM built in. No download, no API key, and nothing leaves your Mac. Why I built it I was making a tool that writes API docs from code, and I didn't want users to install Ollama or paste an API key. Apple's model was already on… 29 arXiv — NLP / Computation & Language research 7h ago Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only… 28 arXiv — NLP / Computation & Language research 7h ago Epstein Files Engine: Agentic Search for Investigative Journalism arXiv:2609.30611v1 Announce Type: cross Abstract: On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times… 24 r/LocalLLaMA community 7h ago I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2… 36 r/LocalLLaMA community 11h ago Qwen company already rushed out a Jev competitor. No open weights yet. It's called decision-model-preview. There is only a docs page. No announcement or anything. I can't post a link because reddit's filters just deletes posts that contain a link to the cloud platform that hosts it. But you'll find the page if you Google the model name.  … 28 r/LocalLLaMA community 12h ago Qwen plays World of Warcraft Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control… 5 r/LocalLLaMA community 12h ago Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM We built an inference engine for MoE models that don't fit in VRAM + RAM. Most of the model stays on the SSD, and experts are read as tokens need them. This started as a proof of concept, and poc worked, we are getting 9-10 tok/s decode on Qwen3.8-Flash-Next NVFP4 (9.06 on the… 6 llama.cpp releases dev-tools 12h ago b11223 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) ( #28876 ) server : allow splitting RANK pooling for causal LLM rerankers Rerank models fall into two categories: bidirectional cross-encoders (BERT, etc.) that require all tokens in a… 26 r/LocalLLaMA community 14h ago Swift 1.5 Qwen3.8 27b (A must-have for low thinking!) Just made this post for those who missed it : https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b UkisAI released their updated Qwen 27B (tuned for token efficiency). I grabbed the IQ4_XS quant to test against Unsloth's Q4_K_S: Low-thinking: UkisAI consistently beat Unsloth in… 9 r/LocalLLaMA community 14h ago ... so, yeah. Finally got 3.8-Flash-Next running on my M4Pro 48GB Mac with https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Dense 3.8-27B is just faster... and maybe better due to quantization level... EDIT: Hold a second, Flash-Next is actually performing faster than 27B… 20 r/LocalLLaMA community 16h ago Don't trust frontier models when asking about budget hardware! Early this year when I was first looking at building up my inference capability you could get the 16GB Tesla P100s for between $60 and $80. Asked claude about it, told me absolutely not worth it. No tensor cores, bad int4/int8, no BF16, not worth it. Needs special power… 36 r/LocalLLaMA community 17h ago Updated from 3x3090(2x3090, 1x3090TI) to 2x5090 Upgraded from 3x RTX 3090s to 2x RTX 5090s on my homelab server and picked up a solid speed jump on top of it from a software update (speculative decoding + NVFP4). Setup: llama.cpp (build b11216), running Qwen3.8-27B (abliterated, Q8_0). Blue = old 3090 setup, green = new 5090… 37 r/LocalLLaMA community 18h ago is switching from llama cpp to vllm worth it I have hp z8 g4 with 512 ram and 1x3090 1x5060 16gb. has anyone made the transition from llama cpp to vllm recently? is it worth it? docker under windows or full linux install? I am mainly interested in the model support, it seems that many new local models are supported day 0… 37 r/LocalLLaMA community 18h ago Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy Meta came out with a banger paper https://arxiv.org/pdf/2606.00206 , but it did not look at various quantizations supported in llama.cpp. So I did a run on 50 random MATH-500 questions ( https://huggingface.co/datasets/HuggingFaceH4/MATH-500 ) and ran it on various quantizations… 10 r/LocalLLaMA community 18h ago How do you guys give your models web browsing capabilities? I am using the Deepseek Harness, which has webfetch plugins by default. While it can help browse the internet, I myself have to give it specific URLs to search. But, apparently you can give it the full web browser and search engine capabilities. Unfortunately, it apparently… 33 r/LocalLLaMA community 20h ago We released VeriLoop E2 (27B, Apache-2.0). The design question behind it: should an LLM be allowed to commit its own state? Disclosure: I’m one of the authors of VeriLoop E2, a 27B model post-trained from Qwen3.8-27B. The weights are Apache-2.0; the harness we evaluate it in is not open source (details at the bottom). This is a release post, but rather than a benchmark dump I want to talk about the… 28 r/LocalLLaMA community 21h ago Worth going from Qwen3.8 27B to flash next or maaybe deepseek v4 flash? I am running qwen3.8 27b on my dual rtx 3090 (fp8 quant, unquantized cache, 129k context) and I think it works decently well with hermes, opencode etc. But! I am tempted by the new models coming out such as qwen3.8 flash next, deepseek v4 flash, glm 5.3 flash. However, there is… 20 r/LocalLLaMA community 22h ago The Opus 5.5 posts about motion graphics are cool but Qwen 27B made this on a 4090 Saw the hundreds of tweets where people just keep asking Opus 5.5 for motion graphic videos. Decided to ask qwen to look at them and make its own. Quite amazing what local can achieve. **EDIT** It looks laggy because of reddits .gif limit btw the full high res version (with… 20 r/LocalLLaMA community 1d ago Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_nonogram_logic_puzzle/ . That thread shaped v1.2: Reasoning effort is explicit per run Every prompt and output is public. All current top ranking private and open weight… 7 r/LocalLLaMA community 1d ago Which of the 16gb VRAM qwen3.8 27b’s is the best? I’m having a hard time finding out which one gives you fastest speed, maximum context with best possible quality. I can run unsloth qwen3.8 27b iq4_xs with 65k q8 kv, context without MTP and vision offloaded to cpu. But also kinda slow for agentic work at like 30ish tok/s )I… 30 r/LocalLLaMA community 1d ago Another "Harness matters" post (codex cli > pi and opencode) I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my 2x3090 + Ram! But local LLM never did worked for me outside some demos like build me a "3D Mario Game, multistage" which I've… 27 r/LocalLLaMA community 1d ago Qwen, where's the small stuff? (1B/2B/4B) I know Qwen is a key player in the local LLM space and has consistently introduced truly impactful technologies—like n-gram in Qwen-Next and the recent Qwen 3.8 27B, which is an amazing local model. However, my question is: why are we seeing fewer small-scale models lately—such… 24 r/LocalLLaMA community 1d ago Splash fork optimised for M5 Max: ~1.5× faster (1.25× single request) I’ve spent the last couple of days with Opus 5.5 working on a fork of Inco’s excellent and already blazingly fast Splash engine to optimise it for M5 Max chips. Taking liberties and referring to it as Splish. Roughly the opposite direction to u/Erp4759 ’s great M1 port ( Splash… 36 r/LocalLLaMA community 1d ago When is the next generation of "B tier" models releasing? The only things released in the last couple weeks seem to be bigger models like Qwen, DeepSeek, GLM, etc. Where are the Laguna's, Nemotrons, Olmo's, LongCats, Minimax, etc releases?   submitted by   /u/jazir55 [link]   [comments] 35 r/MachineLearning community 1d ago Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P] Hey r/MachineLearning ! I’ve been working on a new optimizer called Tauon (turns out there is already "teon" but well if you have better idea, - i will gladly accept it! Anyway the core idea of optimizer is about polynomials and orthogonalization just like muon, the whole… 5 TechCrunch — AI news-outlet 1d ago Google tests buying from Walmart-owned Flipkart through Gemini and AI Mode in India The limited test covers select products and users, with a broader rollout planned for later in October. 24 Vercel — AI dev-tools 1d ago Ember-1 from Fireworks now available on AI Gateway Ember-1 from Fireworks is now available on AI Gateway . Ember-1 is a research preview reasoning model built on Kimi K3 for coding and agentic workflows. Fireworks reports approximately 40% fewer generated tokens than Kimi K3 at comparable quality across its evaluations. For… 13 r/LocalLLaMA community 1d ago Koboldcpp v1.122 released   submitted by   /u/Fcking_Chuck [link]   [comments] 18 r/LocalLLaMA community 1d ago Run Qwen3.8+Flash-Next and tiny models on Apple Silicon up to 3x faster Maybe you'll like it? I hope I get to use my self-promotion credit a tiny little bit here after being in the community so long haha. I was the top of MLX.fast for a while and remain the winner on chips below M5. If you have capacity to contribute further enhancements I'd love… 10 r/LocalLLaMA community 1d ago Improved and fixed template for GPT-OSS (again). Includes preserve_thinking and fix for Unsloth-induced bug I posted an updated GPT-OSS template a couple of months ago , which was based on Unsloth's version . It turns out that both Unsloth's version (and, thus, mine) contain a very serious bug that can degrade the model when chat history is replayed and contains previous reasoning… 14 r/LocalLLaMA community 1d ago I added Qwen-Image 2.1 + LoRA support to TensorSharp (GGUF, local inference) I maintain TensorSharp , an open-source inference engine. It can now run Qwen-Image 2.1 locally for text-to-image generation and image editing, with support for its LoRA adapters. I’ve added configs for regular style and editing LoRAs, plus accelerated adapters with their own… 14 r/LocalLLaMA community 1d ago Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? I wish Qwen also released dataset and method to fully train a model ourselves but it is what it is. However, I come here with my stupid question because someone can answer it better. And will the model still be an over thinker of faster inference will make up for that.  … 5 Hacker News — AI on Front Page community 1d ago DeepSeek Elastic Compute (DSec) Article URL: https://arxiv.org/abs/2609.22978 Comments URL: https://news.ycombinator.com/item?id=49859112 Points: 200 # Comments: 61 33 r/LocalLLaMA community 1d ago New to playing around with local ai. why are they free? I am new to playing around with local ai and have a ton to learn about it but I was curious, why are they (who is they?) releasing them for free, don't they want you to pay them to use them, why release free models?   submitted by   /u/poofph [link]   [comments] 20 r/LocalLLaMA community 1d ago 85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri I've been experimenting with ways to run MoE models that don't fit comfortably in RAM, and I ended up making Overspill, a disk tier for FreeToken . The basic idea came from looking at how Colibri handles experts across disk/RAM/VRAM so I took inspiration from the general… 11 r/LocalLLaMA community 1d ago Splash 1.1.0 released, GGUF quants support, MLX import and more On my M5 Pro 64GB I can comfortably work in an agentic setup with the Qwen3.8 27B model in good quality (Unsloth UD-Q4_K_XL) at a decent speed of 50 t/s. Splash combines optimized kernels, excellent speculative decoding, a well-implemented prefix cache, and mixed-weight support… 26 r/LocalLLaMA community 1d ago Just bought a second 3090 but now I don't see the benefits right now. Hi, My local AI server consists of 96GB Ddr5 and one rtx 3090. I got plenty of stuff running, like krea2, qwen image 2.1, minimax h3, ltx 2.5, qwen 3.6, qwen 3.8 q4...,got even qwen 3.8 Flash next running. But now I am thinking of what I can utilize the second card for. I am a… 37 r/LocalLLaMA community 1d ago Qwen3.8-27B IQ3_XXS vs Qwen3.6-35B-A3B Q4_K_M Which one is better for difficult tasks like web scrapping, coding, using tools? Looking for any benchmarks because i couldn't actually find one after quite some digging   submitted by   /u/Loose_Doubt367 [link]   [comments] 30 Don't Worry About the Vase community 1d ago Claude Opus 5.5 Should Raise Your Ambitions When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. 31 r/LocalLLaMA community 2d ago Qwen3.8 flash next + exllamav3 + hermes is amazing I know there is nothing new with what I am saying but I recently started with hermes agent (it’s been a while I wanted to but did not have the time). Qwen3.8fn 6bpw exl3 (from turboderp) on a 6x3090 (I assume lower quants on lower number of gpus work same) gives me around… 28 r/LocalLLaMA community 2d ago I built a tiny (332MB) CPU-friendly model for document sorting that actually knows when to say "none fits" (BeeNara) Hey r/LocalLLaMA ! I wanted to share a small project I’ve been working on called BeeNara Why I built this: I was looking for a way to automatically sort my local documents (invoices, letters, contracts) into my personal folders. While local LLMs are amazing, I noticed that… 35 r/LocalLLaMA community 2d ago Introducing KoboldCpp Agent (and a plea for help) Hello r/localllama once again, it's me your kobold concedo Been a few months since I last posted here, and today I have something new I'd like to share. Specifically, KoboldCpp now ships with a built-in integrated KoboldCpp Agent Harness ! I know it's a little late to the game,… 29 r/LocalLLaMA community 2d ago internlm/Intern-Decision 4B and 0.8B https://huggingface.co/internlm/Intern-Decision-0.8B Intern-Decision-4B Demo | Model Weights | GitHub Intern-Decision-4B is a multimodal structured decision model fine-tuned from Qwen3.5-4B . It accepts a shared state, a schema of named questions, and optional images, and… 6 r/LocalLLaMA community 2d ago Best current Qwen Flash Next Q4-ish? + worth using? Im running a 5090 and 64gb of ram, so im limited on what I can run. I have currently been able to fit the following - Atomic Q4_k_m 4.27bpw @ 31 layers offload Swift IQ4_xs @ 32 layers offload. Im about to try the Unsloth IQ4_xs as well. I could get a "bigger" (non IQ) quant for… 19 r/LocalLLaMA community 2d ago Swift 1.5 27b: Swift Qwen just got faster Enjoy! Fucking loving it.   submitted by   /u/sleight42 [link]   [comments] 11 r/LocalLLaMA community 2d ago Qwen3.8-27B Q4_K_M on 2x3060 For the last few days, I've been using two computers to run multiple agents. My 4x3090 machine is running Qwen 3.8 27B with parallel=2, so I can run two agents at the same time. My pi instances are running on a machine with 2x3060, running/managing smaller models such as Gemma… 31 Ars Technica — AI news-outlet 2d ago Court rules Trump can blacklist Anthropic for refusing to enable Claude features "Overly constrained AI models" could cause military operations to fail, judges say. 6 r/LocalLLaMA community 2d ago How long can I expect to wait until the local ~30B A3B frontier catches up to GLM 5.3 Flash quality? The jump from Qwen3 Coder 30B A3B to current-day Qwen 3.6 35B A3B is crazy, especially with all the fine-tunes, and that was around 6 months (I didn't care for local AI back then, or AI at all, apart from as a toy so I don't know). Is around a year until I will never need cloud… 16 Page 1 of 10 · 500 articles Older →