News / #edge Tag Edge 500 articles archived under #edge · RSS Sign in to follow arXiv — NLP / Computation & Language research 18d ago X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding arXiv:2609.09166v1 Announce Type: new Abstract: This paper investigates collaborative speculative decoding (CoSD), a distributed large language model (LLM) inference framework in which an on-device small language model (SLM) drafts candidate tokens and a server LLM verifies… 30 arXiv — NLP / Computation & Language research 18d ago From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls arXiv:2609.09476v1 Announce Type: cross Abstract: In-vehicle assistants must translate natural-language requests into accurate vehicle function calls under strict memory and latency constraints, making small language models (SLMs) attractive for on-device deployment. For such… 27 arXiv — NLP / Computation & Language research 18d ago PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling arXiv:2609.09662v1 Announce Type: cross Abstract: Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. However, LLMs have heavy… 27 r/LocalLLaMA community 18d ago Local LLM / Qwen 3.8 win I’ve been working on perfecting my setup ever since 3.8 hit, but no longer pay for subscriptions. I had a coding interview today and setup openrouter ahead of time as another option. I figured I could use the latest GLM flash for cheap and it would be fast. Nope. Over thought… 9 r/LocalLLaMA community 18d ago Would you consider 5t/s usable for a local model? I'm able to run qwen3.8 27b in two ways on my system: split between my 3060 12gb and 9070xt running at 20t/s or running off the 780m iGPU and 5400mhz DDR5 at 5t/s. Personally I feel like the 5t/s is still more usable because I have enough RAM to still use my system mostly… 4 r/LocalLLaMA community 18d ago Don't let FOMO win if you're interested in local llm from a hobby/learning aspect Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over. No matter the hobby, it's so easy to get stuck in a trap where we buy more trying to do more only to realize we've lost the fun in it all or even the… 21 TechCrunch — AI news-outlet 18d ago Apple CEO John Ternus says the best AI device is still the iPhone The company also argued that its on-device models offer consumers more privacy. 8 r/LocalLLaMA community 18d ago Desert Ant Labs: On-device intelligence for every product   submitted by   /u/Arcuru [link]   [comments] 38 Hacker News — AI on Front Page community 19d ago Desert Ant Labs: local, fast models that run on device Article URL: https://desertant.com/blog/introducing-desert-ant-labs/ Comments URL: https://news.ycombinator.com/item?id=49624823 Points: 205 # Comments: 61 31 r/LocalLLaMA community 19d ago Qwen3-0.6B (400 MB) on a Samsung Note 8 (2017) phone drives a real desktop Chrome Up front: I'm one of the people building the page-perception layer used here. We started by testing small local models. The result turned out to be more interesting than the original test. 12 small models, 3 verifiable tasks, logs, and offline replay. Setup: Galaxy Note 8 (2017,… 17 r/LocalLLaMA community 20d ago Which local model is actually good at knowing when to stop and ask you a question? I’ve been thinking about this after using more agentic/local coding models. A lot of the newer models are surprisingly good at continuing on their own. But sometimes that seems like the problem. If a requirement is ambiguous, I’d rather the model stop and ask: “Do you mean A or… 18 r/LocalLLaMA community 20d ago I made Warrior Quest, a local LLM-powered dark-fantasy RPG where the model only plays NPCs and the actual game state stays deterministic The LLM is limited to NPC emulation. Game state, world logic, quests, and the authored story are handled by deterministic game systems rather than the LLM. I built it this way because I wanted the freedom of talking to NPCs like you would at a tabletop game, without handing the… 7 r/LocalLLaMA community 20d ago After over a year of my nights and weekends, the Jenny app is done! Hi all! I just wanna say that I am tired lol. Yes, it's another harness, but I spent a lot of time and effort and have forsaken my hobbies to build the Jenny (like XJ-9) app. Jenny is a free, MIT licensed electron desktop app for running local LLMs with tool calling, rollback,… 14 r/LocalLLaMA community 20d ago 9 easy steps for llama.cpp, a local model, Freecad (and pi coding agent) to generate solid objects that sound mechanically good and can be also be 3D printed/milled Quick setup on linux: STEP 0: install/download llama.cpp, Freecad, your favourite gguf model - possibly with multimedia image reading capabilities (I've used Qwen3.8-27B-UD-Q4_K_M and relative mmproj-F16 quantized by Unsloth), uv (, pi.dev) STEP 1: $ cd /your/path/to/ (i.e.… 6 r/LocalLLaMA community 21d ago Easy local Copilot with VS Code and Lemonade Not so long ago I wrote a guide on how to get GitHub Copilot running with a local model in Visual Studio Code. Since then Copilot subscriptions have got much more expensive, local models have got much more powerful and getting local Copilot up and running has got much easier. So… 6 r/LocalLLaMA community 21d ago Best local models for hardware programming? Guys can you tell me which local LLMs are best for hardware programming? like Verilog RTL, UVM, System Verilog?   submitted by   /u/Curious_Cantaloupe65 [link]   [comments] 11 r/LocalLLaMA community 21d ago What is the obstacle in front of Local Frontiers? We've reached a point with local LLMs where models are now very close to (and even reach) the level of models like the Opus, with some minor modifications. While some K3 and GLM 5.3 models are incredible, they are barely as powerful as the Opus or on par with the Fable or Astra.… 7 arXiv — Machine Learning research 21d ago PACE: Propagation-Aware Collaborative Correction for One-Shot Personalized Federated Graph Learning arXiv:2609.04832v1 Announce Type: new Abstract: Client heterogeneity creates both an opportunity and a risk in personalized federated graph learning. Knowledge held by other subgraphs may complement a receiver's Local model, but an incompatible transfer can override reliable… 18 r/LocalLLaMA community 21d ago Trying to create my own server and consuming it for code with my phone remotely (Mac OS) Hi there! I need some help with this. I have a 32gb Macbook Pro with the latest available update of Tahoe. I'm using LMStudio with MLX to serve a local model and I want to expose it so that I can consume it with my phone to code and review stuff when I'm commuting to places.… 28 r/LocalLLaMA community 21d ago 2x R9700, 64 GB DDR5 is an absolute beast machine with vLLM Radiance / R9V and Qwen 3.8 27b and Flash next I've been tinkering with local LLMs since the beginning of the year when I had an Intel Arc B580 and 32 GB of DDR5. Curiosity got the best of me and I bought the first R9700 about half a year ago, also because I wanted to upgrade my gaming graphics for 4k. As the 5090 was about… 14 r/LocalLLaMA community 21d ago Planning to get a cheap-ish GPU. Would appreciate some advice. Hi, I've been wanting to run my own local LLM for some time and I finally saved enough to get a budget GPU. I can spend 700$ at most, and I'm looking for a GPU that can run quantized 30B~ parameter models with decent speed. Being able to run Qwen 3.8 27 B Q4_K_M and similar… 14 r/LocalLLaMA community 22d ago Validate your local LLM advertised KV cache against real pressure; see exactly how old contexts get evicted from cache Hello, I'm a bit obsessed with cache management on local LLMs. For the last few days I've been working on cache management on vLLM with my 2x DGX Spark and DeepSeek v4 Flash 0731. I felt something was off so I investigated, found issues, fixed them, but needed a way to validate… 22 r/LocalLLaMA community 22d ago How does your favourite local model do with the slinky test? Prompt: Make me a single HTML file of a rainbow slinky going down an up-escalator forever. No libraries, just canvas and code. The slinky should be a chain of springs, each coil a different color of the rainbow. It starts folded in half like a horseshoe draped over a step. When… 27 r/LocalLLaMA community 22d ago My local LLM demoscene generator can now watch its own output and rewrite it! I've updated my auto_demo_scener project with Ninfer support and a “rewrite based on video” feature that I thought you might find interesting. The project is basically an endless demoscene machine. A local LLM writes Three.js effects (from a library of editable prompts), you… 23 r/LocalLLaMA community 22d ago Any resource on using Blender with local models, and which models work best? Hey all, I've seen some really fun looking things with people having their local models drive Blender to create pretty cool looking world scenes. Is there a good tutorial on setting up Blender yo be driven by your model? For example, what programming harness, do you use a MCP… 22 r/LocalLLaMA community 22d ago Is there a local LLM or toolchain to edit 3d models? I got a lot of ads for meshy recently and went to try it with hilariously bad results. It apparently can't do anything but decorative figurines. I wanted a body shell for an rc car and it just couldn't generate a car without wheels or bottom chassis. It also looks like it can't… 22 r/LocalLLaMA community 23d ago I've found myself using Local LLM's like 3D printers. Anyone who has a 3D printer and get use of it finds it incredibly useful for those odd jobs around the house, a missing bracket, a cable router, steam deck holder and so on. In the past if I was missing an app or useful software, a game I'd do the lazy thing, even though I can… 22 r/LocalLLaMA community 23d ago Qwen3.8-27b is the first Local model im able to blindly trust You know that thing where you just throw a task at a frontier model and not have to supervise it worrying of it going off course? Qwen3.8-27b has officially gotten me to that point for local work. He has been doing non-stop continuous agentic work for 8+ hours and hasnt screwed… 14 arXiv — Machine Learning research 24d ago LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference arXiv:2609.03079v1 Announce Type: new Abstract: On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload… 36 r/LocalLLaMA community 24d ago Bernie Sanders proposes to ban AI Defined as AI exceeding human cognitive abilities. 20 years in prison. Plenty of local models already fall under that big of an umbrella in some capacities. This is why it's not enough to say that you could torrent open models so who cares what the politicians do. They want you… 4 r/LocalLLaMA community 25d ago I built a local web UI to finetune models on my own text and actually watch the training (works on AMD ROCm) I wanted to do continued pretraining/finetuning of a local model on my own notes and see what's happening while it trains and do it on my AMD card, since most tools assume CUDA. llm-training-panel is a local web UI that: - loads a model from a local dir or a HF id - shows a live… 34 r/LocalLLaMA community 25d ago Can a 4B local model actually feel like an AI assistant? I've been building Arcon around Qwen3-4B + LoRA. Instead of just making it a chatbot, I'm experimenting with persistent memory, personality/mood, internal state, tools, and eventually having it process things before replying. I'm curious what people who've built local agents… 6 r/LocalLLaMA community 25d ago KV cache might be a bigger problem for local models than parameter count Everyone keeps talking about fitting larger models into local hardware, but parameter count isn’t the whole story. For long context inference, KV cache can become the real memory bottleneck. Every new token adds key and value states that need to stay available, so a model that… 8 arXiv — NLP / Computation & Language research 25d ago How Do Prompt Variations Affect Energy Consumption in On-Device LLMs? arXiv:2609.01798v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored. This paper aims to understand how two… 4 r/LocalLLaMA community 25d ago Chrome browser add-on that uses local LLMs to move thousands of unsorted bookmarks into a smart list of automatically calculated categories Is anyone interested in my Chrome browser add-on that uses local LLMs to move thousands of unsorted bookmarks into a smart list of automatically calculated categories?   submitted by   /u/paranoidray [link]   [comments] 7 r/LocalLLaMA community 25d ago What do you do in the meantime when your favourite local model is thinking and working hard with your harness? Real question, whether your are vibe coder or expert or whatever, what do you do? Read carefully each single word? Plan the next steps? Do the gym?   submitted by   /u/takoulseum [link]   [comments] 14 r/LocalLLaMA community 25d ago GLM 5.3 Flash makes a black hole Minecraft mod running locally on 4x RTX PRO 6000 WS saw the post the other day where people said Minecraft clones aren't impressive anymore, because at this point the whole thing might as well be in the training data. so i tried something slightly different, which is asking a local model to write a mod for the real game, using… 5 r/LocalLLaMA community 26d ago Qwen 3.8 Flash Next for Creative Writing? As we all know on of the best local models for creative writing is gemma 4 31b and Muse Glimmer 30b. However, ive been a happy user of Qwen3.8 Flash Next and I wanted to know how well Qwen 3.8 Flash next is doing in terms of creative writing (preferably German).   submitted… 19 Hacker News — AI on Front Page community 26d ago My local model setup on an M4 Pro Mac Mini Article URL: https://lws.io/blog/my-local-model-setup/ Comments URL: https://news.ycombinator.com/item?id=49529132 Points: 244 # Comments: 151 36 r/LocalLLaMA community 26d ago All currently popular local models in one table + Opus 4.8 results If you are thinking what model will fit best your HW specs and tasks you are doing here is one table with all currently popular models that still can be considered as local. LLM Test Scores Feature DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next… 24 r/LocalLLaMA community 27d ago Which current local models that can run within 128GB generate the best SVG pelicans? I used a famous Simon Willison's pelican riding a bicycle prompt on the biggest local LLMs that can run on 128GB Apple Silicon. U used quantizations by Unsloth. Qwen3.8 Flash-Next gives a lot of details. DeepSeek V4 Flash is strangely underwhelming. Qwen3.8 27B still rocks, and… 6 r/LocalLLaMA community 27d ago First time running local models Sad that I only have 12gb of vram but this ik_llama is so fast   submitted by   /u/Needausernameplzz [link]   [comments] 36 r/LocalLLaMA community 27d ago Sliding-window beats linear attention Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic attention with sliding window attention + attention sinks and no post training. This could be big for memory constrained local LLM… 25 r/LocalLLaMA community 27d ago The Chrono Trigger plot challenge - Crono awakens in his modest bedroom of 2095... A hallucinated event, Crono awakens in his modest bedroom of 2095... I am using local models since 2025 January. My daily driver is qwen 3.6 35B A3B, which works pretty well for coding tasks, but I always benchmark the models for lexical knowledge as well, where the models… 38 r/LocalLLaMA community 29d ago I fine-tuned a 0.8B local model for dictation cleanup. It matched a hosted frontier model on this narrow task I built SpeakoFlow Mini, an Apache-2.0 fine-tune of Qwen3.5-0.8B for dictation cleanup. It is not a chat model or a general rewriter. It takes speech-to-text output, applies corrections the speaker actually made, and leaves everything else alone. That last part is harder than it… 21 arXiv — Machine Learning research 1mo ago A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models arXiv:2608.26926v1 Announce Type: new Abstract: Small language models (sLLMs) are nowadays hosted on devices with limited memory and computational budget. In an autoregressive setup, inference is memory-bandwidth bound: uniform quantization is often detrimental to such models,… 11 r/LocalLLaMA community 1mo ago Ornith-1.5-35B-A3B on 8 GB VRAM: I think I've found my sweet spot A few days ago I posted asking what people considered the best local model for an 8 GB VRAM GPU . At the time, my personal sweet spot was Qwen3.6-35B-A3B , for agentic coding with Pi.dev. Well… Thanks to the suggestions in that thread, I think I've found something even better.… 25 r/LocalLLaMA community 1mo ago No, Engrams won't let you run 1T models locally. It does something even better. Ever since Qwen 3.8 Flash Next dropped, there's a misconception going around that N-gram tables will let people run 1T+ parameter models on a single server with 980B parameters offloaded to SSD. I'm here to disappoint you: it won't. But what it will actually do for local models… 31 r/LocalLLaMA community 1mo ago I used local Qwen 27b to build a harness and replace OpenCode Sharing my harness for running local LLMs that I built using Qwen 3.x 27B (> 90% locally built) under my supervision - not vibe-coded. Its free, no telemetry, and open-source (AGPL). Works on Windows, Linux (sorry, no Mac yet). I use it for my own coding + mixed workflows. How… 13 r/LocalLLaMA community 1mo ago Local LLM harness for reverse engineering software? Does anyone know of a decent reverse engineering harness/workspace setup for local models that I can just point the model at and have it go to work until it's reversed most if not all of the functions in a binary, even if it takes days? Of course, I'm willing to put effort into… 13 Page 2 of 10 · 500 articles ← Newer Older →