News / #edge Tag Edge 433 articles archived under #edge · RSS Sign in to follow r/LocalLLaMA community 3h ago 2x R9700, 64 GB DDR5 is an absolute beast machine with vLLM Radiance / R9V and Qwen 3.8 27b and Flash next I've been tinkering with local LLMs since the beginning of the year when I had an Intel Arc B580 and 32 GB of DDR5. Curiosity got the best of me and I bought the first R9700 about half a year ago, also because I wanted to upgrade my gaming graphics for 4k. As the 5090 was about… 14 r/LocalLLaMA community 4h ago Planning to get a cheap-ish GPU. Would appreciate some advice. Hi, I've been wanting to run my own local LLM for some time and I finally saved enough to get a budget GPU. I can spend 700$ at most, and I'm looking for a GPU that can run quantized 30B~ parameter models with decent speed. Being able to run Qwen 3.8 27 B Q4_K_M and similar… 14 r/LocalLLaMA community 11h ago Validate your local LLM advertised KV cache against real pressure; see exactly how old contexts get evicted from cache Hello, I'm a bit obsessed with cache management on local LLMs. For the last few days I've been working on cache management on vLLM with my 2x DGX Spark and DeepSeek v4 Flash 0731. I felt something was off so I investigated, found issues, fixed them, but needed a way to validate… 22 r/LocalLLaMA community 1d ago How does your favourite local model do with the slinky test? Prompt: Make me a single HTML file of a rainbow slinky going down an up-escalator forever. No libraries, just canvas and code. The slinky should be a chain of springs, each coil a different color of the rainbow. It starts folded in half like a horseshoe draped over a step. When… 27 r/LocalLLaMA community 1d ago My local LLM demoscene generator can now watch its own output and rewrite it! I've updated my auto_demo_scener project with Ninfer support and a “rewrite based on video” feature that I thought you might find interesting. The project is basically an endless demoscene machine. A local LLM writes Three.js effects (from a library of editable prompts), you… 23 r/LocalLLaMA community 1d ago Any resource on using Blender with local models, and which models work best? Hey all, I've seen some really fun looking things with people having their local models drive Blender to create pretty cool looking world scenes. Is there a good tutorial on setting up Blender yo be driven by your model? For example, what programming harness, do you use a MCP… 22 r/LocalLLaMA community 1d ago Is there a local LLM or toolchain to edit 3d models? I got a lot of ads for meshy recently and went to try it with hilariously bad results. It apparently can't do anything but decorative figurines. I wanted a body shell for an rc car and it just couldn't generate a car without wheels or bottom chassis. It also looks like it can't… 22 r/LocalLLaMA community 1d ago I've found myself using Local LLM's like 3D printers. Anyone who has a 3D printer and get use of it finds it incredibly useful for those odd jobs around the house, a missing bracket, a cable router, steam deck holder and so on. In the past if I was missing an app or useful software, a game I'd do the lazy thing, even though I can… 22 r/LocalLLaMA community 2d ago Qwen3.8-27b is the first Local model im able to blindly trust You know that thing where you just throw a task at a frontier model and not have to supervise it worrying of it going off course? Qwen3.8-27b has officially gotten me to that point for local work. He has been doing non-stop continuous agentic work for 8+ hours and hasnt screwed… 14 arXiv — Machine Learning research 2d ago LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference arXiv:2609.03079v1 Announce Type: new Abstract: On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload… 36 r/LocalLLaMA community 3d ago Bernie Sanders proposes to ban AI Defined as AI exceeding human cognitive abilities. 20 years in prison. Plenty of local models already fall under that big of an umbrella in some capacities. This is why it's not enough to say that you could torrent open models so who cares what the politicians do. They want you… 4 r/LocalLLaMA community 3d ago I built a local web UI to finetune models on my own text and actually watch the training (works on AMD ROCm) I wanted to do continued pretraining/finetuning of a local model on my own notes and see what's happening while it trains and do it on my AMD card, since most tools assume CUDA. llm-training-panel is a local web UI that: - loads a model from a local dir or a HF id - shows a live… 34 r/LocalLLaMA community 3d ago Can a 4B local model actually feel like an AI assistant? I've been building Arcon around Qwen3-4B + LoRA. Instead of just making it a chatbot, I'm experimenting with persistent memory, personality/mood, internal state, tools, and eventually having it process things before replying. I'm curious what people who've built local agents… 6 r/LocalLLaMA community 3d ago KV cache might be a bigger problem for local models than parameter count Everyone keeps talking about fitting larger models into local hardware, but parameter count isn’t the whole story. For long context inference, KV cache can become the real memory bottleneck. Every new token adds key and value states that need to stay available, so a model that… 8 arXiv — NLP / Computation & Language research 3d ago How Do Prompt Variations Affect Energy Consumption in On-Device LLMs? arXiv:2609.01798v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored. This paper aims to understand how two… 4 r/LocalLLaMA community 3d ago Chrome browser add-on that uses local LLMs to move thousands of unsorted bookmarks into a smart list of automatically calculated categories Is anyone interested in my Chrome browser add-on that uses local LLMs to move thousands of unsorted bookmarks into a smart list of automatically calculated categories?   submitted by   /u/paranoidray [link]   [comments] 7 r/LocalLLaMA community 4d ago What do you do in the meantime when your favourite local model is thinking and working hard with your harness? Real question, whether your are vibe coder or expert or whatever, what do you do? Read carefully each single word? Plan the next steps? Do the gym?   submitted by   /u/takoulseum [link]   [comments] 14 r/LocalLLaMA community 4d ago GLM 5.3 Flash makes a black hole Minecraft mod running locally on 4x RTX PRO 6000 WS saw the post the other day where people said Minecraft clones aren't impressive anymore, because at this point the whole thing might as well be in the training data. so i tried something slightly different, which is asking a local model to write a mod for the real game, using… 5 r/LocalLLaMA community 4d ago Qwen 3.8 Flash Next for Creative Writing? As we all know on of the best local models for creative writing is gemma 4 31b and Muse Glimmer 30b. However, ive been a happy user of Qwen3.8 Flash Next and I wanted to know how well Qwen 3.8 Flash next is doing in terms of creative writing (preferably German).   submitted… 19 Hacker News — AI on Front Page community 4d ago My local model setup on an M4 Pro Mac Mini Article URL: https://lws.io/blog/my-local-model-setup/ Comments URL: https://news.ycombinator.com/item?id=49529132 Points: 244 # Comments: 151 36 r/LocalLLaMA community 5d ago All currently popular local models in one table + Opus 4.8 results If you are thinking what model will fit best your HW specs and tasks you are doing here is one table with all currently popular models that still can be considered as local. LLM Test Scores Feature DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next… 24 r/LocalLLaMA community 5d ago Which current local models that can run within 128GB generate the best SVG pelicans? I used a famous Simon Willison's pelican riding a bicycle prompt on the biggest local LLMs that can run on 128GB Apple Silicon. U used quantizations by Unsloth. Qwen3.8 Flash-Next gives a lot of details. DeepSeek V4 Flash is strangely underwhelming. Qwen3.8 27B still rocks, and… 6 r/LocalLLaMA community 6d ago First time running local models Sad that I only have 12gb of vram but this ik_llama is so fast   submitted by   /u/Needausernameplzz [link]   [comments] 36 r/LocalLLaMA community 6d ago Sliding-window beats linear attention Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic attention with sliding window attention + attention sinks and no post training. This could be big for memory constrained local LLM… 25 r/LocalLLaMA community 6d ago The Chrono Trigger plot challenge - Crono awakens in his modest bedroom of 2095... A hallucinated event, Crono awakens in his modest bedroom of 2095... I am using local models since 2025 January. My daily driver is qwen 3.6 35B A3B, which works pretty well for coding tasks, but I always benchmark the models for lexical knowledge as well, where the models… 38 r/LocalLLaMA community 7d ago I fine-tuned a 0.8B local model for dictation cleanup. It matched a hosted frontier model on this narrow task I built SpeakoFlow Mini, an Apache-2.0 fine-tune of Qwen3.5-0.8B for dictation cleanup. It is not a chat model or a general rewriter. It takes speech-to-text output, applies corrections the speaker actually made, and leaves everything else alone. That last part is harder than it… 21 arXiv — Machine Learning research 9d ago A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models arXiv:2608.26926v1 Announce Type: new Abstract: Small language models (sLLMs) are nowadays hosted on devices with limited memory and computational budget. In an autoregressive setup, inference is memory-bandwidth bound: uniform quantization is often detrimental to such models,… 11 r/LocalLLaMA community 9d ago Ornith-1.5-35B-A3B on 8 GB VRAM: I think I've found my sweet spot A few days ago I posted asking what people considered the best local model for an 8 GB VRAM GPU . At the time, my personal sweet spot was Qwen3.6-35B-A3B , for agentic coding with Pi.dev. Well… Thanks to the suggestions in that thread, I think I've found something even better.… 25 r/LocalLLaMA community 10d ago No, Engrams won't let you run 1T models locally. It does something even better. Ever since Qwen 3.8 Flash Next dropped, there's a misconception going around that N-gram tables will let people run 1T+ parameter models on a single server with 980B parameters offloaded to SSD. I'm here to disappoint you: it won't. But what it will actually do for local models… 31 r/LocalLLaMA community 10d ago I used local Qwen 27b to build a harness and replace OpenCode Sharing my harness for running local LLMs that I built using Qwen 3.x 27B (> 90% locally built) under my supervision - not vibe-coded. Its free, no telemetry, and open-source (AGPL). Works on Windows, Linux (sorry, no Mac yet). I use it for my own coding + mixed workflows. How… 13 r/LocalLLaMA community 10d ago Local LLM harness for reverse engineering software? Does anyone know of a decent reverse engineering harness/workspace setup for local models that I can just point the model at and have it go to work until it's reversed most if not all of the functions in a binary, even if it takes days? Of course, I'm willing to put effort into… 13 Ars Technica — AI news-outlet 11d ago IBM's new Granite 4.2 models ride the wave of interest in local LLMs The focus is on agentic capability and predictable enterprise deployment. 30 arXiv — Machine Learning research 11d ago Renormalization Group Flow Matching for Scalable Local Generative Modeling arXiv:2608.23696v1 Announce Type: new Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. Global approaches can capture full structural coherence but suffer from high computational costs, while local models are… 26 Hugging Face Daily Papers research 12d ago MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Abstract MobilePA-Bench is an interactive sandbox benchmark that evaluates mobile planning agents on tool-calling, sub-agent collaboration, memory usage, and composite skill invocation under real runtime constraints. Generated by thinkingmachines/Inkling-Small As on-device LLM… 13 r/LocalLLaMA community 13d ago Please join r/LowEndLocalAI, a community for running local LLMs on low spec hardware If you’re trying to run local LLMs on a normal laptop, an older desktop, integrated graphics, limited VRAM, or simply the hardware you already own, r/LowEndLocalAI is meant for you. The idea is simple: What useful things can we do with the hardware we already have? I’ve been… 24 r/LocalLLaMA community 13d ago MobileMoE - a facebook Collection MobileMoE is a family of on-device Mixture-of-Experts (MoE) language models with sub-billion active parameters, designed to push the quality–efficiency Pareto frontier for on-device LLMs, including three model scales (S/M/L): 0.3B/0.5B/0.9B active parameters (1.3B/2.8B/5.3B… 7 r/LocalLLaMA community 13d ago What's the best local model you've found for 8 GB of VRAM? I'm curious what other people are using for local LLM coding / agentic coding with only 8 GB of VRAM . My current setup is: Intel Core i7-11800H RTX 3070 Laptop , 8 GB VRAM 32 GB DDR4 RAM openSUSE Tumbleweed / KDE Unsloth Studio pi.dev as the coding agent After testing quite a… 33 arXiv — Machine Learning research 13d ago Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI arXiv:2608.21172v1 Announce Type: new Abstract: Federated fine-tuning enables large language models to adapt on edge devices without centralizing private data, but practical deployments must address hardware instability and adversarial update corruption together. Thermally… 4 r/LocalLLaMA community 14d ago Qwen 3.8 27B for actual local programming Most YouTube benchmarks only show trivial tasks like generating landing pages or simple Three.js games. Is a local model like Qwen 3.8 27B actually capable of real-world systems programming—such as building GTK4 or Qt 6 applications in Rust or C++ with external libraries?… 10 Hacker News — AI on Front Page community 15d ago Why your local LLM feels dumber than it is Article URL: https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917 Comments URL: https://news.ycombinator.com/item?id=49402232 Points: 201 # Comments: 69 24 r/LocalLLaMA community 15d ago Current best model for narrative, chat, prompt creation (so basically everything except agentic coding)? - 5090 Im looking to set up a new local llm (probably on unsloth studio as that seemed to be doing pretty well last time I tested it). This one won't need to do agentic coding or app building or anything (not this time) but instead more 'text' based tasks such as - being given… 10 r/LocalLLaMA community 15d ago Fixed the MTP head on Ornith1.5 35B A3B. +3% TPS -33% wall clock I love the Ornith 35B local models, 1.0 has been running my HAM radio rig for me. I have a hackRF receiver and a 5 watt quansheng portable the both run headless through the PC. I tried out the new Ornith1.5 build and it was faster and more accurate than 1.0. I read the threads… 30 r/LocalLLaMA community 15d ago This is a great sub, regardless of what complaints people have about it. This is a genuine community of real generally respectful adult human beings. Despite the enthusiasm all of you have for local AI, you can recognize that there are times when local LLMs are flawed, and even how practical they are to use for the majority of people to use. Go over… 12 r/LocalLLaMA community 15d ago How to give a local LLM/agent access to a "real" web browser I can't seem to find a good answer to this, my Hermes agent has access to Firecrawl and some other web scrapers for content extraction, but anyone know of a way to let a local LLM drive a "real" web browser? My wife asked me to have Hermes go and look at her LinkedIn profile,… 21 r/LocalLLaMA community 16d ago Ornith-1.5-35B-A3B-NInfer - 250 tok/s, 5-8k prefill, 5090 I tried this model yesterday, and it felt to me like the best one I've tried for a local model for interactive use; the responses and reasoning are very fast, and it actually performs agentic tasks well. The speed is phenomenal. I am running this on Ninfer for Windows -… 37 Hacker News — AI on Front Page community 17d ago Show HN: I trained a 125M model to autocomplete piano on-device I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model… 38 r/LocalLLaMA community 17d ago TinySearch v0.6.1 - still a lightweight web research tool for local LLMs, now with bring-your-own-browser support Hey everyone, Posted TinySearch here a few versions ago and got a bunch of useful feedback, so figured I'd post an update because the thing has changed quite a bit since then. Repo: [ https://github.com/TinySuiteHQ/TinySearch]() The basic idea is still the same: TinySearch is a… 30 r/LocalLLaMA community 17d ago Qwen3.8-27b has the highest level of "agency" I've ever seen in a local model Off a single prompt, given my credentials and the name of my university, qwen3.8-27b was able to successfully pull my class schedule from the kinda shitty and convoluted web of university websites. It needed no human intervention, and executed 80 tool calls. Another time, I… 28 r/LocalLLaMA community 17d ago Anyone NOT on full auto when coding with local LLMs? Would love to know who's letting a 9B just go ham locally, haha But in all seriousness, how many of you are keeping to manual or manual-ish dev workflows?   submitted by   /u/BatPlack [link]   [comments] 23 NVIDIA Developer Blog official-blog 18d ago Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control Robots need policies that can adapt to their sensors, environments, and tasks while running on onboard computing hardware. World models offer a foundation for... 24 Page 1 of 9 · 433 articles Older →