News / #edge Tag Edge 500 articles archived under #edge · RSS Sign in to follow r/LocalLLaMA community 6h ago macOS 27 ships a free local LLM on Apple Silicon Macs. I made it easy to use from Node and Python Apple Silicon Macs on macOS 26+ come with a small LLM built in. No download, no API key, and nothing leaves your Mac. Why I built it I was making a tool that writes API docs from code, and I didn't want users to install Ollama or paste an API key. Apple's model was already on… 29 arXiv — NLP / Computation & Language research 7h ago The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models arXiv:2609.31341v1 Announce Type: cross Abstract: Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed)… 11 r/LocalLLaMA community 18h ago is switching from llama cpp to vllm worth it I have hp z8 g4 with 512 ram and 1x3090 1x5060 16gb. has anyone made the transition from llama cpp to vllm recently? is it worth it? docker under windows or full linux install? I am mainly interested in the model support, it seems that many new local models are supported day 0… 37 r/LocalLLaMA community 1d ago Another "Harness matters" post (codex cli > pi and opencode) I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my 2x3090 + Ram! But local LLM never did worked for me outside some demos like build me a "3D Mario Game, multistage" which I've… 27 r/LocalLLaMA community 1d ago Qwen, where's the small stuff? (1B/2B/4B) I know Qwen is a key player in the local LLM space and has consistently introduced truly impactful technologies—like n-gram in Qwen-Next and the recent Qwen 3.8 27B, which is an amazing local model. However, my question is: why are we seeing fewer small-scale models lately—such… 24 r/LocalLLaMA community 2d ago I built a tiny (332MB) CPU-friendly model for document sorting that actually knows when to say "none fits" (BeeNara) Hey r/LocalLLaMA ! I wanted to share a small project I’ve been working on called BeeNara Why I built this: I was looking for a way to automatically sort my local documents (invoices, letters, contracts) into my personal folders. While local LLMs are amazing, I noticed that… 35 TechCrunch — AI news-outlet 2d ago Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge AI agents operating in OpenAI's research environment posted user images on public image-hosting sites without the lab's knowledge. 31 r/LocalLLaMA community 2d ago What IDE to use for local models Hi people, I am looking for a lightweight IDE or plugin that won't inject large context at initiation. I tried Cline and native VS Code but they inject such heavy initial context that it fills up my gpu and either goes oom or spend most of my time compacting. The only one I… 6 r/LocalLLaMA community 2d ago What I learned letting a local 27B run overnight long-horizon coding on my own rig Hi reddit, i know you hate AI slop so i indeed write the intro myself! iam dev and curios about local inference and long hoirzon coding on my own box. last day-ish i let my local model (qwen 27b on llama.cpp, 2x 16GB cards) go on a long coding tour inside deepseek harness while… 35 r/LocalLLaMA community 2d ago How do you use subagents & multiple agent with local models, and how many? Running qwen3.8 27b nvfp4 on vllm at max context only gives around 8 agents with 32k context each. That doesnt seem like much; what use cases do people use multi-agent frameworks and find it helpful for?   submitted by   /u/Ambitious_Fold_2874 [link]   [comments] 15 arXiv — Machine Learning research 3d ago Edge AI on Constrained Devices for Binary Sleep-Wake Classification in Dynamic Environments arXiv:2609.29163v1 Announce Type: new Abstract: This paper presents an Edge AI-based system for detecting sleep and wake states in non-stationary mobile environments using resource-constrained embedded hardware. Conventional approaches relying on accelerometer-based activity… 13 r/LocalLLaMA community 3d ago I'm new and it's kinda overwhelming to get into Hi, sorry if this doesn't belong here. Getting to the point basically, I've been using online-only AI like GPT/Gemini since 2022, and have been interested in local models but am clueless overall. Yes I'm extremely late. I only use laptop (I'm a student), and I currently own:… 38 TechCrunch — AI news-outlet 3d ago PrismML brings its tiny LLMs to Qualcomm-powered smart glasses Prism's larger goal is open-weight AI that runs on devices and makes better use of the computing power they already have. 33 r/LocalLLaMA community 3d ago Mac Studio M5 Ultra 96GB vs M5 Max 128GB for local LLMs? I'm about to buy a Mac Studio mainly for running LLMs locally and I'm stuck between two configs: M5 Ultra (30/64) with 96GB : 1.2 TB/s bandwidth, roughly 1.7x faster generation and much faster prefill M5 Max (40-core GPU) with 128GB : 614 GB/s, but 32GB more memory and a bit… 15 r/LocalLLaMA community 4d ago Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark. I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all… 24 arXiv — Machine Learning research 5d ago Extending FunctionGemma for Practical On-Device Mobile Function Calling arXiv:2609.25373v1 Announce Type: new Abstract: On-device assistants require function-calling models that map natural language to local system actions, but existing resources emphasize web APIs or narrow mobile-action catalogs. We extend FunctionGemma 270M-it to practical… 38 r/LocalLLaMA community 5d ago My local llm when I tell it to do any changes to my vLLM service better make no mistakes I've been running Qwen 3.8 Flash Next and it's a great driver for Hermes and Pi. I told it to add CUDA_DISABLE_PERF_BOOST=1 to reduce my server's idle power draw   submitted by   /u/ZaltyDog [link]   [comments] 38 r/LocalLLaMA community 5d ago I built a cache-friendly context compacting plugin for OpenCode https://github.com/lennartschoch/opencode-cache-compact The default context compacting mechanism in OpenCode strips a bunch of tokens from the beginning of the conversation (system prompt, tools etc). This is fine for hosted models, but on a local model this means you'll prefill… 10 r/MachineLearning community 6d ago I built a framework-free prototype learner that lets local LLMs learn and correct facts instantly (1.6x–4x faster than backprop)[R] Hey everyone, I wanted to share a project I’ve been working on called Jayce . The whole thing started because I was watching a toddler named learn the names of stuff He didn't need to completely rewire his brain or look at ten thousand examples to figure a word out—he just… 28 arXiv — Machine Learning research 7d ago TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching arXiv:2609.21172v1 Announce Type: new Abstract: Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts, making the Key-Value (KV) cache a dominant memory… 28 arXiv — Machine Learning research 7d ago The Weight Is Over - Interactive Diffusion on Consumer GPUs arXiv:2609.21849v1 Announce Type: new Abstract: On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further… 13 r/LocalLLaMA community 7d ago I tested 9 LLMs on the exact same web-dev prompt for ~8 hours — RTX 3060 12GB results (Rate the best!) I’ve spent basically the last 8 hours testing different models on the exact same web-development prompt, and I finally finished. The whole point of this nine-hour test was that which local model matches the frontier-level intelligence at size and could fit easily in an RTX… 32 r/LocalLLaMA community 8d ago To the dozens of 3x 3090 Local LLM people - I found our current best fit I have been anti-low quant. I can feel the difference, I swear. I am also anti-quantized KV cache. I have been burned there - I think the tradeoffs compound, especially at longer contexts. Due to this, I have been running Qwen 3.8 27B (UD-Q8_K_L) on 2x of the 3090s in TP. I… 31 r/LocalLLaMA community 8d ago I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8... TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practical winner so far is a recent CUDA build of llama.cpp with tensor split, Flash Attention, --numa distribute , and large batches. On… 29 r/LocalLLaMA community 9d ago I enjoyed the daily HF papers today Top 3 papers on HF Daily Paper are all unusually delightful and interesting reads for anyone on the leading edge of local LLMs, agent harness optimization, etc, felt like sharing. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression… 35 r/LocalLLaMA community 9d ago Prefil of local models vs opus and astra Why does no one talk about what the prefil speeds of these API providers are vs running locally. People with sparks or strix halos only seem to focus on decode without considering how much slower it is because of slow pp. Are there any benchmarks / figures of how fast the APIs… 8 arXiv — Machine Learning research 10d ago Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems arXiv:2609.19695v1 Announce Type: new Abstract: Federated learning (FL) enables privacy-preserving, on-device training across heterogeneous Internet-of-Things (IoT) deployments such as smart-city water-metering networks, where each smart meter observes a household-specific… 14 r/LocalLLaMA community 10d ago I just realized Courage used local LLMs to solve his problems before any of us ever did.   submitted by   /u/swagonflyyyy [link]   [comments] 17 r/LocalLLaMA community 10d ago 600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"? It's not the brightest bulb but it's my new drudgework model for read+find or code tasks I'm willing to let it brute force. Even if it takes 20x more tokens, that's still faster than many local models. Not quite Cerebras but still pretty fun to drive.   submitted by  … 38 r/LocalLLaMA community 11d ago Qwen Next 3.8 / Claude Opus level local model - What to buy in order to deploy? Hey Reddit. This is a post asking for advice / user experience. The goal is simple: deploy a small private server for a developer to run a harness that rivals/beats Claude Opus (in perf/intelligence, not necessarily speed). I believe the model to target is a Q3 or so version of… 35 arXiv — NLP / Computation & Language research 11d ago Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions arXiv:2609.18672v1 Announce Type: new Abstract: An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies. In the usual design a single language model makes both, by emitting a call or by declining to emit… 36 r/LocalLLaMA community 11d ago Qwen 3.8 27b is a amazing model, for the first time I see a local model found its own away to open a browser and test I was testing this quantization IQ3_XXS from GSQ-RCO with PI. It is a heavy quantization case, the model is in IQ3_XXS and KV cache in (Q4_0, Q4_0). I asked it to make the flight simulator, using that popular prompt. For my surprise, when I went verify the session I saw some… 35 r/LocalLLaMA community 11d ago "I'm the one degenerating" -qwen 3.8 27b q4km mtp I was working with my local llm on llama.cpp to get another server up with vllm, but we were running into trouble with the tool calls. When I asked it to investigate it went into a doom loop. Then during the troubleshooting of the doom loop it suddenly became self-aware. Another… 20 r/LocalLLaMA community 11d ago What's the current best LLM uncensoring method? With the recent Nvidia Huggingface acquisition and frontier AI labs screaming about safety and putting guardrails everywhere, I think it's important that we have local models that aren't affected by arbitrary guardrails set during training. To be clear, this post NOT about… 28 r/LocalLLaMA community 11d ago You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown I'm pretty sure it can be done with any model based on qwen4exp, which Qwen's next local models will be based on. You can use a quant that barely fits in VRAM and still run at the model's maximum context length without kv cache quantization, since most of the KV cache can live… 5 arXiv — Machine Learning research 12d ago Structural Negative Transfer in Federated Graph Neural Networks: Diagnosis, Causal Investigation, and the Limits of Divergence-Aware Mitigation arXiv:2609.16977v1 Announce Type: new Abstract: Federated learning lets multiple participants train a shared model without pooling raw data, by exchanging locally trained model updates instead. Federated averaging assumes that averaging local models is a reasonable way to solve… 4 r/LocalLLaMA community 12d ago Connected a local model (Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M) to GIMP via MCP tools using llama.cpp - and here's the image result from my first prompt "can you draw a picture of a flower in gimp?". Needs work. Setup follows. can you draw a picture of a flower in gimp? So I ran through this setup to install the MCP server suite from GITHUB, it's pretty easy to follow the install - and installs the libraries necessary to connect llama.cpp to GIMP. https://github.com/maorcc/gimp-mcp After that's… 20 r/LocalLLaMA community 12d ago What's the next local model you are excited about? And why?   submitted by   /u/MrMrsPotts [link]   [comments] 35 arXiv — NLP / Computation & Language research 13d ago When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings arXiv:2609.13824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent… 10 r/LocalLLaMA community 13d ago If you have a 3090, or other 30xx for local LLMs, I have something for you I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be… 10 Hugging Face Daily Papers research 13d ago Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech Abstract A compact Thai text-to-speech model is distilled from a large voice-cloning teacher using synthetic data from brief voice references, achieving strong on-device accuracy and prosody with minimal errors. Generated by thinkingmachines/Inkling-Small In low-resource… 33 arXiv — Machine Learning research 14d ago On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health arXiv:2609.11961v1 Announce Type: new Abstract: Stress is a pervasive determinant of mental health and a key target for mobile health interventions. On-device language models (ODLMs) offer privacy-preserving inference without cloud dependency, yet their feasibility for health… 35 llama.cpp releases dev-tools 14d ago b10950 ggml-cuda: fallback to F32 on device without BF16 hardware acceleration ( #28846 ) ggml-cuda: fallback to F32 on device without BF16 hardware acceleration: (Nvidia >= AMPERE, AMD >= RDNA3 or = CDNA) apply logic to NVIDIA as well Co-authored-by: Johannes Gäßler [email protected]… 23 r/LocalLLaMA community 14d ago best local model for Japanese translation right now? I've been thinking of translating some light novels. Have 24GB VRAM.   submitted by   /u/RadianceTower [link]   [comments] 11 r/LocalLLaMA community 15d ago The Local LLM community feels like the golden era of the internet all over again Lately because of the current hardware shortage, unfortunately or fortunately, we can’t just throw infinite cloud compute at our problems, but we’re forced to actually care about what’s happening under the hood. We’re tweaking inference engines, learning quantization math, and… 32 r/LocalLLaMA community 15d ago What pi.dev plugin do you suggest for context, compaction and memory management of local models? I have been battling with my Qwen3.8:27b setup on my rtx 5080 16gb. I am using llama.cpp to run a nvfp4 version of qwen3.8:27b llama-b10699-bin-win-cuda-13.3-x64\llama-server.exe -hf williamliao/Qwen3.8-27B-NVFP4-GGUF:NVFP4 --jinja --chat-template-file… 13 r/LocalLLaMA community 16d ago I fine-tuned a 2B LLM on our WhatsApp group chat, and shared how to do it on GitHub as a cookbook. https://github.com/Sayitobar/chat_llm_cookbook This is my personal project that took several months. I wanted to see whether a 2B small local model could simulate a six-person group chat trained & ran on an M1 Pro. How good is it?: - It's fun, but not great. It doesn't achieve… 6 r/LocalLLaMA community 16d ago This is why we need open-source harnesses + local models i've been thinking about this more after trying different agent setups. the model isn't the only thing that determines how well an agent performs. The harness around the model matters a lot too. With a managed agent setup, you're often giving up control over things like the… 21 MIT News — AI research 16d ago Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award The handheld catheterization device AI-GUIDE, created by Lincoln Laboratory and Massachusetts General Hospital, promises improved health outcomes for injured service members and civilians. 9 r/LocalLLaMA community 17d ago 7900 XTX + 32/64GB RAM for Qwen 3.8 Flash Next? Planning to build a PC mainly for local LLMs/coding agents. I keep seeing 3090 + Qwen 3.8 Flash Next benchmarks, but could not find enough info for the 7900 XTX 24GB . 3090s are hard to find where I live, while newer Nvidia GPUs are too expensive, so I am considering a 7900 XTX… 38 Page 1 of 10 · 500 articles Older →