News / #edge Tag Edge 373 articles archived under #edge · RSS Sign in to follow Hacker News — AI on Front Page community 3mo ago Show HN: Find the best local LLM for your hardware, ranked by benchmarks Article URL: https://github.com/Andyyyy64/whichllm Comments URL: https://news.ycombinator.com/item?id=48146369 Points: 224 # Comments: 38 21 r/LocalLLaMA community 3mo ago What is the most unexpected thing you have gotten a local model to do? Most local LLM use cases I see are chat, coding, and RAG. But with vision models getting better and faster on consumer hardware, I feel like there is a lot of untapped territory. I got a local VLM to play a board game by just looking at the screen and it worked way better than I… 25 r/LocalLLaMA community 3mo ago Used over a million tokens in three separate sessions to test Qwen 3.6 35b (new Multi-token Prediction version) In my opinion, MTP models are 100% game changer for local LLMs. In terms of speed, I was getting around 1.5x the tok/sec of previous tests. The project was a test - building a full iterative step-by-step pygame; a small mystery dungeon-style game. At first I set 100-200k context… 28 arXiv — Machine Learning research 3mo ago Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization arXiv:2605.14373v1 Announce Type: new Abstract: Zeroth-Order (ZO) optimization is pivotal for scenarios where backpropagation is unavailable, such as memory-constrained on-device learning and black-box optimization. However, existing methods face a stark trade-off: they are… 7 r/LocalLLaMA community 3mo ago club-5060ti: practical RTX 5060 Ti local LLM notes and configs I put together a small public repo for RTX 5060 Ti 16GB local LLM setups: I took inspiration from the club-3090 repo, but this one is focused on documenting what we’ve actually tested on 5060 Ti hardware so the setup details are easier to share and reproduce. Current seed setup… 6 r/LocalLLaMA community 3mo ago A VERY lightweight open web-search tool for smaller local LLMs Hey everyone, Been playing around with local agent setups lately, mostly Cline/Roo with smaller models, and web search kept annoying me. Not because it doesn’t work, but because it usually throws way too much random page text into the context. small models really don’t handle… 29 r/LocalLLaMA community 3mo ago Got local Qwen 3.5/3.6 generating meeting summaries entirely offline on an M4 Max. Demo with Wi-Fi off. This is the future. I'm the founder behind Hedy, an AI meeting app. I'm a huge supporter of Local AI, and we've been working on making it "consumer friendly". Speech recognition in Hedy has always run on-device (whisper.cpp and now also parakeet). What just shipped is that the rest of the AI… 22 r/LocalLLaMA community 3mo ago Anyone actually using a local LLM as their daily knowledge base? Not for coding, for life stuff. What's your setup? So I've been going down a rabbit hole lately and I can't find many people actually talking about this specific use case. everyone here runs local LLMs for coding, chat, maybe some creative writing. cool. But what about using it as a proper personal knowledge base? like, dump… 24 r/LocalLLaMA community 3mo ago The "the future is fictional" problem of many local LLMs Many local models have a problem (that raised due to excessive RHLF training): They mostly think that everything that is beyond their knowledge cutoff date would be "fictional" or "satirical". To be fair: Even the Gemini API without web access can have this sometimes. But it… 20 r/MachineLearning community 3mo ago Your AI Use Is Breaking My Brain: Why 10 Minutes of Prompting Fries Us[D] It’s 2:30 AM. My youngest just woke up crying for water, completely derailing my train of thought while I was trying to debug a weird edge case in a side project. I stared at my IDE, then at my local model running in the terminal, then back at the IDE. My brain felt like… 26 r/LocalLLaMA community 3mo ago Small local model for questions on German grammar I'm trying to learn German. I use Qwen3.5/3.6 locally, but this is pretty bad for German grammar. Has anyone got a recommendation for a small-ish local model that knows German grammer well and can answer questions on this? EDIT: I give an example output from unquantized Qwen3.5… 38 arXiv — Machine Learning research 3mo ago A Comparative Study of Federated Learning Aggregation Strategies under Homogeneous and Heterogeneous Data Distributions arXiv:2605.11010v1 Announce Type: new Abstract: Federated Learning has emerged as a transformative paradigm for collaborative machine learning across distributed environments. However, its performance is strongly influenced by the aggregation strategy used to combine local model… 17 r/LocalLLaMA community 3mo ago I've seen a lot of folks ask "can local LLMs actually do anything useful?" And I'm here to share my experience. The answer is resoundingly 'yes'. Let me start with the local model I use every day in my AI harness: embedding models. I'm using an embedding model to give my AI's persistent memory system a semantic search protocol that makes its memory… 37 r/LocalLLaMA community 3mo ago Local LLM autocomplete + agentic coding on a single 16GB GPU + 64GB RAM Today I set up a full coding toolbox on a single RTX 5080 (with RAM offloading) that's actually viable. Autocomplete : bartowski/Qwen2.5-Coder-7B-Instruct-GGUF:Q6_K_L Agentic : unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q8_K_XL Why these models: Qwen2.5 is still the best model for infill… 9 Smol AI News news-outlet 4mo ago not much happened today **Gemma 4** was launched by **Google** under an **Apache 2.0 license**, marking a significant open-model release focused on **reasoning, agentic workflows, multimodality, and on-device use**. It outperforms models 10x larger and has immediate ecosystem support including… 35 NVIDIA Developer Blog official-blog 4mo ago Bringing AI Closer to the Edge and On-Device with Gemma 4 The Gemmaverse expands with the launch of the latest Gemma 4 multimodal and multilingual models, designed to scale across the full spectrum of deployments, from... 27 Hugging Face official-blog 4mo ago Welcome Gemma 4: Frontier multimodal intelligence on device Back to Articles Welcome Gemma 4: Frontier multimodal intelligence on device Published April 2, 2026 Update on GitHub Upvote 891 merve merve Pedro Cuenca pcuenq Sergio Paniego sergiopaniego ben burtenshaw burtenshaw Steven Zheng Steveeeeeeen Alvaro Bartolome alvarobartt Nathan… 9 NVIDIA Developer Blog official-blog 4mo ago NVIDIA IGX Thor Powers Industrial, Medical, and Robotics Edge AI Applications Industrial and medical systems are rapidly increasing the use of high-performance AI to improve worker productivity, human-machine interaction, and downtime... 13 NVIDIA Developer Blog official-blog 5mo ago CUDA 13.2 Introduces Enhanced CUDA Tile Support and New Python Features CUDA 13.2 arrives with a major update: NVIDIA CUDA Tile is now supported on devices of compute capability 8.X architectures (NVIDIA Ampere and NVIDIA Ada), as... 18 Import AI news-outlet 5mo ago Import AI 448: AI R&D; Bytedance's CUDA-writing agent; on-device satellite AI If Ukraine is the first major drone war, when will there be the first major AI war? 6 NVIDIA Developer Blog official-blog 5mo ago How to Minimize Game Runtime Inference Costs with Coding Agents NVIDIA ACE is a suite of technologies for building AI agents for gaming. ACE provides ready-to-integrate cloud and on-device AI models for every part of in-game... 23 Google DeepMind official-blog 13mo ago Gemini Robotics On-Device brings AI to local robotic devices We’re introducing an efficient, on-device robotics model with general-purpose dexterity and fast task adaptation. 33 Google DeepMind official-blog 15mo ago Announcing Gemma 3n preview: Powerful, efficient, mobile-first AI Gemma 3n is a cutting-edge open model designed for fast, multimodal AI on devices, featuring optimized performance, unique flexibility with a 2-in-1 model, and expanded multimodal understanding with audio, empowering developers to build live, interactive applications and… 23 Page 8 of 8 · 373 articles ← Newer