News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow r/LocalLLaMA community 4h ago Planning to get a cheap-ish GPU. Would appreciate some advice. Hi, I've been wanting to run my own local LLM for some time and I finally saved enough to get a budget GPU. I can spend 700$ at most, and I'm looking for a GPU that can run quantized 30B~ parameter models with decent speed. Being able to run Qwen 3.8 27 B Q4_K_M and similar… 14 llama.cpp releases dev-tools 4h ago b10826 cuda: fixes races in mmid and mmf ( #28475 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/45586039 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework… 14 r/LocalLLaMA community 8h ago 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics This comparison was requested by a few people, and certainly we were all eager to see the final results. The comparison had taken 11 days and the GPU was crunching numbers for ~167 hours. We've been comparing different abliterated models from huggingface to see if they really… 10 r/LocalLLaMA community 10h ago 48 tg/s 440 prefill on my grandma's cluster (2xP40) (sort of) TL;DR: switching KV cache to f16 may give a boost in speed if using MTP and ngrams. I have a self-built "AI mega-cluster" with 2x P40s on a cheap Chinese motherboard and a Xeon CPU (around $1,100 to build, including water cooling for the GPUs). I was normally getting up to 15… 31 r/LocalLLaMA community 19h ago Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant So after all my work, yeah, Raymond did it better, so I ported his work over, extended it turboX, extended it multiple other models (he had only Qwen models), and benchmarked the crap out of it to make sure it was worth it still. So really the credit goes to Raymond (… 9 r/LocalLLaMA community 21h ago My only real use case for a local AI use is document management, how much VRAM do I realistically need for a good experience? I just want to use paperless-ai and be able to ask questions relative to it. Bonus points if I could use it with home assistant but that's not the focus. I just can't see needing a 32 GB VRAM GPU for just that, but I don't want to buy a GPU only to find out that "yeah, it's… 11 Simon Willison community 22h ago Introducing GPT-6 Astra for developers Introducing GPT-6 Astra for developers Blink and you'll miss it, but there's a familiar creature at 1m59s : Across the board, Astra has more attention to detail, better understanding of the user's prompt, and can build more sophisticated outputs. In particular, it excels at… 7 r/LocalLLaMA community 23h ago Is 3090 + 5070 & 5060s a good idea? I have a 5070 Ti and two 5060 Ti (all 16Gb cards). I planned to add another 5070 Ti giving me two pairs of 32Gb each but the NVidia prices have just jumped by 25% where I am and I've found a 3090 Founders Edition for a good chunk cheaper than the 5070 would cost. It's 8Gb more… 8 r/LocalLLaMA community 1d ago My local LLM demoscene generator can now watch its own output and rewrite it! I've updated my auto_demo_scener project with Ninfer support and a “rewrite based on video” feature that I thought you might find interesting. The project is basically an endless demoscene machine. A local LLM writes Three.js effects (from a library of editable prompts), you… 23 r/LocalLLaMA community 1d ago Best way to run Qwen3.8-27B on a system with a RTX 5090 + RTX 5070 Ti (32GB + 16GB)? I have a system with 2 GPUs and 48GB VRAM total, a RTX 5090 + RTX 5070Ti. What would you say is the best way to run Qwen3.8-27B on that system with the best quality and 262k context? Would just the normal llama.cpp work with how it detects and does its own magic with dual CPU… 28 The Information — AI news-outlet 1d ago Tech’s Phone-Free, Viking-Filled Wedding of the Summer • Our very first Fall Preview package ! Nvidia’s new sales chief, Anthropic’s mega IPO—and 12 other things that will define Silicon Valley over the next three months • Plus, Recommendations—our weekly pop culture picks: “ Dolly Parton’s America ,” “ Money to Burn ” and “ Coyote… 15 r/LocalLLaMA community 1d ago Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal? Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4 — is this config optimal? Hardware CPU: Intel Core i5-12600K RAM: 128 GB DDR4 @ 3600 MHz GPU: NVIDIA RTX 3090, 24 GB VRAM OS: Windows 11 llama.cpp: freshly compiled from today's master (build b10794, Sep 4… 30 r/LocalLLaMA community 1d ago NVIDIA PAIR — Your Personal AI Cluster That is interesting, I got bunch of old hardware I could connect, wonder what the speed would looks like.   submitted by   /u/SpendLucky1273 [link]   [comments] 22 r/LocalLLaMA community 1d ago Help me understand gguf size/ctx size Let's say I have 2x 16Gb GPUs and I want to run Qwen3.8 27B. Monitor is ran by the integrated GPU so both 16Gb GPUs are almost fully free. I load the UD-Q4_K_S on one card at 15.4Gb. I then load the context on the other card? Would that be the most efficient way? Or should I aim… 20 r/LocalLLaMA community 2d ago Am I the only one having these problems with downloading models from HF? https://preview.redd.it/m5kved2a5knh1.png?width=469&format=png&auto=webp&s=780011abb19cb79fa4dc64032c86ff858d879786 I don't have problems with Nvidia buying HF, but I have problems with the fact that lately HF became almost unusable. It is around one month that I experience big… 24 r/LocalLLaMA community 2d ago I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code ( C code), the results were not completely unexpected but some quants were definitely underwhelming. TLDR : Best overall: bartowski/Qwen3.8-27B-IQ4_XS . Best… 32 NVIDIA Developer Blog official-blog 2d ago Building a Memory-Driven Agent with NVIDIA NemoClaw Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it... 26 The Information — AI news-outlet 2d ago Abu Dhabi’s G42 Considers US Ownership to Safeguard AI Chip Access G42, the Abu Dhabi artificial intelligence firm at the center of the United Arab Emirates’ AI push, has discussed handing majority control to American companies in order to keep buying the most advanced chips next year, including Nvidia’s most advanced H100 chips, Bloomberg… 12 r/LocalLLaMA community 2d ago Georgi Gerganov on the Nvidia acquisition Link: https://x.com/ggerganov/status/2095897173376618881   submitted by   /u/CombinationKitchen76 [link]   [comments] 27 NVIDIA Developer Blog official-blog 2d ago Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run... 17 TechCrunch — AI news-outlet 2d ago Apple’s Ternus era begins as Nvidia bets on the whole AI stack It’s officially the Ternus era at Apple.   Tim Cook stepped down as CEO this week, handing the company to former hardware chief John Ternus, whose first memo promised a “huge launch next week” — timing that puts Apple’s next iPhone event… 5 llama.cpp releases dev-tools 2d ago b10798 common : make build info output stream configurable ( #28322 ) Let llama_print_build_info write to a caller-provided FILE* instead of hardcoding stderr. The parameter defaults to stderr so existing callers keep their current behavior. The version command in llama-app now passes… 23 The Information — AI news-outlet 2d ago Nvidia’s New Sales Chief, Anthropic’s Mega IPO—and 12 More Things That Matter in Tech This Fall This summer wasn’t exactly uneventful in Silicon Valley: No stretch of time in the ever-changing AI age could possibly be described as such. But with everyone back at their desks, the next few months ahead seem absolutely chockablock with action—most notably, Anthropic’s… 30 r/LocalLLaMA community 2d ago NVIDIA's $12,930,300,000.00 acquisition of Hugging Face contains an easter egg. The first 6 numbers of the acquisition price represent the decimal conversion of Unicode character U+1F917. The 🤗 emoji. From: Polymarket on 𝕏: https://x.com/Polymarket/status/2095646821485842805 Julien Chaumond on 𝕏: https://x.com/julien_c/status/2095822387895824836   submitted by   /u/Nunki08 [link]   [comments] 30 The Information — AI news-outlet 2d ago DeepSeek Plans Major Huawei Chip Order in New AI Data Center DeepSeek plans to install at least 160,000 Huawei AI chips at a data center in Inner Mongolia, Northern China, Bloomberg reported, citing people familiar with the situation. The project would support China’s push to replace Nvidia silicon in light of U.S. chip restrictions.… 11 r/LocalLLaMA community 2d ago We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0) I'm one of the developers. We said in August it would go open source in September and it did last night. MIT or Apache-2.0, pick one. The repo you see is our internal repo, kernels included, so from now on everything happens in public. It's an inference engine in Rust and C++… 25 r/LocalLLaMA community 2d ago On GPT-6 Astra 98.6% ARC AGI-3: don't fall for the hype Here is the news you may have missed: Nvidia already demonstrated 100% on ARC AGI-3 , using their novel harness AVO:… 6 The Information — AI news-outlet 2d ago Nvidia Discusses $2.5 Billion Investment in Mira Murati’s Thinking Machines Lab Mira Murati’s Thinking Machines Lab is in talks to raise between $5 billion and $6 billion at a pre-money valuation of at least $40 billion, with venture firm Accel in discussions to lead the round, The Information reported . Chipmaking giant Nvidia is expected to contribute… 5 llama.cpp releases dev-tools 2d ago b10796 src : add n_expert_used_max function ( #28323 ) src : add n_expert_used_max function With Commit c61b98b ("model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support ( #25444 )") it is now possible for each layer to have a specific number of experts but there are a… 38 arXiv — Machine Learning research 2d ago RATL: Learning from Retrieved Residuals for Robust Multivariate Time-Series Forecasting arXiv:2609.03937v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) complements parametric models with retrieved external evidence. The same idea is attractive for continuous-output regression, but directly reusing retrieved target values is often not robust… 27 arXiv — Machine Learning research 2d ago OSR: Output Space Redistribution for Adaptive Label Removal in Classification Models arXiv:2609.03972v1 Announce Type: new Abstract: Label removal occurs frequently in classification systems with evolving taxonomies, where categories must be dynamically updated or eliminated. To accommodate such changes, classification models must adapt accordingly. Existing… 27 arXiv — Machine Learning research 2d ago You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring arXiv:2609.03035v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fooled by obfuscation and steganography, motivating… 38 arXiv — NLP / Computation & Language research 2d ago Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards arXiv:2609.03181v1 Announce Type: new Abstract: We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters… 9 arXiv — NLP / Computation & Language research 2d ago How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models arXiv:2609.03322v1 Announce Type: new Abstract: Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate… 18 arXiv — NLP / Computation & Language research 2d ago Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT arXiv:2609.03366v1 Announce Type: new Abstract: Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output may be ungrounded, incomplete, or unfaithful to the decision process. Achieving accountability requires verified… 33 arXiv — NLP / Computation & Language research 2d ago Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation arXiv:2609.04048v1 Announce Type: new Abstract: Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic… 9 vLLM releases dev-tools 2d ago v0.29.0rc3 [CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1… 33 r/LocalLLaMA community 2d ago UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM, everything else on the GPUs. Since… 26 The Information — AI news-outlet 2d ago Nvidia’s Hugging Face Embrace Burnishes AI Kingmaker Status Nvidia’s AI chip design rival, Broadcom, may be forecasting rocketing growth . But Nvidia CEO Jensen Huang needn’t worry. He’s deftly parlayed Nvidia’s early AI winnings into what’s likely to be an enduring role as kingmaker in the sector, as he demonstrated by confirming the… 36 r/LocalLLaMA community 2d ago Qwen 3.8 27B Vs. Qwen 3.6 27B on oMLX Quality: 81.1 → 87.7 (+8%) Speed: 35 → 29 tok/s (−16%) Runtime: 8m51s → 44m39s (5x longer) Output tokens: 18K → 78K (🤯) Noticeably better quality, but you're paying for it with tokens and time. Full benchmark results (all hardware, all quants): llm-bench.io Qwen3.6-27B Vs.… 28 TechCrunch — AI news-outlet 3d ago Meta is paying to peek at how you use their latest AI model For its new Muse Spark model, intended for operating coding and other agents, Meta is offering an explicit discount averaging out to about 95% for users who "contribute" to the development of future models by sharing their prompts and model outputs. 13 r/LocalLLaMA community 3d ago Micron Explores Near-GPU NAND Flash to Run Bigger LLMs I would be really curious about this especially on unified memory devices.   submitted by   /u/giveen [link]   [comments] 25 r/LocalLLaMA community 3d ago "ModelScope" Is a Hugging Face Alternative now that Nvidias deal is a Go I liked the Nvidia that focused on just GPUs for gaming, not on the Nvidia of today which seem want power consolidation. Modelscope is another platform for those that simply want to know an alternative if things go south. However, time will tell what happens to huggingface after… 6 NVIDIA Developer Blog official-blog 3d ago NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents.... 8 Ars Technica — AI news-outlet 3d ago Nvidia buys Hugging Face, the GitHub of AI, for $13 billion Nvidia says Hugging Face will stay open even as the chipmaker takes control of a key AI hub. 20 TechCrunch — AI news-outlet 3d ago Nvidia confirms it will buy Hugging Face for $12.9 billion Nvidia said Hugging Face hosts over 3 million models and is used by over 18 million developers. 12 The Information — AI news-outlet 3d ago Nvidia to Buy Hugging Face for $12.9 Billion Nvidia has agreed to buy AI model platform Hugging Face for $12.9 billion, the latest step by the chip giant to use its financial heft to exert greater control in the AI ecosystem. Nvidia’s announcement confirms The Information’s report last week that Nvidia and Hugging Face had… 28 r/LocalLLaMA community 3d ago It's official! Nvidia to acquire Hugging Face for 12.9 billion dollars.   submitted by   /u/SarcasticBaka [link]   [comments] 8 Hacker News — AI on Front Page community 3d ago Nvidia to Acquire Hugging Face Article URL: https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/ Comments URL: https://news.ycombinator.com/item?id=49548952 Points: 211 # Comments: 58 16 Hacker News — AI on Front Page community 3d ago Nvidia to acquire Hugging Face https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face... Comments URL: https://news.ycombinator.com/item?id=49548952 Points: 271 # Comments: 84 19 Page 1 of 10 · 500 articles Older →