News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow llama.cpp releases dev-tools 6d ago b11090 cuda: fix sm_70 tile compilation error ( #29224 ) The 5-argument load_ldmatrix added in 1884824 only defines tile<16,8>, so the Volta tile<8,4> does not match. See #29222 for details. Building on 1884824 , generalize the tile shape of the 5-argument load_ldmatrix from <16,8> to… 14 r/LocalLLaMA community 6d ago Wow, Mimo 2.6 pro seems to be pretty good, but requires more prompting than Sol Although it makes mistakes and use more tokens, but after some reprompting , It(max) gives really good outputs like on par with 5.6 sol and close to astra High in one task.. IT will likely vary on tasks. I hope ds v4.1pro is gonna be really good or at least as good as Kimi K3.… 27 NVIDIA Developer Blog official-blog 6d ago Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability... 18 r/LocalLLaMA community 6d ago Huawei shelves global AI chip rollout as China's own demand outstrips supply — AMD and Nvidia no longer have to worry.   submitted by   /u/fallingdowndizzyvr [link]   [comments] 29 r/LocalLLaMA community 6d ago A better coder for the small-GPU/small-RAM crowd! I’ve been working on making small models more capable at agentic coding and work, because most people in the world don’t have the sort of hardware needed to run 3.8-27B, or even 35B-A3B or 9B dense, and I want to extend local agentic coding capability to less privileged users.… 21 r/LocalLLaMA community 6d ago [MASSIVE RELEASE] Supra2-IMG - a tiny 100M text-to-image model - SOTA quality and open release! Hey everyone! It has been quite a while since the last SupraLabs model - but today we've something special for y'all: Supra2-IMG It's a 100M parameter DiT text-to-image model trained entirely from scratch in under 10 hours on a single H100 on Runpod. It can generate… 32 NVIDIA Developer Blog official-blog 6d ago Turn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2 Weather-sensitive industries increasingly have access to observations that offer an earlier, more local view of changing conditions. Energy companies collect... 13 llama.cpp releases dev-tools 6d ago b11071 ci : Upgrade CUDA to 13.4 for Ubuntu CUDA Release Builds ( #29202 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48951751 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel… 35 r/LocalLLaMA community 6d ago 16GB (and in many cases 12GB) is the max vram most people will ever reasonably have This sub is, needless to say very niche and skewed towards the high end. There are tons of extremely high end setups here with multiple gpu's etc. Even 24GB is out of reach of most people financially, forget about the 3x3090 or 5090 or even higher setups. Macs/Strix Halo/dgspark… 25 r/LocalLLaMA community 6d ago Deterministic Kittens: fun with Qwen Image 2.1 on an M2 Macbook Pro Having a blast with Qwen Image 2.1 on my M2 Macbook Pro with 32GB RAM, even though it takes 16 minutes to complete the recommended number of iterations with the resolution reduced to 1024x1024. And those are some hot minutes! Above are some entertaining outputs. I strongly… 9 llama.cpp releases dev-tools 7d ago b11069 cuda : tune MMVQ to MMQ crossover for SM70 (Volta) ( #28912 ) tune MMVQ to MMQ crossover for SM70 (Volta) Signed-off-by: Yangyu Chen [email protected] Apply suggestion from @JohannesGaessler Apply suggestion from @JohannesGaessler Apply suggestion from @JohannesGaessler… 8 r/LocalLLaMA community 7d ago Gewell - Gemma4 inference engine # the What An engine to run Gemma 4 31B on blackwell under massive concurrency and rather specific workload patterns. I've been waiting for someone to do ninfer but for gemma, and, well, ended up having to do it myself. More models and potentially more gpus are likely to be… 36 llama.cpp releases dev-tools 7d ago b11067 webgpu : add fused gdn + cpy ( #28976 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48873164 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:… 20 vLLM releases dev-tools 7d ago v0.30.0: [Build] Fix DeepGEMM CUDA 12.9 release builds (#57554) Signed-off-by: khluu [email protected] Co-authored-by: OpenAI Codex [email protected] Co-authored-by: Jee Jee Li [email protected] (cherry picked from commit eb87980 ) 15 r/LocalLLaMA community 7d ago What's the verdict on Ternary Bonsai 2 27B? Is it worth switching to from Ornith 1.5 9B? Or is there another similarly sized model that's beating both of them? (I head K2 is good but the KV cache is HUGE). Also btw this is for GPU poor folks so please don't suggest some 30B model.   submitted by   /u/PotterSkxawng… 4 r/LocalLLaMA community 7d ago Where are the current GPU VRAM sweet spots? I have been reasonably satisfied with my single R9700 (32GB) as I can run practical quants of Qwen 3.8-27B at good speeds, as well as other similar models in its weight class (Gemma 4 is still my go-to for general knowledge, until I see something better - has that happened?).… 31 arXiv — Machine Learning research 7d ago IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts arXiv:2609.21346v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many experts contribute knowledge to its output, execution is how many are actually… 17 arXiv — Machine Learning research 7d ago MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems arXiv:2609.21533v1 Announce Type: new Abstract: LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs… 15 arXiv — Machine Learning research 7d ago The Weight Is Over - Interactive Diffusion on Consumer GPUs arXiv:2609.21849v1 Announce Type: new Abstract: On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further… 13 arXiv — Machine Learning research 7d ago Available Guardrails: Certifying Selective Prediction across ML Systems arXiv:2609.22048v1 Announce Type: new Abstract: A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every… 34 arXiv — NLP / Computation & Language research 7d ago From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost arXiv:2609.21117v1 Announce Type: new Abstract: AI productivity is often measured by task completion time, economic value, or improvements in outcome quality. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not… 37 arXiv — NLP / Computation & Language research 7d ago Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency arXiv:2609.21227v1 Announce Type: new Abstract: Factual hallucination is commonly defined by incorrect factual outputs. We study a paraphrase-induced hallucination setting, where a model answers a factual question correctly in its original form but generates an incorrect answer… 4 arXiv — NLP / Computation & Language research 7d ago The Hidden Cost of Digits: Number Normalization and WER in ASR Systems arXiv:2609.21084v1 Announce Type: cross Abstract: Modern automatic speech recognition (ASR) systems trained on extremely large datasets can produce transcripts with numbers written in Arabic numerals. This creates a need for fair comparison with models that output verbatim texts… 9 arXiv — NLP / Computation & Language research 7d ago CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation arXiv:2609.21793v1 Announce Type: cross Abstract: Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine… 35 arXiv — NLP / Computation & Language research 7d ago VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits arXiv:2505.10202v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success but face significant computational and memory challenges, particularly due to their extensive output vocabularies. The final linear projection layer, mapping hidden… 9 arXiv — NLP / Computation & Language research 7d ago Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA arXiv:2602.00279v2 Announce Type: replace Abstract: Reliable uncertainty quantification (UQ) is essential for safe deployment of large language models (LLMs) in scientific question answering, where long-form outputs exceed practical human verification at scale. We introduce the… 17 arXiv — NLP / Computation & Language research 7d ago JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems arXiv:2604.23478v3 Announce Type: replace Abstract: Large language models are widely used to judge the output of other language models, yet whether a judge returns the same verdict when the same request is worded differently remains largely unexamined. We study that question… 23 llama.cpp releases dev-tools 7d ago b11065 CUDA: tune FA for Gemma 4 on Ampere or newer ( #29152 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48802880 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS… 28 r/LocalLLaMA community 7d ago The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks TL;DR: Local agent loop, ~21 days, one RTX 3090. Task was pretty much "build a CUDA inference engine for optimized for yourself on this GPU arch." Got working kernels and benches, not a win over llama.cpp. ~12 human messages. Compaction ate ~83 hours. Old joke: you don’t… 16 Hacker News — AI on Front Page community 7d ago Samsung is expected to more than double output of its HBM4 and HBM4E DRAM Article URL: https://en.sedaily.com/finance/2026/09/20/samsung-to-double-hbm4-output-next-year-sources-say Comments URL: https://news.ycombinator.com/item?id=49778029 Points: 283 # Comments: 190 25 The Information — AI news-outlet 7d ago How Nvidia Is Trying to Solve the Data Center Power Bottleneck Nvidia is tracking “every single gigawatt of land, power and shell around the world, literally everything on the planet,” CEO Jensen Huang said at Goldman Sachs’ annual tech conference earlier this month. “We know where everything is.” Why track power so intently? Nvidia sees… 33 r/LocalLLaMA community 7d ago focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737) I forked llama.cpp to implement Declarative Attention (arXiv:2609.02737, Google DeepMind and KAIST AI). The model declares in its own output which context chunks it needs <focus magic\_chunks="N">), and the engine listens and restricts what the following tokens can attend to. No… 9 llama.cpp releases dev-tools 8d ago b11062 CUDA: enable sparse fa for qwen4 ( #28770 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48739244 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:… 35 r/LocalLLaMA community 8d ago Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3) Hey everyone, After wrestling with hardware and engine issues for days, I finally got Qwen 3.8 Next running properly on my multi-GPU rig. Thought I’d share the setup journey, benchmarks, and thermal results for anyone trying something similar. Seeing all the ongoing memes on… 16 r/LocalLLaMA community 8d ago Radeon RX 10800 XT can outperform the RTX 5090 by 15-25% in 4K gaming and local AI https://en.gamegpu.com/news/zhelezo/radeon-rx-10800-xt-mozhet-obojti-rtx-5090-na-15-25-v-igrakh-v-v-4k-i-lokalnom-ii The more competition, the better!   submitted by   /u/Lumpy_Phase_9539 [link]   [comments] 11 r/LocalLLaMA community 8d ago I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8... TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practical winner so far is a recent CUDA build of llama.cpp with tensor split, Flash Attention, --numa distribute , and large batches. On… 29 r/LocalLLaMA community 9d ago General warning about Clore.AI Hello, I know some of us may be tempted to rent out our expensive GPUs to recoup some of the cost of self-hosting, and it should be obvious that this can be a risky decision. I decided to try hosting my rig on clore.ai briefly to see what kind of revenue it could bring in,… 26 r/LocalLLaMA community 9d ago Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI access to the project files so it could search through them for bugs and stuff. I already had an old computer that I decided not to… 20 llama.cpp releases dev-tools 9d ago b11047 cuda : fix CUB argsort corruption caused by in-place keys ( #28389 ) argsort_f32_i32_cuda_cub called the one-shot DeviceRadixSort::SortPairs API with d_keys_in == d_keys_out (temp_keys, temp_keys). CUB's internal double-buffer ping-pong requires distinct key buffers: with… 22 r/LocalLLaMA community 9d ago Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings: explicit effort level on every role (unset ones defaulted to xhigh) thinking_token_budget of 7500 maxTokens 8k → 32k (file writes were getting cut off) tool output over 10… 23 r/MachineLearning community 9d ago How is RLCD (jev) RL? [D] Just saw the YouTube presentation and I was left wondering this question. If jev only outputs Choice, Score, or Noul … well those are all perfectly differentiable. (Cross entropy or mse) I don’t know if I’m missing something or if adding RL is just for marketing. Like what would… 34 The Information — AI news-outlet 9d ago Nscale IPO Files to Go Public, Shows Huge Revenue Jump And Steep Losses Nvidia-backed startup Nscale disclosed Friday that revenue had surged 10 times in the first half of this year from the year-ago period, but losses mounted as the cloud startup spent billions on data centers and the AI chips they’ll house. The two-year-old UK startup, a spinout… 12 r/LocalLLaMA community 9d ago Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark Planning, coding, testing = 8h total. Stack: VSCode Copilot in autopilot mode + SGLang Stats: ∼10k lines generated, ∼800k tokens consumed Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...   submitted by… 20 TechCrunch — AI news-outlet 9d ago Open or closed AI? Nvidia’s Nader Khalil and Sydney Sykes take on one of the decisions shaping next-gen startups at TechCrunch Disrupt 2026 Nvidia's Nader Khalil and Sydney Sykes discuss one of the decisions shaping next-gen startups on the Builders Stage at TechCrunch Disrupt 2026. 37 llama.cpp releases dev-tools 9d ago b11037 ggml-webgpu: fix supports_op condition for GET_ROWS ( #28978 ) fix get_rows vec4 handling Add src strides checking to vec4_aligned of get_rows and the new test case. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48438627 macOS/iOS:… 31 r/LocalLLaMA community 10d ago Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. Been building this for a few months, mostly for myself, and it just got a proper release so figured I'd post it. It's a native GGUF inference runtime with OpenAI/Anthropic-compatible APIs and a chat UI. The whole point is one consumer NVIDIA card + lots of RAM: MoE models that… 21 arXiv — Machine Learning research 10d ago Compressed Active Subspaces for Scalable Bayesian Inference arXiv:2609.19539v1 Announce Type: new Abstract: Active subspace methods provide a framework for quantifying predictive uncertainty in high-dimensional models by identifying and performing inference along parameter directions that have the greatest influence on the model output.… 5 arXiv — Machine Learning research 10d ago The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability arXiv:2609.19616v1 Announce Type: new Abstract: Complexity measured from generated code is failure-dependent: a difficult prompt can yield a short failing program and be assigned low output complexity. We introduce a six-dimension prompt-side structural-complexity index scored… 14 arXiv — Machine Learning research 10d ago How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU? arXiv:2609.20250v1 Announce Type: new Abstract: Zero-shot essay scoring with large language models is usually demonstrated with proprietary API models, yet the settings where automated scoring is most needed, such as public schools grading thousands of essays under strict… 5 arXiv — NLP / Computation & Language research 10d ago Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds arXiv:2609.19149v1 Announce Type: new Abstract: Subliminal learning shows that language models can transmit a hidden trait through outputs that appear unrelated to it. One proposed explanation, token entanglement, links animal and number tokens through the model's output… 16 Page 3 of 10 · 500 articles ← Newer Older →