News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow llama.cpp releases dev-tools 14d ago b10201 ggml-webgpu: improve flash_attn_vec for quantized KV at long contexts ( #25956 ) improve fa of quantized kv cache Fix some bugs and some comments. fix v type check and some comments Fix build error caused by rebasing editorconfig checking pass Website: https://llama.app… 13 arXiv — NLP / Computation & Language research 14d ago Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories arXiv:2607.27595v1 Announce Type: new Abstract: Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how… 16 arXiv — NLP / Computation & Language research 14d ago (Towards) Scalable Reliable Automated Evaluation with Large Language Models arXiv:2607.28282v1 Announce Type: new Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated metrics often fail to capture the complexity and variability inherent in… 10 arXiv — NLP / Computation & Language research 14d ago IFHierBench: Hierarchical Instruction Following for Large Language Models arXiv:2607.27912v1 Announce Type: cross Abstract: Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the output satisfying specific constraints. Modern deployments increasingly handle… 30 r/LocalLLaMA community 14d ago Open Source Ternary LLM Engine in Rust/CUDA for Quantization, Serving, and Training of models on consumer GPUs, called Tritium (Apache 2.0) This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing models to ternary (1.58 bit) with as minimal of loss as possible, a process that can provide even more than 10x reductions in VRAM usage and much… 25 r/LocalLLaMA community 14d ago Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus) Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without compromising others or making things bulky I’ve been noticing a massive… 14 NVIDIA Developer Blog official-blog 14d ago Run High-Performance Core Math at Scale with NVIDIA nvmath-python NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users... 25 r/LocalLLaMA community 14d ago AMD Lucebox Beats Nvidia DGX Spark by 3.63x on DeepSeek V4 Flash Hey fellow llamas, sorry for posting again this week but i thought this was interesting to showcase to share with y'all. Lucebox partnered up with AMD to bring heterogenous consumer hardware to life. We worked really hard on this, and were able to have Lucebox (AMD Radeon AI PRO… 31 NVIDIA Developer Blog official-blog 14d ago NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We... 9 Hugging Face official-blog 14d ago GPU Management: Why Idle GPUs Are the New Grounded Aircraft Back to Articles a]:hidden"> GPU Management: Why Idle GPUs Are the New Grounded Aircraft Team Article Published July 30, 2026 Upvote - Erick Lachmann ErickvL Dharma-AI Gabriel Pimenta de Freitas Cardoso GabrielPimenta99 Dharma-AI Gustavo Lucchetti gustavolucchetti Dharma-AI… 5 r/LocalLLaMA community 14d ago PR for running Ternary-Bonsai-8B-Q2_0.gguf in llama.cpp with CUDA support just got merged Time to see what it's capable of   submitted by   /u/413205 [link]   [comments] 37 r/LocalLLaMA community 14d ago unsloth/Qwen3.6-27B-NVFP4 vs. Intel/Qwen3.6-27B-int4-AutoRound vs. nvidia/Qwen3.6-27B-NVFP4 -- which one to choose? Are there any benchmarks on these 4 bit quants, like how Artificial Analysis runs a slew of various benchmarks? If not, how can I run one (5x over for consistency) on them? I'm also very interested in hallucinations, as community discussions seem to point them out.  … 15 llama.cpp releases dev-tools 15d ago b10188 metal: fix memory unwire if model is freed without any GPU operations ( #26082 ) metal: fix memory leak if model is freed without any GPU operations metal: run dummy work only if residency sets are used metal: wrap function in #if defined metal: measure system-wide wired memory… 6 r/LocalLLaMA community 15d ago 4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s I've been benchmarking a two-card box for a few weeks and I still can't quite get over some of these numbers, so I'm dumping them here. Box: RTX 4090 (24GB) + RTX 5060 Ti (16GB), i9-13900K, 64GB DDR5. WSL2 with 47GB allocated to the VM, CUDA 12.8 (12.8 specifically,13.1… 21 arXiv — Machine Learning research 15d ago From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs arXiv:2607.26571v1 Announce Type: new Abstract: The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. However, direct measurement of inference energy often… 14 arXiv — Machine Learning research 15d ago Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models arXiv:2607.26117v1 Announce Type: cross Abstract: Self-repair - returning a failed program to the model together with its test output and asking for a correction - is a standard component of code agents, and is almost always evaluated against a baseline that does not retry at… 10 arXiv — NLP / Computation & Language research 15d ago Contrastive ESA: Human Evaluation of Multiple Translations at Once arXiv:2607.26640v1 Announce Type: new Abstract: Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that… 29 Hugging Face Daily Papers research 15d ago CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Abstract Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level… 5 r/LocalLLaMA community 15d ago 3090 owners, what vram tempature do you get under ai load? Hello Can you please share the tempature you get on your rtx 3090 under active llm load? Im trying to findout if my rtx 3090's tempatures are healthy or not please share VRAM Tempature only, you can track it via gpu-z on windows   submitted by   /u/Whole_Alternative_18… 9 r/LocalLLaMA community 15d ago Budget Inference: A GPU for dense models vs. More RAM for MoE models? Hi all, I’m building a budget inference machine primarily for personal use (chat/assistant tasks, possibly some RAG). I'm torn between two hardware paths and would love input from anyone who has actually benchmarked these setups. The Dilemma: Option A (GPU for dense models): Buy… 35 NVIDIA Developer Blog official-blog 15d ago How to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo Guardrails Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source... 33 Dwarkesh Podcast news-outlet 15d ago Why compute might get 10x+ more expensive in coming years If a human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot price. 19 r/LocalLLaMA community 15d ago Those who use many layers in CPU/RAM and some in GPU - what are your specs and speeds? I am trying to figure out if it's worth upgrading my RAM, but I've noticed that some MoE models don't seem to do well with many layers shared from VRAM --> CPU/RAM. This may be something on my end; a software config or perhaps my specific hardware config. This made me curious as… 17 r/LocalLLaMA community 15d ago The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU). On the CPU, batch 1 is memory bandwidth bound. But if token/s = bandwidth / (bytes_per_weight * active_weights_per_token) the total number of parameters doesnt slow down the generation speed. So building the architecture aroud a small batch "active parameters per token" (ternary… 17 r/MachineLearning community 16d ago Vendor-agnostic ML inference on production edge devices [R] I work on PostSlate, a video editing tool, and this comes out of our own work. We run ML models on-device, face detection and embedding among other things, which means we can't assume anything about the user's GPU. NVIDIA discrete, AMD, Intel integrated, Apple Silicon, all of… 7 r/LocalLLaMA community 16d ago I tried running a 1.56TB MoE model on a 6GB RTX 4050 Laptop, Here’s the result The Test Bench Setup I tested running a massive 1.56TB Mixture-of-Experts (MoE) checkpoint (96 shards, 93 layers, 896 experts/layer, ~4.46 bits/param MXFP4) on a budget gaming laptop. Laptop: HP Victus 15 GPU: NVIDIA RTX 4050 Laptop (6GB GDDR6, 96-bit interface @ 192 GB/s… 21 r/LocalLLaMA community 16d ago Nvidia is expected to raise GeForce RTX GPU prices again by up to 30%   submitted by   /u/ab2377 [link]   [comments] 21 r/LocalLLaMA community 16d ago SK Hynix stock fell some 40% in the last 30 days, finally cheap RAM and GPUs again? They actually halted trading on the Korean stock exchange today. Finally some hope? And do you think the ruptures in the Korean market will finally free up supply again, and we can finally go back to normal? Or are we doomed to continue the hardware-starved life we endured for… 28 NVIDIA Developer Blog official-blog 16d ago Developing Healthcare Robotics with GPU-Native Medical Physics Simulation Unlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation.... 28 llama.cpp releases dev-tools 16d ago b10172 ggml-webgpu: Fix some binding alias issues to support all archs, fix recurrent-state-rollback test ( #25931 ) Add overlap glu variant to support all archs, fix recurrent-state-rollback test format Fix all arch overlapped ranges format diagnose bus error on apple ci More testing… 25 llama.cpp releases dev-tools 16d ago b10166 ggml : set output of view src ( #25729 ) llama-graph: set_outputs to t->view_src change set_output to GGML_ASSERT about views not being outputs sampler : avoid views in outputs cont : fix dist sampler cont : consistent logits handling ggml : set output of view src graph :… 28 llama.cpp releases dev-tools 16d ago b10164 ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration ( #22675 ) ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration cuda: added SSD CICD fixes for CUDA / HIP / MUSA / MSVC. ggml-cuda: review comments fixed. ggml-cuda: Fuse M matrix materialization… 23 r/MachineLearning community 17d ago NeurIPS 2026 AI-generated reviews [D] I'm really confused about what the point of the prompt injection was (speaking as an author). Is it just a study? I would really prefer that they took action against the AI-generated reviews. Obviously, we cannot assume that the reviewers were copy-pasting the output from the… 5 r/LocalLLaMA community 17d ago I've been tracking RTX 5090 prices across EU stores since March, it's up €1,061 and still climbing Been running a GPU price tracker ( https://www.pricesquirrel.com ) since March, covering 20+ EU stores, recently added RAM, SSDs and CPUs too. Every GPU tier has gotten cheaper since launch. The RTX 5090 has done the exact opposite. The data: The ASUS TUF Gaming RTX 5090 OC was… 34 Hugging Face Daily Papers research 17d ago Characterizing Warp Divergence from Pascal to Blackwell Abstract Since Volta introduced Independent Thread Scheduling (ITS), NVIDIA GPUs have been widely assumed to handle warp divergence in a fixed manner. We test this assumption across Ampere, Hopper, and datacenter and consumer Blackwell GPUs, using pre-ITS Pascal as a baseline.… 23 r/MachineLearning community 17d ago Are single GPU research still published in ML/DL and its applications nowadays? Which are the most notable recent ones? [D] ML research is progressing at breakneck speed where frontier labs in both academia and industry have access to considerably large computes (GPUs). Where do small labs or independent researchers go in this context? Have you come across recent works in ML/DL and its applications… 26 Hugging Face Daily Papers research 17d ago Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels Abstract Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the… 13 r/LocalLLaMA community 17d ago microsoft/VibeVoice-ASR-BitNet VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 GB to 1.58 GB while achieving 1.6–2.3× faster inference than Whisper.cpp with… 17 arXiv — Machine Learning research 17d ago In-Context Learning as Implicit Policy Gradient arXiv:2607.23153v1 Announce Type: new Abstract: Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings,… 35 arXiv — Machine Learning research 17d ago When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation arXiv:2607.23390v1 Announce Type: new Abstract: When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system over a fixed horizon as a pure schedule selecting fields from a declared… 4 arXiv — Machine Learning research 17d ago Restoration Flow Matching-Based Channel Refinement and Equalization Correction for MIMO Semantic Communications arXiv:2607.23615v1 Announce Type: new Abstract: In multiple-input multiple-output (MIMO) semantic communication, imperfect channel state information (CSI) and equalization mismatch can seriously degrade semantic reconstruction quality. To address this issue, we propose a unified… 26 arXiv — NLP / Computation & Language research 17d ago Not All LLM Reasoning is Visible in the Chain-of-Thought arXiv:2607.22925v1 Announce Type: new Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically… 16 arXiv — NLP / Computation & Language research 17d ago Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models arXiv:2607.23067v1 Announce Type: new Abstract: Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature layers. However, DoLa's dynamic layer selection relies solely on… 31 arXiv — NLP / Computation & Language research 17d ago LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction arXiv:2607.23420v1 Announce Type: new Abstract: Large language models show strong promise for information extraction (IE), but existing reflection-based correction methods are often misaligned with structured extraction outputs. Free-form self-reflection can flag an error, yet… 31 arXiv — NLP / Computation & Language research 17d ago A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever arXiv:2607.23806v1 Announce Type: new Abstract: Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows… 31 arXiv — NLP / Computation & Language research 17d ago Understanding Tone-Dependent Inference Cost in Large Language Models arXiv:2607.23915v1 Announce Type: new Abstract: We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570… 37 arXiv — NLP / Computation & Language research 17d ago Pointer-Augmented Autoregressive Generation of Patent Claims with Joint Topology and Content Decoding arXiv:2607.24040v1 Announce Type: new Abstract: Autoregressive decoders emit flat token sequences and cannot enforce hierarchical constraints across output segments, a limitation that becomes acute in patent claim generation, where a claim set forms a dependency forest whose… 38 arXiv — NLP / Computation & Language research 17d ago CAGE: Cognitive Attribution Graphs for Faithful Inline Citation Generation in Long-Form Question Answering arXiv:2607.24236v1 Announce Type: new Abstract: Long-form question answering increasingly relies on retrieved evidence to make LLM outputs verifiable, with inline citations tracing claims to source documents. However, existing systems often attach citations that are topically… 12 arXiv — NLP / Computation & Language research 17d ago Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets arXiv:2607.24268v1 Announce Type: new Abstract: Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation… 6 arXiv — NLP / Computation & Language research 17d ago Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs arXiv:2607.24371v1 Announce Type: new Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic coding, CPT for procedure billing, and HL7 FHIR for data exchange. While large… 22 Page 5 of 10 · 500 articles ← Newer Older →