News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow r/LocalLLaMA community 6d ago parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM High-performance inference of NVIDIA's Parakeet TDT 0.6B V2 English transcription model, in the browser. Check out the live demo: https://parakeet.narcotic.sh/ A fully custom, dependancy-free implementation with raw WebGPU compute shaders and SIMD WebAssembly audio frontend. 1… 19 r/LocalLLaMA community 6d ago Am I just hallucinating Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running the same prompts), it's all just vibes. Some context, I'm running the latest… 27 llama.cpp releases dev-tools 6d ago b10307 sycl: fix UE4M3 parsing ( #25608 ) The NVFP4 quantization format stores a scaling factor for every group of 16 weights, packed into a single UE4M3 byte. The SYCL GPU code was converting these scale values using the E4M3 path, but that's signed , and these are unsigned values.… 13 r/LocalLLaMA community 6d ago RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Specific GPUs GitHub : https://github.com/humza-khalid/12vhpwr-guard Reddit thread : https://www.reddit.com/r/nvidia/comments/1vglua1/i_built_a_free_open_source_tool_that_shuts_your/   submitted by   /u/pmttyji [link]   [comments] 9 r/LocalLLaMA community 6d ago Anyone running DeepSeek-V4-Flash-0731 on MI325X with vLLM? Mine is behaving completely broken Is anyone here successfully running DeepSeek-V4-Flash-0731 locally with vLLM , especially on AMD MI325X? My setup: GPU: 1x AMD Instinct MI325X Model: deepseek-ai/DeepSeek-V4-Flash-0731 vLLM: 0.26.0 ROCm image --tokenizer-mode deepseek_v4 --reasoning-parser deepseek_v4… 20 llama.cpp releases dev-tools 6d ago b10301 cuda: fix warnings for unused variable/function ( #26688 ) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU)… 36 r/LocalLLaMA community 7d ago DS4 Flash incoming price increase "we've been able to reproduce their current prices even on rented GPUs" https://preview.redd.it/kvfk26z2uwhh1.png?width=598&format=png&auto=webp&s=356a8793a6c31bc563d552aaa5a73112ced7372e https://preview.redd.it/xthbu87auwhh1.png?width=598&format=png&auto=webp&s=08f686fee339905a33609a0346f13163aedc2671 Hello, I've seen these tweets from dax… 16 Hugging Face Daily Papers research 7d ago Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains Abstract Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation… 5 arXiv — Machine Learning research 7d ago PPDL: LLM-Based Flows as Probabilistic Programs arXiv:2608.05234v1 Announce Type: new Abstract: Building reliable applications that leverage large language models (LLMs) remains a significant challenge. While LLMs offer impressive capabilities across diverse tasks, their outputs often lack accuracy and provide no clear… 10 arXiv — Machine Learning research 7d ago Why the Third Axis Is Freedom arXiv:2608.05423v1 Announce Type: new Abstract: In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a model that produces one common answer can outperform a model retaining a broader repertoire.… 9 arXiv — Machine Learning research 7d ago Hybrid Probabilistic Zonotopes for Identifiable and Refinable Predictive Uncertainty arXiv:2608.05454v1 Announce Type: new Abstract: Probabilistic prediction heads in neural networks typically output either a Gaussian mixture or a single conformal region. Neither separates the distinct sources of uncertainty often present in real prediction tasks: a discrete… 34 arXiv — Machine Learning research 7d ago Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers arXiv:2608.05472v1 Announce Type: new Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set. We study the… 34 arXiv — Machine Learning research 7d ago Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation arXiv:2608.05819v1 Announce Type: new Abstract: Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network contraction can reduce this cost by exploiting circuit structure, although its… 34 r/LocalLLaMA community 7d ago Dual 3090 setup: 400 pp t/s to 1600 pp t/s on Qwen 3.6 27B... with slightly lower tps. First of all, my setup: Ryzen 9 5950x DDR4 3200Mhz 64gb (2x32) Dual 3090s, no NVLINK Runtime: llama.cpp Nvidia Drivers 610 Windows 11 25H2 Qwen 3.6 27B Q8 I've been using llama-server with --split-mode tensor for a couple months now, since it gave a pretty nice 10%-20% boost in… 35 r/LocalLLaMA community 7d ago 🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp 🐦⬛ Magpie-TTS Multilingual 🦜 Nemotron Speech Streaming EN 0.6B 🦜 Nemotron-3.5 ASR Streaming 🦜 Parakeet CTC 1.1B 🦜 Parakeet TDT 0.6B v3 🥦 NanoCodec Merged PR https://huggingface.co/nvidia/magpie_tts_multilingual_357m#run-magpietts-locally-with-nemo-speechcpp I run open… 9 r/LocalLLaMA community 7d ago I thought Deepseek was the answer since I cannot afford GPU for local LLM   submitted by   /u/HsSekhon [link]   [comments] 18 Ars Technica — AI news-outlet 7d ago Anthropic will design its own hardware to power Claude Anthropic and OpenAI are racing to scale up while reducing dependence on Nvidia. 35 r/LocalLLaMA community 7d ago nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlx nvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it plainly, the vision and audio towers need a runtime that implements the… 20 r/LocalLLaMA community 7d ago I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLM I'm the author, so discount the enthusiasm accordingly. This is an unaffiliated community port, not endorsed by the vLLM project, which it uses to verify its correctness. What started it: I love vLLM, but a vLLM install here is 9.1 GiB of virtualenv, and I wanted to embed… 36 r/LocalLLaMA community 7d ago nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blue (RGB) document image and a task prompt, the model produces formatted text and… 38 r/LocalLLaMA community 7d ago Best llama cpp flags to run Deepseek-flash 0731 Hi all. These are my system specs: dual xeon e5 2696 v2 , 160gb DDR3 ram ECC(1600mhz), 3 gpus: 3060 12gb, p100 16gb, 3050 6gb. And a 400gb nvme sdd RAID0, 3000 mb/s. The model is Deepseek-flash-0731 UD_8_X_XL, loseless, 161gb. Now, I'm not too knowledgeable about llama cpp… 12 r/LocalLLaMA community 8d ago I get that AI labs need to make money, but zero-warning price spikes are a nightmare for production builds Seen a ton of posts today about the DeepSeek API price hike. Half the feed is doom-posting, the other half is explaining basic GPU economics. Honestly, I get the cost side. Sub-cent tokens were never gonna last forever. But what actually sucks is the zero-day notice. Dropping a… 23 arXiv — NLP / Computation & Language research 8d ago Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap arXiv:2608.04160v1 Announce Type: new Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. We test whether the… 34 arXiv — NLP / Computation & Language research 8d ago Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs arXiv:2608.04488v1 Announce Type: new Abstract: Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high VRAM, time, and energy costs. Parameter-Efficient Fine-Tuning (PEFT) of Small… 8 arXiv — NLP / Computation & Language research 8d ago Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv:2608.04899v1 Announce Type: new Abstract: Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extremely sparse outputs. For instance, Qwen3-32B… 27 arXiv — NLP / Computation & Language research 8d ago FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables arXiv:2608.04077v1 Announce Type: cross Abstract: Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner… 29 arXiv — NLP / Computation & Language research 8d ago Simile Understanding in Text-to-Image Models: An Evaluation Framework arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models… 12 r/LocalLLaMA community 8d ago Could we have a --disk-moe or --n-disk-moe like --cpu-moe or --n-cpu-moe so we can use disk/cpu/gpu ? Explicit title, It would be nice to have the ability to have 3 tiers moe offload :(   submitted by   /u/storm1er [link]   [comments] 38 r/LocalLLaMA community 8d ago 40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s) daily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster I would wager that compared to a naive kernel anyone can write it's more in the range of 10-20% faster e2e in reality, if at all, but hey, it's free and open!… 8 Hugging Face Daily Papers research 9d ago PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs Abstract Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly.… 35 arXiv — NLP / Computation & Language research 9d ago Sphere Retraction Normalizations arXiv:2608.02668v1 Announce Type: cross Abstract: Residual connections are the de facto mechanism for training deep neural networks stably. Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state… 35 arXiv — Machine Learning research 9d ago Output-Aware Rotation for INT2 KV-Cache Quantization arXiv:2608.02691v1 Announce Type: new Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods… 23 arXiv — Machine Learning research 9d ago Inverted Detection and Control in Steering Vectors arXiv:2608.02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption underpinning SVs is that they are linearly discriminative with respect to the… 20 arXiv — Machine Learning research 9d ago Provably Learning Multi-Head Attention with Queries arXiv:2608.03294v1 Announce Type: new Abstract: We study the problem of learning multi-head softmax attention from black-box input-output access. The learner may query arbitrary real-valued token sequences and observe only the scalar output at the final token. Recent work gives… 6 arXiv — NLP / Computation & Language research 9d ago Stuck on "A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model arXiv:2608.02689v1 Announce Type: new Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budget, and ask a simple question: what exactly does the conversion break? After… 38 arXiv — NLP / Computation & Language research 9d ago ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection… 21 arXiv — NLP / Computation & Language research 9d ago VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations arXiv:2608.03675v1 Announce Type: new Abstract: Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in high-stakes domains such as human and veterinary medicine. However, this does… 29 arXiv — NLP / Computation & Language research 9d ago VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs arXiv:2608.03810v1 Announce Type: new Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a… 11 r/LocalLLaMA community 9d ago PSA Update CUDA from 13.2 to 13.3 to solve DeepSeek V4 Flash 0731 Looping Problem! So one of yall mentioned that cuda 13.1 or 13.2 is broken for unsloth so I looked in to it, and they were right. I had 13.2 installed, after I switched to 13.3 no more looping!!! Before the cuda update, the model was literally unusable. A few minutes into the run it would start… 20 r/LocalLLaMA community 9d ago DeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark https://github.com/yhfgyyf/vllm-deepseek-v4-sm89 I couldn't believe that someone actually got vLLM working with this particular set of GPUs, but here it is. The video is from right after I got it working with 64k context, but it is now running with 256k.   submitted by  … 26 r/MachineLearning community 9d ago I Compressed Bad Apple into a 3MB Neural Network [P] I trained a small MLP to memorize the classic Bad Apple animation, ~2.7 billion pixels of video compressed into 790k parameters (3.2 MB float32, 1.6 MB float16). The network takes a 3D coordinate (t, y, x)- frame index and pixel position- and outputs a grayscale value between 0… 12 llama.cpp releases dev-tools 9d ago b10275 server: decode Windows OEM output to UTF-8 in built-in tools ( #26597 ) a child process writes in the OEM code page, which is not UTF-8 on a western Windows install, so accented output reaches the JSON layer as invalid bytes and gets replaced there, silently losing the… 38 TechCrunch — AI news-outlet 9d ago Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending against AI agents. 5 r/LocalLLaMA community 9d ago Hugging Face CEO says China is winning the AI race and dominating on open models This is something that was spoken here and there, and now it is like writing on the wall. The main additional point is that China has created an independent supply chain. Starting from raw materials and home-made lithography equipment, through their own GPU manufacturing, and to… 28 r/LocalLLaMA community 9d ago A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM A new llama.cpp PR (#26563) adds a heatmap that tracks which MoE experts are used most often. Instead of keeping every expert on the GPU or offloading all of them, it caches the frequently selected experts in VRAM while the cold experts continue running on the CPU. The author’s… 29 r/LocalLLaMA community 9d ago Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible. I run my inference machine in the living room, so noise and heat output are a significant concern. Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% less t/s in decode and 8.8% in prefill (which is already very fast). Well worth… 6 NVIDIA Developer Blog official-blog 9d ago Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super Autonomous vehicle (AV) development often relies on separate models for trajectory generation, high-level intent prediction, scene understanding, and data... 10 r/LocalLLaMA community 9d ago Llama.cpp PR 8% speed boost Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and observed a 4% increase inference speed boost. Pretty exciting to see 84 tok/s max on… 19 r/LocalLLaMA community 9d ago Optimised DSv4-Flash for 2x GH200: 10,000 tok/s PP, >300 tok/s TG on SGLang There are some PRs to use and a nice trick to speed up PP on really longs contexts in my write up. Hope it helps! TL;DR: On this dual GH200 box, you build vLLM v0.26.0 from source, add the merged DSV4 cache-layout patch (PR #48993), disable async scheduling, and run DSpark at 6… 25 Hacker News — AI on Front Page community 9d ago DeepSeek V4 Flash on a Single AMD MI300X Article URL: https://github.com/ryanzhou/deepseek-v4-flash-mi300x Comments URL: https://news.ycombinator.com/item?id=49166386 Points: 290 # Comments: 64 32 Page 3 of 10 · 500 articles ← Newer Older →