News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow arXiv — NLP / Computation & Language research 12d ago StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation arXiv:2609.16340v1 Announce Type: new Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which the newer model may already surpass. Moreover, collecting fresh… 9 arXiv — NLP / Computation & Language research 12d ago Challenges of Auditing: Variability in Outputs of Large Language Models for Health arXiv:2609.16590v1 Announce Type: new Abstract: People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations… 29 arXiv — NLP / Computation & Language research 12d ago An Empirical Study of Counterfactual Self-Explanations in LLMs arXiv:2609.17119v1 Announce Type: new Abstract: Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a… 25 arXiv — NLP / Computation & Language research 12d ago Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs arXiv:2609.17327v1 Announce Type: new Abstract: This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages. We employ several fine-tuned vision-language… 25 arXiv — NLP / Computation & Language research 12d ago Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation arXiv:2609.16145v1 Announce Type: cross Abstract: We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable… 17 r/LocalLLaMA community 12d ago LACT PR to let NVIDIA gpus go lower than stock VBIOS limit (so below 400W for 5090, or below 250W for 6000 PRO MaxQ) Hallo guys, hoping you're doing fine. This PR let for users that use LACT, go lower than the stock VBIOS limit on their cards, and such being more power efficient. This let you go down to 30W, or lower if the VBIOS on your card goes lower than that. On a RTX 6000 PRO the min it… 29 TechCrunch — AI news-outlet 12d ago We don’t need AI regulation — leave safety to us, Nvidia’s Jensen Huang says AI isn't some kind of new form of "alien mind," according to Jensen Huang. It's just hardware and software, so safety can be engineered by each AI product maker. 21 Hacker News — AI on Front Page community 12d ago Building a Linux GPU Driver for the M4 Mac Mini in One Month Article URL: https://codyho.dev/blog/gpu-driver/ Comments URL: https://news.ycombinator.com/item?id=49717638 Points: 211 # Comments: 131 30 NVIDIA Developer Blog official-blog 12d ago How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin Power is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each component of that platform must maximize... 32 NVIDIA Developer Blog official-blog 12d ago How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories For operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale AI training, every GPU in the cluster... 37 NVIDIA Developer Blog official-blog 12d ago Scaling Federated Learning Across Docker, Kubernetes, and Slurm with NVIDIA FLARE Federated learning (FL) projects often begin with a straightforward setup: one server, a few clients, and one dataset at each site. As those projects grow, the... 36 r/LocalLLaMA community 12d ago ByteShape Qwen 3.8 27B: To KL Diverge or Not to KL Diverge, Part 2: Metric Boogaloo Hey r/LocalLLaMA , We’ve released our full ShapeLearn GGUFs for Qwen 3.8 27B. Blog / Download models TL;DR 3.84 bpw (GPU-5) reaches 99.63% of BF16’s aggregate score of 8 benchmarks, being the most accurate quant we’ve evaluated; 3.23 bpw (GPU-4) reaches 98.72%. These average… 6 llama.cpp releases dev-tools 12d ago b10984 cuda: support row-contiguous SUM_ROWS ( #26308 ) cuda: support row-contiguous SUM_ROWS organize the code and add GGML_OP_MEAN to support row-contiguous tensors using the same shared kernel, and add a test to MEAN permute/slice Keep original comments and add if/else branch… 4 llama.cpp releases dev-tools 13d ago b10981 OpenVINO: optimize stateful decode and GPU MoE inference ( #28638 ) exclude GPU/NPU failing POOL_2D case Fix pool case ggml-openvino: fix stateful decode for Gemma-4 per-layer-type head sizes ggml-openvino: fix MSVC narrowing error in permute ggml-openvino: classify… 20 TechCrunch — AI news-outlet 13d ago Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear Salesforce Koa is built on Nvidia's open-weight Nemotron model and is trained to do sales, marketing and customer-support tasks. 14 llama.cpp releases dev-tools 13d ago b10977 ci: Bump CUDA Windows x64 builds to 13.4.1 ( #28930 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47583571 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS… 23 r/LocalLLaMA community 13d ago Dual AMD Radeon AI Pro R9700 or dual NVIDIA or RTX 3090. I’m building a dual GPU box. Originally the goal was 30b at FP8 or 70b at Q4. I planned to run two R9700’s, but I found a pair of NIB RTX 3090’s near me for $1600 each. Discount if I buy both. RTX 40/50 series are off the table. I would love an RTX PRO 6000, but that is also off… 16 arXiv — Machine Learning research 13d ago Generative Interpretability via Scalable Neuro-Symbolic Models arXiv:2609.13529v1 Announce Type: new Abstract: As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability,… 22 arXiv — Machine Learning research 13d ago AttnFuse: A Composable DSL for Compiling Attentions to Fused GPU Kernels arXiv:2609.13612v1 Announce Type: new Abstract: Modern AI systems are built on the Transformer architecture, whose core operation, attention, accounts for the majority of computation and memory cost. Researchers continually propose new attention variants to improve quality,… 7 arXiv — Machine Learning research 13d ago An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS arXiv:2609.13624v1 Announce Type: new Abstract: Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are… 20 arXiv — Machine Learning research 13d ago Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging arXiv:2609.13773v1 Announce Type: new Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human… 38 arXiv — Machine Learning research 13d ago Affinity-Aware Sharding for Delayed Tensor Parallelism arXiv:2609.13846v1 Announce Type: new Abstract: Delayed Tensor Parallelism (DTP) removes the blocking all-reduce of tensor-parallel Transformer inference. Every device adds its own partial output to its residual stream (and broadcasts it) immediately, but only gathers (receives)… 25 arXiv — Machine Learning research 13d ago Communication-Efficient LLM Adaptation over Decentralized GPU Meshes arXiv:2609.14339v1 Announce Type: new Abstract: Decentralized training enables large-model training over low-end GPUs and internet-grade connections, but communication along both data-parallel and pipeline-parallel axes becomes the primary bottleneck. We study post-pretraining… 11 arXiv — NLP / Computation & Language research 13d ago In the Blind: Building Pseudo-References for MT Evaluation arXiv:2609.13611v1 Announce Type: new Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these… 19 arXiv — NLP / Computation & Language research 13d ago ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation arXiv:2609.13737v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk… 13 arXiv — NLP / Computation & Language research 13d ago Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth arXiv:2609.13745v1 Announce Type: new Abstract: Vision--language models (VLMs) can answer chart questions accurately, but output accuracy does not show how they combine the evidence needed to recover an exact value. We study vertical-bar value reading with controlled… 32 arXiv — NLP / Computation & Language research 13d ago Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning arXiv:2609.14896v1 Announce Type: new Abstract: A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring… 27 arXiv — NLP / Computation & Language research 13d ago Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation arXiv:2609.14963v1 Announce Type: new Abstract: As large language models become capable translators of classical texts, a key challenge is deciding which outputs need expert review when no human reference exists. This study tests reference-free error triage through… 21 Hugging Face Daily Papers research 13d ago Kaininja: Extending Native 3D Generators to the Part Level Abstract KaiNinja extends a native 3D generator to part-level outputs using a dual-volume representation that resolves interface conflicts, improving both part and whole-object fidelity without segmentation. Generated by thinkingmachines/Inkling-Small Native 3D generators turn… 4 llama.cpp releases dev-tools 13d ago b10975: cuda : enable i16 and i32 for DUP (#28897) cuda : enable i16 and i32 for DUP docs : update ops table for DUP on CUDA 5 TechCrunch — AI news-outlet 13d ago Nvidia CEO Jensen Huang tells Trump ‘we’re not going to let [an AI slowdown] happen’ Though Elon Musk and Sam Altman have supported Dario Amodei's calls to slow the pace of AI development, Jensen Huang seems to feel differently. 35 r/LocalLLaMA community 13d ago Nvidia's RTX 5090 vanishes from online retail in the US — third-party sellers now demand as much as $9,500 for Nvidia's fastest GPU   submitted by   /u/Norwood_Reaper_ [link]   [comments] 17 r/LocalLLaMA community 13d ago NVIDIA Unveils RTX PRO 5500 "Blackwell" Workstation GPU with 84 GB GDDR7 Memory   submitted by   /u/Lumpy_Phase_9539 [link]   [comments] 32 r/MachineLearning community 13d ago How to automatically find the batch size when using Accelerate with FSDP2? [D] Hi, For single-GPU training, I’m using Hugging Face SFTTrainer with auto_find_batch_size=True, which automatically reduces the batch size after a CUDA OOM until it finds a batch size that works. I would like to have similar behavior when training on multiple GPUs on a single… 17 llama.cpp releases dev-tools 13d ago b10969: ci : add ubuntu-cuda builds to release (#28186) release : add ubuntu-cuda build job (12.8/13.3, x64+arm64) Add GCC 14 for CUDA arm64 builds in CI Eplicit bash Install git for CCCL fetch Install git before we clone/checkout Match CI names for WIndows Whitelist llama.cpp repo to git Use $GITHUB_WORKSPACE Also ship dependent… 20 r/LocalLLaMA community 13d ago Coding Agent running entirely in the browser with Pi + MiniCPM5-2B (webGPU)   submitted by   /u/paf1138 [link]   [comments] 28 NVIDIA Developer Blog official-blog 13d ago Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE... 29 r/LocalLLaMA community 13d ago For the GPU poor. K2 Horizon 7B ranks between qwen 3.6 27B and qwen 3.6 35BA3b on the Artificial Analysis Intelligence Index. From initial testing it seems pretty solid so far. Asked it to compile the latest llama.cpp for CUDA and its doing well so far. If this thing holds up to its score then its SHOCKINGLY good for its size. https://huggingface.co/IFM/K2-Horizon-7B-GGUF   submitted by  … 19 llama.cpp releases dev-tools 13d ago b10956 sycl: rfc: Use radix select for top_k ( #28670 ) sycl: GPU-resident TOP_K for large k, parallelised over the device The SYCL backend refused GGML_OP_TOP_K above k = 32 and let it fall back to the CPU, a backend round-trip per call. The limit was not conservatism: the scan-merge… 30 The Information — AI news-outlet 13d ago Anthropic Data Fears Prompt Nvidia, Palantir and Booz Allen to Restrict Model Use As paranoia rises over whether Anthropic or OpenAI could learn from their customers’ intellectual property, large firms involved in sensitive corporate work, such as Palantir Technologies, Nvidia and Booz Allen Hamilton, have started demanding new guarantees or reducing or… 6 r/LocalLLaMA community 14d ago What are the current best retail GPUs for max VRAM at a reasonable price? I am considering dumping my ChatGPT Plus subscription and go full local, but to do so I would first need to reach a decent result for quality (and reasonable speed). My 4090 fried itself out of nowhere, so I am not stuck with a 3070 until I get something better. I am kind of… 16 r/LocalLLaMA community 14d ago Running Qwen 3.8 next on 16vram+32ram - A useful/fun post for the gpu poors Hello Reddit. Posting this for fun. I thought it was a lonely and silly journey to set up Qwen 3.8 Next on a system that doesn't really run it properly—it was a challenge that might help the community. I have yet to benchmark this specific REAP version versus Qwen 3.8 27B QK4,… 7 arXiv — Machine Learning research 14d ago 3D Digital Twin Visualization of Multiclass GRF-Based Gait Disorder Classification arXiv:2609.12442v1 Announce Type: new Abstract: Automated gait analysis requires accurate classification and interpretable outputs. We propose an integrated framework for classifying healthy gait and multiple musculoskeletal impairment groups using bilateral ground reaction… 34 arXiv — NLP / Computation & Language research 14d ago SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data arXiv:2609.12353v1 Announce Type: new Abstract: Large language models trained recursively on their own or other models' outputs undergo model collapse, in which distributional tails and factual accuracy deteriorate while fluency survives. Prior work diagnoses collapse after… 24 arXiv — NLP / Computation & Language research 14d ago AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization arXiv:2609.12471v1 Announce Type: new Abstract: We introduce AMDKernelVault, an open HIP and Triton kernel corpus and training framework for recent AMD CDNA GPUs. Existing LLM-based kernel agents are largely CUDA/NVIDIA-centric and often depend on repeated frontier-LLM calls for… 20 arXiv — NLP / Computation & Language research 14d ago The House with a Million Windows: Interactive Fiction for Narrative Restorying arXiv:2609.12537v1 Announce Type: new Abstract: AI-assisted writing can flatten meaning in human storytelling, enabling the production of homogeneous outputs without the intentional effort and sense-making writing entails. To address this challenge, we present The House with a… 20 arXiv — NLP / Computation & Language research 14d ago Confidence-Gated Transductive Test Generation for Code Reranking arXiv:2609.12489v1 Announce Type: cross Abstract: Test case synthesis is crucial for evaluating and ranking programs generated by large language models (LLMs). However, constructing high-quality test cases remains challenging because reliable expected outputs are often difficult… 6 arXiv — NLP / Computation & Language research 14d ago UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking arXiv:2505.15063v3 Announce Type: replace Abstract: The rapid adoption of Large Language Models (LLMs) has raised important concerns about the factual reliability of their outputs, particularly in low-resource languages such as Urdu. Existing automated fact-checking systems are… 21 arXiv — NLP / Computation & Language research 14d ago Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation arXiv:2510.18019v3 Announce Type: replace Abstract: Multilingual watermarking aims to make large language model (LLM) outputs traceable across languages, yet current methods still fall short. Despite claims of cross-lingual robustness, they are evaluated only on high-resource… 29 arXiv — NLP / Computation & Language research 14d ago Decomposing LLM-Judge Uncertainty to Target Expert Labels arXiv:2609.06444v3 Announce Type: replace Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce,… 27 Page 5 of 10 · 500 articles ← Newer Older →