News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow arXiv — NLP / Computation & Language research 10d ago Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds arXiv:2609.19149v1 Announce Type: new Abstract: Subliminal learning shows that language models can transmit a hidden trait through outputs that appear unrelated to it. One proposed explanation, token entanglement, links animal and number tokens through the model's output… 16 r/LocalLLaMA community 10d ago Installing 6 GPUs in a standard case rather than using an open-frame chassis. CPU : epyc 7262 MB : ROMED8-2T GPU : V100 16GB PCIE x6 I have built a server with six V100 GPUs. I am now testing it and plan to eventually run Qwen3.8-Next-Flash configured with TP2 and PP3. Because open-frame or server-style cases are large and unattractive, I chose to install… 8 r/LocalLLaMA community 10d ago Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU. The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence. -… 7 Hacker News — AI on Front Page community 10d ago Bend – A language that blocks AI mistakes via proof, on CPU and GPU Article URL: https://bend-lang.com/ Comments URL: https://news.ycombinator.com/item?id=49746163 Points: 250 # Comments: 132 19 r/LocalLLaMA community 10d ago AMD Plans 10% Price Hike Across GPUs, Chipsets, and Possibly CPUs Great news! AMD is also considering accepting payment in organs! Slightly less sarcastically, grab what you can, while you can. Waiting is becoming very costly almost by the day   submitted by   /u/FullstackSensei [link]   [comments] 27 TechCrunch — AI news-outlet 10d ago Huawei plans Q1 2027 launch of new AI chip as it takes on Nvidia Huawei is accelerating the launch of its next-generation Ascend 960DT AI chip as it pushes to compete with Nvidia and close China’s AI computing gap with the U.S. 38 TechCrunch — AI news-outlet 10d ago Google, Nvidia and Anthropic want Emerald AI to find space on the grid for more data centers A new coalition that includes Google, Nvidia, Anthropic and Emerald AI wants to find 100 GW of grid capacity for new data centers. 22 r/LocalLLaMA community 10d ago We built an open-source GPU profiler you point an AI agent at, instead of reading traces yourself We've been tuning vLLM/SGLang/llama.cpp setups for a long time and got tired of the profiling part: nsys trace, open the GUI, squint, change a flag, repeat. The profilers assume a human is looking at the timeline. These days the thing doing our tuning is usually an agent, and it… 26 r/LocalLLaMA community 11d ago China's Huawei says AI chip demand outstrips supply as it steps up Nvidia challenge   submitted by   /u/sunychoudhary [link]   [comments] 11 r/LocalLLaMA community 11d ago Intel releases OpenVINO 2026.4 UPDATE https://github.com/ggml-org/llama.cpp/pull/29009 MERGED More Gen AI coverage and frameworks integrations to minimize code changes New models supported: On CPU: Gemma-3n On CPU, GPU: Kokoro-82M, Qwen3-VL-4B with eagle3, Qwen3-ASR, Muse Glimmer 30B, Qwen 3.8 27B, Gemma4… 22 The Information — AI news-outlet 11d ago Huawei Speeds Up AI Chip Launch to Challenge Nvidia Huawei Technologies plans to launch its new AI chip in the first quarter of 2027, nine months earlier than planned, as it steps up its efforts to challenge Nvidia. Tao Wang, Huawei’s deputy chairman and rotating chairman, said at the Huawei Connect conference in Shanghai on… 15 Hugging Face Daily Papers research 11d ago Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Abstract Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators,… 16 arXiv — Machine Learning research 11d ago Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs arXiv:2609.17564v1 Announce Type: new Abstract: Deploying retrieval-augmented generation (RAG) on commodity GPUs such as the NVIDIA T4 (16 GB VRAM) exposes a practical failure mode we call the Compression Paradox: neural prompt compression can add key-value (KV) cache contention… 7 arXiv — Machine Learning research 11d ago Bias Amplification in Multi-Agent Network: How Biased Agents Shape Opinions and Rhetoric arXiv:2609.18306v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in applications involving interaction between agents, where their output plays a role in collective reasoning and decision-making processes. Despite significant research into… 16 arXiv — Machine Learning research 11d ago The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations arXiv:2609.18560v1 Announce Type: new Abstract: Some aspects of AI development resemble a population process in which models are specialised, retrained on the output of peers, or combined by averaging weights. These practices lead to generations of models, in the biological… 18 arXiv — Machine Learning research 11d ago Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence arXiv:2609.18612v1 Announce Type: new Abstract: We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its input (reading) and output… 7 arXiv — Machine Learning research 11d ago WaveTLM: Reliable Time-Series Language Modeling through Task Compilation arXiv:2609.18812v1 Announce Type: new Abstract: Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while hallucinating the required object:… 37 arXiv — Machine Learning research 11d ago TwinMark: A Unified Watermark for Provable Survival Under Feature and Logit Distillation arXiv:2609.19011v1 Announce Type: new Abstract: We propose TwinMark, a watermarking scheme that reads a single SHAKE128 secret through two complementary linear functionals of model-output summaries: a covariance projector against the carrier-set covariance (cov-Feat) and a… 5 arXiv — NLP / Computation & Language research 11d ago Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks arXiv:2609.17552v1 Announce Type: new Abstract: Moral-reward RL can make language-model agents more cooperative, but whether that alignment survives adversarial persona pressure is unknown. Such attacks are realistic: retrieved context, tool outputs, or multi-turn framing can… 10 arXiv — NLP / Computation & Language research 11d ago Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents arXiv:2609.17708v1 Announce Type: new Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry.… 35 arXiv — NLP / Computation & Language research 11d ago A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality arXiv:2609.18005v1 Announce Type: new Abstract: Large language model optimization is an active research area, spanning quantization of model weights, early-exit methods for skipping layers, and speculative decoding. Each track uses its own quality measures, typically an… 35 arXiv — NLP / Computation & Language research 11d ago Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers arXiv:2609.18204v1 Announce Type: new Abstract: Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces… 31 arXiv — NLP / Computation & Language research 11d ago SEA-LION-v4.8: A Technical Report arXiv:2609.18310v1 Announce Type: new Abstract: We introduce Nemotron-SEA-LION-v4.8, a family of Southeast Asian Languages in One Network (SEA-LION) built upon NVIDIA Nemotron 3. The family includes 30B-A3B and 120B-A12B models, with both continued-pretrained base checkpoints… 14 arXiv — NLP / Computation & Language research 11d ago Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models arXiv:2609.18320v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are… 35 arXiv — NLP / Computation & Language research 11d ago TalkMatrix: Generating Character Dialogue that is Both Consistent and Diverse arXiv:2609.19022v1 Announce Type: new Abstract: Candidate-based decoding typically selects a completion for each prompt independently, but many applications require a collection of outputs that satisfies global, non-decomposable requirements. We formulate this setting as… 19 arXiv — NLP / Computation & Language research 11d ago Entropy in Conversational AI: Structured Unpredictability as Inferrable Interiority arXiv:2609.19044v1 Announce Type: new Abstract: Sampling can increase response diversity without producing history-dependent behavior. We formalize a different design target, structured unpredictability, as conditional dependence between an output and a persistent hidden state… 25 r/LocalLLaMA community 11d ago Upgraded my local setup with 2 rtx pros and it's amazing. Follow up post of https://www.reddit.com/r/LocalLLaMA/s/nGMyKswrch . Thanks everyone who replied. I didn't change the specs. Might be loosing some of the memory bandwidth but will scale in future if I need to. It took me 2.5 days to build it because one of the GPU connected to… 35 r/LocalLLaMA community 11d ago I built a native Vulkan training backend for 143 modern Transformer architectures — no CUDA or PyTorch required I've been working on something that started as the training backend for my Hierarchos architecture, but it has grown into a much broader project: https://github.com/necat101/Hierarchos-Native Hierarchos Native now includes a native Rust + Vulkan backend for training and… 14 r/LocalLLaMA community 11d ago i left gpu poor range i am not GPU poor anymore https://preview.redd.it/a2tcwu9kjxph1.png?width=1173&format=png&auto=webp&s=58dfd2a26b48825d64e80cd4358584d524565bb7   submitted by   /u/paulqq [link]   [comments] 8 r/LocalLLaMA community 11d ago Enable CUDA graph for MTP draft by gaugarg-nv · Pull Request #28549 · ggml-org/llama.cpp one more MTP speedup   submitted by   /u/jacek2023 [link]   [comments] 10 llama.cpp releases dev-tools 11d ago b11007 Enable CUDA graph for MTP draft ( #28549 ) Improve CUDA graph usage for MTP Rename field Address review feedback Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47982528 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon… 13 r/LocalLLaMA community 11d ago Qwen3.8 Flash on 12GB VRAM - 15 tokens/s Achieved steady 15 tokens/second output and 100-120 promp processing per second with 12GB GPU (RTX 5070 SFF) with Qwen3.8-Flash-Next-GSQ-RCO-GGUF at 3 bpw (IQ3_XXS which maches BF16 on AIME25). Around 76GB full gguf - while only 47GB needs to be sharded (loaded into VRAM+RAM)… 31 r/LocalLLaMA community 11d ago What's the current best LLM uncensoring method? With the recent Nvidia Huggingface acquisition and frontier AI labs screaming about safety and putting guardrails everywhere, I think it's important that we have local models that aren't affected by arbitrary guardrails set during training. To be clear, this post NOT about… 28 NVIDIA Developer Blog official-blog 11d ago Translating CUDA Tile Operations from Python to Rust Using Agentic AI cuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to... 34 r/LocalLLaMA community 11d ago Qwen3.8-27B uncensored Q6_K at 156K context on one RTX 5090, 140-190 tok/s with DFlash2 Setup for one long agentic coding session (tools + vision) on a single 5090: Qwen3.8-27B RVN Heretic (ARA abliterated) at Q6_K, 159744 context , 1 slot, q8_0 K/V, DFlash2 speculative decoding, vision projector on the GPU, Sharp chat template. All numbers measured 2026-09-16.… 18 r/LocalLLaMA community 11d ago Qwen3.5 4B + grabbing logits is almost "Jev"? Or even just Qwen Reranker? Saw the new Introducing System One Models & Jev - TypeSafe AI Blog Jev model which just outputs probabilities given choices and I thought it sounded a something the Qwen reranker models could do? Then I tried implementing it and got even better results with Qwen 4B + just assign… 7 TechCrunch — AI news-outlet 11d ago Robots are waiting for a ChatGPT moment: Nvidia’s Les Karpas explains why at TechCrunch Disrupt 2026 The robotics industry is still waiting for their breakthrough into day-to-day life. Nvidia's Les Karpas has an answer as to why at TechCrunch Disrupt 2026. Register before September 25 to save up to $200 on your pass. 9 llama.cpp releases dev-tools 11d ago b11002 CUDA/HIP: improve access patterns in im2col ( #28013 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47937267 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS… 10 r/LocalLLaMA community 11d ago Apple May Return to Server Market With Nvidia Technology Apple is considering offering an AI server built around "M8" series chips and has discussed incorporating Nvidia networking hardware, The Information reports. The system would be sold to outside customers, potentially bringing Apple back into a business it left behind when it… 30 r/LocalLLaMA community 11d ago LocalJev? Jev is a model to produce structured output (choices) from input text. It apparently can play (not run!) Doom. https://typesafe.ai/blog/introducing-system-one-models-and-jev Is there already a open implementation of this kind of model?   submitted by   /u/SomewhereAtWork… 7 The Information — AI news-outlet 11d ago Apple Considers Return to Server Market, Has Talked With Nvidia to Use Network Tech Apple has been working on plans for an enterprise server using its own chips that could incorporate networking equipment from Nvidia , as the Mac maker looks to capitalize on surging interest in its computers for AI. Apple has been developing the server with plans to sell it to… 19 The Information — AI news-outlet 11d ago Nvidia CEO to Attend Trump’s State Dinner for Xi Nvidia CEO Jensen Huang plans to attend President Donald Trump’s state dinner for Chinese leader Xi Jinping in Washington next week, according to a person familiar with the matter. His attendance comes as AI chips and the pace of AI development become growing sources of tension… 26 llama.cpp releases dev-tools 12d ago b10993 ci : add self-hosted webgpu to hf-jobs ( #28712 ) add self-hosted vulkan and webgpu to hf-jobs try t4-medium cont : adjust cpu backend threads try t4-small again restore cm jobs Co-authored-by: Georgi Gerganov [email protected] Website: https://llama.app Attestations:… 29 r/LocalLLaMA community 12d ago Is an X399 rig still viable? I saw a couple of these motherboards with 4x pcie 3.0 x16 slots paired with a 2nd gen threadripper 2920x/2950x for like $350 and I was wondering if they are still worth using for multi gpu builds?   submitted by   /u/No-Orchid-6159 [link]   [comments] 23 arXiv — NLP / Computation & Language research 12d ago Z-Loss Backward Geometry in Dense Output Heads and Sparse Routers arXiv:2609.16179v1 Announce Type: cross Abstract: Z-loss has been widely applied to the logits of language-model output heads and sparse mixture-of-experts routers. Z-loss constrains the softmax log-normalizers of these output heads and routers, thereby limiting large-logit… 34 arXiv — Machine Learning research 12d ago A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure arXiv:2609.16606v1 Announce Type: new Abstract: Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a… 12 arXiv — Machine Learning research 12d ago LCAP: Population-Informed Latent Chip Adaptation from Few Output Probes for Photonic Neural Networks arXiv:2609.16823v1 Announce Type: new Abstract: Photonic neural networks (PNNs) offer efficient analog inference, but parameters optimized under ideal device models can degrade after fabrication, creating a persistent simulation-to-hardware (sim-to-real) gap. When many… 8 arXiv — Machine Learning research 12d ago High-Fidelity Digital Twin Data Models by Randomized Dynamic Mode Decomposition and Deep Learning with Applications in Fluid Dynamics arXiv:2609.17101v1 Announce Type: new Abstract: The purpose of this paper is the identification of high-fidelity digital twin data models from numerical code outputs by non-intrusive techniques (i.e., not requiring Galerkin projection of the governing equations onto the reduced… 33 arXiv — Machine Learning research 12d ago Nonsmooth Optimization via Orthogonalized Momentum arXiv:2609.13677v1 Announce Type: cross Abstract: Modern real application problems involve matrix-valued parameters, yet conventional optimizers treat them as vectors, thereby motivating matrix-aware methods that exploit input-output geometry, such as Muon which orthogonalizes… 4 arXiv — NLP / Computation & Language research 12d ago Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks arXiv:2609.16023v1 Announce Type: new Abstract: Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively, however, is expensive, and rubric-based evaluation has become the… 25 Page 4 of 10 · 500 articles ← Newer Older →