News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow r/LocalLLaMA community 3d ago MiniMax M3.1 (Space Bunny Alpha) thinks in caveman mode The CoT of MinimMax M3.1, currently available in openrouter and opencode under the guise of "Space Bunny Alpha", has the familiar look of caveman mode in order to save tokens. This has no impact on the final output. (note: in the first screenshot, pi-caveman is set to off; in… 30 llama.cpp releases dev-tools 4d ago b11158 vulkan: tune KHR cooperative matrix support for Adreno GPUs ( #29328 ) Enable coopmat support for Vulkan backend Fixed the mul_mat_s Removed the debug statement Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49831613 macOS/iOS: macOS… 35 llama.cpp releases dev-tools 4d ago b11157 cuda : add conv3d with implicit GEMM ( #29137 ) cuda : add conv3d with implicit GEMM cuda : refine conv3d implicit GEMM and handle empty kernels Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49789544 macOS/iOS: macOS Apple Silicon… 21 r/LocalLLaMA community 4d ago Folks, have you purchased the Mac M5 Ultra with 256GB yet? We need serious benchmarks, because we only get YouTube clowns influencers results Ok, so the Mac M5 Ultra (256GB) hit the market, but the only publicly available benchmarks material are flashy YouTube "clown influencers" videos. We need serious numbers to evaluate whether Apple’s silicon can actually compete with Nvidia’s current GPU‑centric workflows or not.… 11 llama.cpp releases dev-tools 4d ago b11154 test-save-load-state : print a per-model results table in --models mode ( #29316 ) test-save-load-state : print a per-model results table in --models mode in --models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead,… 35 r/LocalLLaMA community 4d ago My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context These boards cost me $115 each and I have them connected using llama.cpp with Vulkan and RPC on Bazzite. The boards have roughly 27GB of combined GPU memory and communicate over 1gb Ethernet. For around $300 including psu I’m loving the performance. I have a few more and want to… 22 arXiv — Machine Learning research 4d ago EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management arXiv:2609.27547v1 Announce Type: new Abstract: Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource… 7 arXiv — Machine Learning research 4d ago NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers arXiv:2609.27735v1 Announce Type: new Abstract: Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to… 29 arXiv — Machine Learning research 4d ago Binary Quantized Neural Network Training Is W[1]-Hard Parameterized by Input and Output Dimensions arXiv:2609.27932v1 Announce Type: new Abstract: Ganian et al. (ICLR 2026) proved that quantized neural network training is fixed-parameter tractable when parameterized jointly by architecture treewidth, input dimension $\alpha$, and output dimension $\omega$, and left open… 11 arXiv — NLP / Computation & Language research 4d ago COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference arXiv:2609.26913v1 Announce Type: new Abstract: No single Large Language Model (LLM) is uniformly reliable across queries, motivating multi-model inference systems that either route among models or combine their outputs. However, routing stops after selecting an initial model,… 7 arXiv — NLP / Computation & Language research 4d ago Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond arXiv:2609.27650v1 Announce Type: new Abstract: Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a… 25 arXiv — NLP / Computation & Language research 4d ago Predicting Quantization Price for Selecting PTQ Configurations Before Deployment arXiv:2609.28270v1 Announce Type: new Abstract: Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods… 20 arXiv — NLP / Computation & Language research 4d ago Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study arXiv:2609.26823v1 Announce Type: cross Abstract: Post-training quantization of speech language models is often summarized with text-output scores and nominal bit widths. Those numbers alone do not establish behavior that depends on information missing from a transcript, or… 38 Ollama releases dev-tools 4d ago v0.34.4-rc1: mlxrunner: Update XGrammar to 0.2.7 for structured outputs We pick up schema fixes for typed dictionary values and short arrays. 12 Ollama releases dev-tools 4d ago v0.34.4: mlxrunner: Update XGrammar to 0.2.7 for structured outputs We pick up schema fixes for typed dictionary values and short arrays. 18 NVIDIA Developer Blog official-blog 4d ago Validate GPU Cluster Readiness Before AI Workloads Land A GPU cluster can pass every health check and still fail to run an AI workload. Even when every GPU, network link, and pod reports healthy, a 512-GPU training... 28 Hugging Face official-blog 4d ago How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows Back to Articles a]:hidden"> How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows Enterprise + Article Published September 23, 2026 Upvote - Johnny Nuñez Cano johnnynv nvidia Asier Arranz asiernvidia nvidia Rishabh Chadha rchadha-nv nvidia… 9 r/LocalLLaMA community 4d ago M5U base 96GB inference numbers for Q3.8FN after 112M tokens TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted So the good news is that I got the base model on launch day with only 64 core GPU. All benchmarks are for current maxed out model,… 34 r/LocalLLaMA community 4d ago Jev isn't new tech. Its marketing targets people who think AI started with LLMs. I keep seeing Jev presented as some new class of decision model, but most of what’s being advertised is just normal classifier behavior with modern zero-shot capabilities. It outputs probabilities over constrained choices, doesn’t generate autoregressively, can’t output an… 10 llama.cpp releases dev-tools 4d ago b11140 CUDA: enable sparse-fa for dsv4 prefill (again) ( #29298 ) CUDA: enable sparse-fa for dsv4 prefill (again) CUDA: unroll the query loop of the sparse mask scan The query loop of flash_attn_mask_to_sparse_indices has a runtime trip count, which keeps the unrolled scan over the… 35 r/LocalLLaMA community 4d ago Streaming Nemotron 3 Diarization I’ve been playing with Nemotron 3 Diarization , and it fills a gap I’ve had with local voice agents: keeping track of who is speaking. It’s a diarization model, so it gives you speaker labels rather than transcriptions or people’s names. It can stream its output and track up to… 4 r/LocalLLaMA community 4d ago Perhaps the highest quality mainline quants of Qwen3.8 27B? I am proud to release these quants of Qwen 3.8 27B. They beat the excellent ISTA and Unsloth quants byte-for-byte on three corpora. Both KLD and top 1% were tested 3x. It took a week of continuous GPU and CPU time to generate these, all done on a single Strix Halo.… 4 llama.cpp releases dev-tools 4d ago b11132 model : support Gemma4 DSpark draft backbone ( #29226 ) dspark: add Gemma 4 draft support Add GGUF conversion and runtime support for full-attention and SWA Gemma 4 DSpark drafts, including tied output weights and boolean backbone metadata. Assisted-by: Codex dflash: infer Gemma… 8 Hugging Face official-blog 4d ago **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** Back to Articles a]:hidden"> Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization Enterprise + Article Published September 23, 2026 Upvote 2 Francesco fciannella nvidia Ivan Medennikov imedennikov nvidia Taejin Park taejinp nvidia Adi-… 12 llama.cpp releases dev-tools 5d ago b11124 cuda: top-k MoE should always fire ( #28432 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49495074 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework… 23 arXiv — NLP / Computation & Language research 5d ago Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages arXiv:2609.25721v1 Announce Type: cross Abstract: Language models are increasingly trained on the outputs of other models, forming chains that we call lineages, in which a trait present in one generation can pass to the next. Prior work on subliminal learning has shown that a… 11 arXiv — NLP / Computation & Language research 5d ago Training a Language Model End-to-End in Rust: An Experience Report arXiv:2609.25008v1 Announce Type: new Abstract: I pretrained a language model end-to-end in Rust - alone, with no team, no PyTorch, and no Python in the training path - for $164 in rented GPU time. I report that as an achievement, not a recommendation: the more useful… 17 arXiv — NLP / Computation & Language research 5d ago Matryoshka attribution: Learning to attribute language model outputs to representations and weights arXiv:2609.25518v1 Announce Type: new Abstract: Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle… 36 arXiv — NLP / Computation & Language research 5d ago Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference arXiv:2609.25537v1 Announce Type: new Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing… 24 arXiv — NLP / Computation & Language research 5d ago Modality-Gated Deep Adapters: Adding a Modality to a Frozen Embedding Model with Exact Preservation arXiv:2609.26182v1 Announce Type: new Abstract: Multimodal embedding models are deployed at scale: retrieval indices, benchmark results, and behavioral audits all depend on the base model's exact outputs. Extending such a model to a new modality with existing parameter-efficient… 29 arXiv — NLP / Computation & Language research 5d ago Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion arXiv:2609.26381v1 Announce Type: new Abstract: Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems. Recent vision-based parsers improve accuracy, but need GPUs and may introduce noise into the extracted text.… 34 arXiv — NLP / Computation & Language research 5d ago Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding arXiv:2609.26638v1 Announce Type: new Abstract: Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed. Unlike open-ended text generation,… 5 arXiv — NLP / Computation & Language research 5d ago Efficient Iterative Retrieval with Heterogeneous Batching arXiv:2609.25405v1 Announce Type: cross Abstract: Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these… 23 LangChain releases dev-tools 5d ago langchain-anthropic==1.7.3 Changes since langchain-anthropic==1.7.2 release(anthropic): 1.7.3 ( #40773 ) chore(model-profiles): refresh anthropic model profile data ( #40772 ) fix(anthropic): auto-route with_structured_output to method="json_schema" for fable and opus 5.5 ( #40766 ) chore(anthropic):… 38 r/LocalLLaMA community 5d ago My local llm when I tell it to do any changes to my vLLM service better make no mistakes I've been running Qwen 3.8 Flash Next and it's a great driver for Hermes and Pi. I told it to add CUDA_DISABLE_PERF_BOOST=1 to reduce my server's idle power draw   submitted by   /u/ZaltyDog [link]   [comments] 38 The Information — AI news-outlet 5d ago Wall Street’s GPU Futures Push Stalls at the CFTC The launch of futures for rental prices on Nvidia GPUs, which are seen as important to make AI compute an investable asset class, is hitting a snag. CME Group had hoped to launch contracts as soon as early October, but it won’t have regulatory approval by then. The Commodity… 15 llama.cpp releases dev-tools 5d ago b11112 server: support input_image in function_call_output ( #20663 ) ( #22575 ) server: support input_image in function_call_output ( #20663 ) server: fix if statement spacing server: avoid repeated type lookup Website: https://llama.app Attestations:… 34 r/LocalLLaMA community 5d ago MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap Some people here say MiMo-V2.6 is bad with tools and are going back to GLM-5.3-Flash. I spent today running MiMo-V2.6-Flash-RL as the backend for an agent harness, on 2× DGX Spark with vLLM, using the tonyd2wild recipe. Most of the "tool problems" I hit turned out to be serving… 8 llama.cpp releases dev-tools 5d ago b11109 metal : gate mul_mm_id src1 rescale behind ggml_prec ( #29029 ) metal : gate mul_mm_id src1 rescale behind ggml_prec Assisted-by: Claude Fable 5.1 ggml-webgpu: reject MUL_MAT_ID when src1 precision is F32 cuda/vulkan: reject MUL_MAT_ID in supports_op when src1 prec is F32 fix… 10 NVIDIA Developer Blog official-blog 5d ago Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing As large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated... 33 r/LocalLLaMA community 5d ago Best value 24/32 GB GPU? Currently sitting on some PC parts that I’m planning to sell off, and i’m considering using the money to upgrade my current gpu (12 gb vram) but if i do so i want to make sure i do my homework. Seen a lot of things thrown around this sub, like a 3090 or 9700, but i honestly… 26 NVIDIA Developer Blog official-blog 5d ago Topology-Aware Workload Scheduling with NVIDIA Topograph AI factories are power-limited systems that deliver maximum value when fully optimized. GPU workload placement is a key optimization. Poor workload placement... 5 TechCrunch — AI news-outlet 5d ago Five AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda At TechCrunch Disrupt 2026, five sessions across the AI Stage and Real World AI Stage cover AI safety, featuring leaders from Anthropic, NVIDIA, AWS, Waabi, and more. Register now to save up to $200 before Sept 25. 6 NVIDIA Developer Blog official-blog 5d ago What’s New for Game Developers: DLSS 5 with 3D-Guided Neural Rendering, NVIDIA ACE Updates, and New RTX Kit Capabilities NVIDIA DLSS 5 introduces DLSS 3D-Guided Neural Rendering and granular controls that help game developers add lifelike lighting and material detail while... 25 NVIDIA Developer Blog official-blog 5d ago Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS GPU acceleration can speed up compute-intensive robotics workloads, but a fast CUDA kernel alone does not guarantee a fast ROS 2 graph. As messages move between... 28 arXiv — Machine Learning research 6d ago CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds arXiv:2609.22183v1 Announce Type: new Abstract: Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown. We present CleanScore, a black-box audit using scored outputs only. Each benchmark question… 36 arXiv — Machine Learning research 6d ago Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles arXiv:2609.22220v1 Announce Type: new Abstract: Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these… 38 arXiv — NLP / Computation & Language research 6d ago Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems arXiv:2609.22204v1 Announce Type: new Abstract: This exploratory pilot study evaluates the scope and perceived accuracy of personal information output from ongoing conversational interactions in generative AI systems using GPT-5.2 Instant and GPT-5.2 Thinking, categorized into… 38 arXiv — NLP / Computation & Language research 6d ago Beyond Final-Token Classification: Heterogeneous Readouts for Evidence-Grounded Suicide Risk Detection arXiv:2609.22767v1 Announce Type: new Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label psychosocial factor detection, and extraction of supporting phrases. We introduce… 4 arXiv — NLP / Computation & Language research 6d ago Attributable Post-Rationalization in RAG Citations: A Controlled Reproduction and an RLVR Comparison arXiv:2609.23053v1 Announce Type: new Abstract: A RAG system can hand you the right answer and cite a source it did not actually use. Models output these unfaithful citations via post-rationalization: they write the answer first and then attach a citation to whatever passage… 24 Page 2 of 10 · 500 articles ← Newer Older →