News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow Hugging Face Daily Papers research 17d ago A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever Abstract Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is… 6 r/LocalLLaMA community 17d ago OpenAI management decided earlier today not to join the "Open Secure AI Alliance", founded by Nvidia CEO Jensen Huang. The decision was shared internally and reportedly met with backlash from employees.   submitted by   /u/KickLassChewGum [link]   [comments] 6 r/LocalLLaMA community 17d ago Viable ways to run K3 locally just curious how would people run it cheap if they really want kimi k3. dgx spark / strix halo clusters optane persistent memory platform + some gpus mac studio clusters orange pi 6 clusters ssd streaming + gpus multiple ddr3 + connectx 5 rdma clients two dgx stations power 10… 37 NVIDIA Developer Blog official-blog 17d ago NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they... 9 Marcus on AI community 17d ago Circular financing ain’t what it used to be Nvidia is falling. The mood has changed. 32 TechCrunch — AI news-outlet 17d ago Ilya Sutskever’s Safe Superintelligence partners with Nvidia to scale its AI research After two years in stealth, Safe Superintelligence has announced a long-term partnership with Nvidia as it prepares to scale to its next phase. 38 r/LocalLLaMA community 17d ago Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend & maybe results by next week. Weights are supposed to hit Hugging Face today… 18 r/LocalLLaMA community 17d ago Nvidia CEO Jensen Huang defends Open Source AI by saying distillation is fundamental to learning Nvidia CEO Jensen Huang “Distillation - learning from AI, learning from other people, and learning from other sources of knowledge, is fundamental to intelligence. We are constantly learning from one another. AI also has to learn from something.” Since using AI Desktop 98 , I… 23 llama.cpp releases dev-tools 17d ago b10152 fit : count nextn (MTP) blocks in n_gpu_layers so front layers stay on GPU ( #26177 ) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64… 8 Ars Technica — AI news-outlet 18d ago Artist sues AI meme generator for selling deeply personal comic as ad template Meme generator may have screwed up by using templates in outputs, expert says. 6 Hugging Face official-blog 18d ago NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics Back to Articles a]:hidden"> NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics Enterprise + Article Published July 27, 2026 Upvote 1 Lukas Zbinden lzbinden nvidia Javier Gamazo javirk1 nvidia Mostafa Toloui mtoloui nvidia Sean Huver shuver… 12 r/LocalLLaMA community 18d ago I want to run Kimi K3 at home, so I’m trying to make 2.8T-scale experimentation cheaper Hey r/LocalLLaMA , I’m a retired engineer with a background in distributed computing, currently running a 1-person startup. Like many people here, I’d love to experiment with 2T+ MoE models locally. The problem is that I don’t have an H100 cluster in my living room. So I’ve been… 12 Hugging Face Daily Papers research 18d ago Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Abstract Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in… 35 arXiv — Machine Learning research 18d ago RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this… 31 arXiv — Machine Learning research 18d ago DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection arXiv:2607.22035v1 Announce Type: new Abstract: Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain… 25 arXiv — Machine Learning research 18d ago Interior interpretability with attention rollout: contraction and propagation profiles in Transformers arXiv:2607.22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce… 9 arXiv — Machine Learning research 18d ago FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs arXiv:2607.21624v1 Announce Type: cross Abstract: Transformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains… 20 arXiv — NLP / Computation & Language research 18d ago Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination arXiv:2607.22067v1 Announce Type: new Abstract: The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on… 26 NVIDIA Developer Blog official-blog 18d ago NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware... 26 r/LocalLLaMA community 18d ago ~20s that you'll never get back There is no quality or value to this post, however, I hope you might find humor in this broken output from my local Qwen TTS setup. Evidently I messed something up.   submitted by   /u/Full_Dimension_3495 [link]   [comments] 34 r/MachineLearning community 18d ago I want to use AI coding agents for machine learning projects [D] I'm a software engineer who mainly builds softwaes/applications, and I'm starting to work on machine learning projects. Since ML workloads often require GPUs, I know services like Google Colab and Kaggle exist. but, I'm looking for something a bit different. Is there a platform… 21 r/LocalLLaMA community 19d ago Built a system with four P100 GPUs. https://preview.redd.it/qdhag3xmgkfh1.jpg?width=4000&format=pjpg&auto=webp&s=31f2e1cf513407a640e68eb88e19497c0cebfda1 I have built a system with four P100s, and ultimately, I plan to house six of them in a standard case.… 20 r/LocalLLaMA community 19d ago POCKET-35B agentic model on cpu 59 t/s https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF A 35B model that runs on your PC with no GPU — and on your phone. Just stock llama.cpp . No fork , no CUDA, no cloud. The POCKET lineup — pick by your device Repo File Size Runs on Best for Korean PPL* POCKET-35B-GGUF Q4_K_M 21… 8 r/LocalLLaMA community 19d ago GLM 5.2 and ik_llama.ccp Running GLM-5.2 (the new glm-dsa arch), Unsloth UD-Q4_K_XL, on a 4-socket Xeon E7-8880 v4 box with 1TB RAM and a single RTX 3060 12GB. ik_llama.cpp, experts on CPU (--cpu-moe), 24 attention layers on the GPU. Works great at 8k context — rock solid, ~3.7 tok/s gen. Problem: the… 5 r/MachineLearning community 19d ago Understanding GPU Inference Workloads [D] Hey everyone, I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here. If you've used online services like runpod or vast.ai , your perspective is extremely valuable. Please share… 20 r/MachineLearning community 19d ago Advice on how to land GPU/ML Systems interviews [D] Hi, I’m a recent MS in Data Science grad seeking entry-level GPU/ML Systems or Inference Engineering roles . (Currently on F1 OPT). Background : CUDA, Metal, and Triton work - built a FlashAttention kernel in Metal for Apple Silicon (tiled, online softmax, fp16 with simdgroup… 21 r/LocalLLaMA community 19d ago MI50 power curve tests tests done power limiting the GPU on LACT - real power usage varies wildy at 20W it ranges from 25W to 56W same behavior happens on every setting prompt for the test runs: https://github.com/lukesdevlab/youtube/blob/main/prompts/agent-maze.txt analysis by mimo 2.5 Key Findings:… 25 r/LocalLLaMA community 19d ago Mobile Offline LLMs: What do you use them for? I've spent the last year or so playing around with open source MLX and GGUF models on iPhone hardware. Given the limitations in memory, GPU/CPU/ANE, and in turn the context window I've been trying to figure out the best use cases for them. I've also done a lot of testing with… 13 r/LocalLLaMA community 19d ago Benchmarks: TensorSharp vs. llama.cpp Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like Gemma4, DiffusionGemma, Qwen3.6 with multi-modal (image, vision, audio), Qwen… 38 r/LocalLLaMA community 19d ago LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend Everything runs through WebGPU, in-browser or in electron/tauri apps. It's fully portable and supports either Nvidia and Apple Silicon (Metal). The actual kernels are optimized for the specific hardware of the device. The Nvidia kernels are aggressively fused into a multi-pass… 20 r/LocalLLaMA community 20d ago ModelExpress: Distributing Model Artifacts at the Speed of Light - NVIDIA Technical Blog We cut DeepSeek-V4 Pro startup from 8 minutes to under 2 minutes by moving weights over the fastest path to GPU memory with GPU-to-GPU RDMA . This was achieved using NVIDIA ModelExpress (MX), the weight distribution and cache management service in NVIDIA Dynamo, and this same… 9 r/LocalLLaMA community 20d ago PSA: DO NOT use Intel consumer platforms for multi-GPU setups Since a lot more people are trying to build their own multi-GPU machines, I thought I should help to prevent a common mistake people make with building multi-GPU machines. Which is using an Intel consumer platform like Z890 for multi-GPU setups. Although the CPU provides 24 PCIe… 23 Ollama releases dev-tools 20d ago v0.32.4-rc0: model: add Laguna MLX support (#17237) model: add Laguna MLX support Add Laguna XS 2, XS 2.1, and S 2.1 support to the MLX model and create paths. Read the source config to apply one quantization policy across dense and routed MoE layers. Keep the tied output head and router at source precision, quantize supported… 19 Hugging Face Daily Papers research 20d ago NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Abstract Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent… 34 r/LocalLLaMA community 20d ago It appears that the anti opensource AI lobby is far outgunned already The earlier post on this subreddit by 20+ companies signing the petition including Microsoft, Meta, Nvidia, YC ( https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/ ) etc plus this https://xcancel.com/elonmusk/status/2080672505660834163 And the entire… 32 TechCrunch — AI news-outlet 20d ago As US weighs response to Chinese AI, industry urges against broad open-weight restrictions AI companies including Nvidia and Mistral urge policymakers to avoid broad restrictions on open-weight AI models as Washington debates responses to Chinese AI and alleged model distillation. 5 r/LocalLLaMA community 20d ago More than 20 companies including NVIDIA, Meta, Microsoft, Palantir, and Hugging Face have signed a letter urging policymakers to avoid premature restrictions on open weight models. The Open Letter was initiated by Microsoft and published today: “ Open Weights and American AI Leadership ”. It argues against broad or premature restrictions on open-weight models and explicitly says policymakers should distinguish legitimate model distillation from… 5 Hacker News — AI on Front Page community 20d ago Nvidia, Microsoft, Meta warn against overregulating open-weight models Letter: https://images.nvidia.com/pdf/Open-Weights-and-American-AI-L... [pdf] https://x.com/JensenHuang/status/2080643682408321103 , https://xcancel.com/JensenHuang/status/2080643682408321103 https://www.wired.com/story/silicon-valley-is-completely-div... ,… 13 Hacker News — AI on Front Page community 21d ago Half-Life 2 running natively on HaikuOS Article URL: https://discuss.haiku-os.org/t/haiku-nvidia-porting-nvidia-driver-for-turing-gpus/16520?page=18 Comments URL: https://news.ycombinator.com/item?id=49034868 Points: 283 # Comments: 53 15 llama.cpp releases dev-tools 21d ago b10106 CUDA: fix external compilation of q1_0 MMQ ( #25778 ) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu… 4 Hugging Face Daily Papers research 21d ago SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Abstract We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while… 19 arXiv — Machine Learning research 21d ago Chronofy: A Temporal-Logical Decay Architecture for Information Validity in Time-Aware Retrieval-Augmented Generation arXiv:2607.20560v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems retrieve and integrate external knowledge to ground large language model (LLM) outputs. However, current RAG architectures treat all retrieved facts as equally valid regardless of… 7 arXiv — Machine Learning research 21d ago Detecting Neural Network Failures through Spectral Analysis of Internal Activations arXiv:2607.20590v1 Announce Type: new Abstract: Neural network misclassifications exhibit characteristic spectral instability in internal activations that is invisible at the output layer. This phenomenon is identified and formalized as Spectral Drift -- the frequency-domain… 25 arXiv — NLP / Computation & Language research 21d ago GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries arXiv:2607.20757v1 Announce Type: cross Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in-training by introducing a LogSumExp term to the loss that breaks the… 24 arXiv — Machine Learning research 21d ago Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation arXiv:2607.20908v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code generation. However, existing RLVR approaches primarily rely on outcome-based… 7 arXiv — Machine Learning research 21d ago Best-of-Evidence: Best-of-N Selection under Partial Verification arXiv:2607.20950v1 Announce Type: new Abstract: BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reliably. Many vision-language tasks instead provide only partial… 7 arXiv — Machine Learning research 21d ago Nipping the Butterfly Effect in the Bud: Self-Output Fine-Tuning for Autoregressive Weather Prediction arXiv:2607.21080v1 Announce Type: new Abstract: Long-horizon weather forecasting is a fundamental challenge in atmospheric science, for which autoregressive Deep Learning Weather Prediction (DLWP) has emerged as the primary paradigm. Although the autoregressive pipeline is… 34 arXiv — NLP / Computation & Language research 21d ago More Is Not More: What Matters for Diversity in LLM Opinions? arXiv:2607.20429v1 Announce Type: new Abstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion… 30 arXiv — NLP / Computation & Language research 21d ago Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception arXiv:2607.20444v1 Announce Type: new Abstract: Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal. Yet it remains unclear how confidently models deceive and whether higher… 24 arXiv — NLP / Computation & Language research 21d ago Response drift across frontier large language models arXiv:2607.20454v1 Announce Type: new Abstract: All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human… 31 Page 6 of 10 · 500 articles ← Newer Older →