News / #hardware Tag Hardware 470 articles archived under #hardware · RSS Sign in to follow NVIDIA Developer Blog official-blog 10d ago How to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure Running a dedicated Kubernetes cluster per team often results in more isolation than an organization requires. While one cluster can be successfully shared... 9 r/LocalLLaMA community 10d ago "Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out there and share knowledge if there is any interest. I also wasn't satisfied with… 32 Hugging Face Daily Papers research 11d ago Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Abstract Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and… 8 arXiv — Machine Learning research 11d ago Topology-Aware Data Movement for Disaggregated GPU Inference arXiv:2607.28633v1 Announce Type: new Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate GPU pools, the KV cache must be transferred between them. For a 70B model this is… 6 arXiv — NLP / Computation & Language research 11d ago Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM arXiv:2607.28635v1 Announce Type: new Abstract: In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not adequately capture minority topics. To tackle this challenge, our paper presents… 34 arXiv — NLP / Computation & Language research 11d ago TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models arXiv:2607.28896v1 Announce Type: cross Abstract: Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about… 26 r/LocalLLaMA community 12d ago Setting up of a 16xGB10 (DGX Spark) cluster Preparing this to be able to run locally frontier level open models. Deepseek v4 pro, Kimi K3, future ones like GLM 5.5 and Minimax M4. 16x Asus GX10 linked by mikrotik crs804-4ddq with 4 breakout cables of 400 to 100gbit. Most probable I will be running 2 models on 8x cluster… 13 r/LocalLLaMA community 12d ago Real-world reality check on Qwen for autonomous coding agents TLDR below 👇🏼 I’ve seen a lot of hype around Qwen 3.6 35B and 3.5 120B lately, especially regarding coding and tool-use capabilities. On this subreddit it is the defacto recommended model for everyone without a Datacenter at home. I’ve been running Qwen 3.5 120B… 9 r/MachineLearning community 12d ago Bytedance is using seedance 2.5 to automatically generate animated study guides in gauth. interesting use case for ai video [N] bytedance is using seedance 2.5 to automatically generate animated study guides in gauth. interesting use case for ai video saw this business insider article about how bytedance integrated their seedance 2.5 video model into their study app (gauth). basically generates animated… 34 r/LocalLLaMA community 12d ago EU AI Act takes effect tomorrow, August 2, 2026. 🤡 Basically you now have to mark all AI generated images, audio, video and text as AI generated. :P   submitted by   /u/xoxaxo [link]   [comments] 32 r/LocalLLaMA community 13d ago With release of Deepseek V4 I wanted see how the model sizes are trending over time. The trend is that by this time next year, we probably will have Opus 4.5 level models on consumer grade laptops! I was surprised to see that Deepseek V4 Flash is extremely smart and small enough to fit in setup that can be built with < $50,000. Expensive, but not a datacenter. So I wanted to see the trend over time of model sizes and their scores and created above plots using Opus/Sonnet… 32 r/MachineLearning community 13d ago Learning path to fully understand the Kimi K3 technical report?[D] Hi everyone, Can anyone suggest a learning path to fully understand the technical report for Kimi K3? My background: - I've taken a graduate-level deep learning course. - I understand the Transformer architecture, attention, and the basics of LLMs. - I'm familiar with DeepSeek's… 15 TechCrunch — AI news-outlet 13d ago SpaceX won’t remove all of xAI’s unpermitted turbines for another year SpaceX is building a new power plant for xAI's Colossus data centers, but it won't remove existing, unpermitted turbines for many more months. 15 arXiv — Machine Learning research 14d ago DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement arXiv:2607.27761v1 Announce Type: new Abstract: In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across different views often suffer from misalignment, leading to the partial view… 14 arXiv — Machine Learning research 14d ago Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata arXiv:2607.28338v1 Announce Type: new Abstract: Clustered Federated Learning (CFL) addresses data heterogeneity in federated settings by grouping clients with similar data distributions to enable effective training. Existing methods face a trade-off between privacy preservation,… 15 TechCrunch — AI news-outlet 14d ago Investors love AI, as long as you’re a cloud host Amazon isn't slowing down on data center spending — but investors don't seem to mind. 29 NVIDIA Developer Blog official-blog 14d ago NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We... 9 TechCrunch — AI news-outlet 14d ago Nscale buys Anyscale as it seeks to own more of the AI compute stack British AI neocloud Nscale is buying software startup Anyscale, which helps companies scale their AI workloads across data centers and servers. 12 arXiv — Machine Learning research 15d ago FloDR: An invertible dimensionality reduction method based on a normalising flow arXiv:2607.26278v1 Announce Type: new Abstract: It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point… 22 arXiv — Machine Learning research 15d ago PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems arXiv:2607.26710v1 Announce Type: new Abstract: The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling. Under stringent grid… 6 arXiv — Machine Learning research 15d ago Archetypes or ability? Clustering for modelling student mathematical competence arXiv:2607.26063v1 Announce Type: cross Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different… 36 arXiv — Machine Learning research 15d ago Randomizing the Number of Centers in k-means++ arXiv:2607.26202v1 Announce Type: cross Abstract: The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $\Theta(\log k)$. We consider the same… 17 r/LocalLLaMA community 15d ago Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar? I bought an RTX 5090 last year just to run 27B models natively. I even fine-tuned it with my own data using LoRA, building RAGs and was pretty damn happy with the results at first. But, Q8 quantization 130k context was barely squeezing through. Naturally, I bought two RTX 6000… 37 r/LocalLLaMA community 15d ago "Uncensored" LLMs are measurably more optimistic than their base models Hi. Many people think uncensored models are basically the same model that just doesn't refuse, but... I was recently checking whether uncensored models would give me better answers for stock market predictions (my idea was: the uncensored one will tell you the truth and won't be… 38 arXiv — NLP / Computation & Language research 16d ago An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data arXiv:2503.07303v3 Announce Type: replace Abstract: Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulaic texts, which are characterized by repetition and constrained expression, tend… 37 r/MachineLearning community 16d ago Would you take an MLE role with 50% base salary increase with rigid pto policy and no 401K match? [D] I got an offer for an MLE role where basically i will develop ML/GENAI solutions for clients completely and then hand them over and move to other projects. I think its a consulting ML role, although from a salary perspective its amazing and base is 50% more and if i add my new… 23 TechCrunch — AI news-outlet 16d ago Data centers may face temporary power cuts to prevent blackouts on largest US grid The largest grid operator in the U.S. says it will cut power to large data centers to prevent blackouts starting next year. 7 Hugging Face Daily Papers research 16d ago Characterizing Warp Divergence from Pascal to Blackwell Abstract Since Volta introduced Independent Thread Scheduling (ITS), NVIDIA GPUs have been widely assumed to handle warp divergence in a fixed manner. We test this assumption across Ampere, Hopper, and datacenter and consumer Blackwell GPUs, using pre-ITS Pascal as a baseline.… 23 arXiv — Machine Learning research 17d ago What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation arXiv:2607.22781v1 Announce Type: new Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph models, for example, can achieve high future-link AUC while basic graph… 34 r/LocalLLaMA community 17d ago I ran the 35B agentic comparison someone asked for (stock vs Ornith vs KAT-Coder, 120 runs) Someone in the comments of my 27B post-train bakeoff asked for the 35B version, so I ran it. Same setup as last time: fresh Coder workspaces on my k8s cluster, each driving my own agent (Hermes) headlessly, models on llama.cpp via llama-swap on one 5090, every call traced… 21 Ars Technica — AI news-outlet 17d ago Verizon seeks AI profits with mini data centers, $1B dark fiber deal with Google Telecom expects AI revenue from dark fiber deals and retrofitted data centers. 23 r/LocalLLaMA community 17d ago Viable ways to run K3 locally just curious how would people run it cheap if they really want kimi k3. dgx spark / strix halo clusters optane persistent memory platform + some gpus mac studio clusters orange pi 6 clusters ssd streaming + gpus multiple ddr3 + connectx 5 rdma clients two dgx stations power 10… 37 r/LocalLLaMA community 17d ago You can now fine-tune my 3.96M-parameter TTS on your own voice or language When I released Inflect v2 last week, I thought most people would ask whether a TTS model this small actually sounded decent. Instead, I kept getting two questions: “Can I train it on my own voice?” “Can I move it to another language?” At the time, my answer was basically:… 16 r/LocalLLaMA community 17d ago Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend & maybe results by next week. Weights are supposed to hit Hugging Face today… 18 r/LocalLLaMA community 18d ago I want to run Kimi K3 at home, so I’m trying to make 2.8T-scale experimentation cheaper Hey r/LocalLLaMA , I’m a retired engineer with a background in distributed computing, currently running a 1-person startup. Like many people here, I’d love to experiment with 2T+ MoE models locally. The problem is that I don’t have an H100 cluster in my living room. So I’ve been… 12 arXiv — Machine Learning research 18d ago RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this… 31 arXiv — NLP / Computation & Language research 18d ago Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA arXiv:2607.21861v1 Announce Type: new Abstract: We study baking documents directly into the weights of a 4-bit Gemma-4-e4b model via LoRA, so a system can answer questions about a corpus closed-book: no retrieval and no context-window budget. Across roughly 100 training runs… 22 r/LocalLLaMA community 18d ago Will prices finally go down? I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments into AI data centers and had no use for then, had to rent them, same thing with… 38 r/LocalLLaMA community 18d ago Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Claude code (with DS in CLIProxyAPI ) takes nearly 4 times longer than the fastest… 38 r/LocalLLaMA community 18d ago 90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B Last week someone here said ThinkingCap and Fable Fusion "really do beat the OG" for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the story: each run was a fresh Coder workspace on my k8s cluster driving my own… 17 r/LocalLLaMA community 18d ago World's First(?) Underwhelming AMD Ryzen AI Halo Cluster LTT Labs recently received the Linux version of the AMD Ryzen AI Halo for testing, but it turns out that AMD had intended to send the Windows version. Through this stroke of misfortunate, we were fortunate enough to have two Ryzen AI Halos for a short period of time and the… 14 TechCrunch — AI news-outlet 19d ago One fallen power line exposed a growing AI data center problem. Here’s how to fix it. A close call in Northern Virginia revealed just how poorly data centers respond to grid disruptions. Here's how to fix the problem. 34 Ars Technica — AI news-outlet 20d ago AI firms want more data centers; Trump's EPA may give neighbors less say Rule would allow states to decide how much—if any—public input there can be. 12 Hacker News — AI on Front Page community 20d ago IRGC claims it destroyed Amazon's Bahrain data center Article URL: https://houseofsaud.com/irgc-claims-destroyed-amazon-bahrain-data-center/ Comments URL: https://news.ycombinator.com/item?id=49033240 Points: 262 # Comments: 332 8 arXiv — Machine Learning research 21d ago External Clustering Validation by the Homogeneity-Parsimony Trade-off arXiv:2607.20799v1 Announce Type: new Abstract: Scalar metrics are often used to evaluate clusterings against known classes, but they can obscure a fundamental trade-off: clusterings should be informative about class labels while avoiding unnecessary fragmentation. Here we… 34 arXiv — Machine Learning research 21d ago Regularized Optimization on Grassmann Manifold: Theory, Algorithm and Applications arXiv:2607.21039v1 Announce Type: new Abstract: Spectral methods are among the most widely used techniques for community detection, clustering, and graph learning. Their performance, however, critically depends on the accurate estimation of the underlying spectral subspace and… 7 arXiv — Machine Learning research 21d ago CASC: Causal Adversarial Subspace Clustering for Multivariate Spatiotemporal Data arXiv:2607.21088v1 Announce Type: new Abstract: Deep subspace clustering plays a critical role in applications involving multivariate spatiotemporal data, such as sea ice monitoring, disease spread analysis, and tracking neuro-degeneration over time. Despite recent advances,… 29 arXiv — Machine Learning research 21d ago Semantic-Aware Task Clustering for Constructive and Cooperative Multi-Tasking arXiv:2607.21426v1 Announce Type: new Abstract: Cooperative multi-task semantic communication (CMT-SemCom) improves task execution performance by leveraging shared representations. However, as we demonstrated in [1], cooperative multi-tasking can be either constructive or… 16 r/MachineLearning community 21d ago NeurIPS E and D, Average rating 3 and average confidence 4, I can rebuttal and address all their concerns? Do I still have a decent shot or unlikely ?[R] NeurIPS E and D track review are out today and the average rating I received is a 3 and confidence is a 4. I can correct and address all their concerns. Do I still have a genuine shot of getting in or is it basically impossible at this point since none of my scores are a 4 or 5?… 37 arXiv — Machine Learning research 22d ago SCPP: A Unified Python Library for Soft Clustering arXiv:2607.19620v1 Announce Type: new Abstract: In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a canonical, scikit-learn-compatible estimator interface that standardizes model training,… 8 Page 2 of 10 · 470 articles ← Newer Older →