News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow r/LocalLLaMA community 11h ago 48 tg/s 440 prefill on my grandma's cluster (2xP40) (sort of) TL;DR: switching KV cache to f16 may give a boost in speed if using MTP and ngrams. I have a self-built "AI mega-cluster" with 2x P40s on a cheap Chinese motherboard and a Xeon CPU (around $1,100 to build, including water cooling for the GPUs). I was normally getting up to 15… 31 r/LocalLLaMA community 1d ago My local LLM demoscene generator can now watch its own output and rewrite it! I've updated my auto_demo_scener project with Ninfer support and a “rewrite based on video” feature that I thought you might find interesting. The project is basically an endless demoscene machine. A local LLM writes Three.js effects (from a library of editable prompts), you… 23 r/LocalLLaMA community 1d ago AMD unveils Threadripper Halo Station AMD Threadripper Halo Station CPU Ryzen Threadripper PRO 9995WX (Zen 5, "Shimada Peak") 96 cores / 192 threads Up to 5.4 GHz boost 384 MB L3 cache 350 W TDP 8-channel DDR5 128 PCIe 5.0 lanes System Memory 2 TB DDR5 (as shown at IFA) Accelerators 2 x Liquid Cooled AMD Instinct… 13 r/LocalLLaMA community 1d ago NVIDIA PAIR — Your Personal AI Cluster That is interesting, I got bunch of old hardware I could connect, wonder what the speed would looks like.   submitted by   /u/SpendLucky1273 [link]   [comments] 22 The Information — AI news-outlet 2d ago DeepSeek Plans Major Huawei Chip Order in New AI Data Center DeepSeek plans to install at least 160,000 Huawei AI chips at a data center in Inner Mongolia, Northern China, Bloomberg reported, citing people familiar with the situation. The project would support China’s push to replace Nvidia silicon in light of U.S. chip restrictions.… 11 arXiv — Machine Learning research 2d ago Selective Hypergraph Refinement for Frozen Graph Clustering arXiv:2609.03265v1 Announce Type: new Abstract: Existing graph-clustering methods typically improve clustering performance by optimizing model parameters and node representations. Effective means of further improving the clustering results of an already trained and frozen model,… 7 arXiv — Machine Learning research 2d ago Geometry-Aware Graph Construction via Adaptive Spectral Bandwidth Control arXiv:2609.03306v1 Announce Type: new Abstract: Kernelized graph methods - spectral clustering, diffusion maps, and sparse kernel -regression graphs - that use Gaussian kernels depend on the choice of Gaussian bandwidth sigma, which governs the spectral character of the local… 28 arXiv — NLP / Computation & Language research 2d ago Fixed Suffix Dependency Ratio: Quantifying the Dual-Track Mechanism of Gender Assignment in Latvian Loanwords arXiv:2609.03930v1 Announce Type: new Abstract: Existing research has repeatedly observed the tendency for English loanwords to cluster in the masculine gender across different recipient languages, yet the origin of this pattern remains difficult to determine, as fixed… 14 TechCrunch — AI news-outlet 2d ago Crusoe reportedly raises $3B at a $30B valuation The round came together after the data center developer reportedly secured a $13 billion contract with Jane Street. 32 r/MachineLearning community 3d ago Grounding LLMs with JEPA-based world models trained in simulation — has this been tried? [D] LLMs describe physics well but don't "understand" it in any grounded sense — they've learned statistical relationships between tokens like "falls" and "gravity", not actual physical intuition. This is basically the Mary's Room problem: Mary knows every physical fact about color… 22 arXiv — Machine Learning research 3d ago Tri-Band Channel Measurement-Enabled Multi-Layer Digital Twin for Terahertz Wireless Data Centers arXiv:2609.01699v1 Announce Type: new Abstract: The rapid growth of AI computing has driven increasing demands for flexible and high-capacity data-center interconnections. Owing to its ultra-wide bandwidth and high spatial reuse capability, terahertz (THz) communication has… 33 arXiv — Machine Learning research 3d ago Toward Explainable and Policy-Aware AI for Carbon Credit Price Prediction: A Research Framework for Emerging Carbon Markets arXiv:2609.01765v1 Announce Type: new Abstract: Carbon markets put a price on emissions, yet that price remains hard to forecast. Work in this area clusters on the EU and Chinese schemes, compresses regulatory text into a sentiment score, and reports accuracy without calibration… 38 arXiv — Machine Learning research 3d ago Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks arXiv:2609.01942v1 Announce Type: new Abstract: Bitcoin's pseudonymous nature makes it challenging to analyze user-level activity, since a single user may control multiple identifiers (addresses). Existing heuristic-based methods attempt to identify addresses belonging to the… 22 arXiv — Machine Learning research 3d ago Differentiable Electricity-Market Clearing for Gradient-Based Planning arXiv:2609.02646v1 Announce Type: new Abstract: Planning a large data center is difficult because a facility big enough to matter changes the electricity prices it will pay. Those prices are set by market clearing, a constrained optimization problem solved anew in every… 21 arXiv — Machine Learning research 3d ago Private Computation Space: Experience with Trusted Multi-Cluster Federated Learning for Agriculture arXiv:2609.01667v1 Announce Type: cross Abstract: Artificial Intelligence has shown to help improve agricultural practices, yet adoption remains limited: 69% of U.S. farmers have privacy concerns with sharing their data, and these concerns must be addressed before adoption is… 5 arXiv — NLP / Computation & Language research 3d ago IDEEA: training-free Input-Dependent stEEring via Activation cluster matching arXiv:2609.02089v1 Announce Type: new Abstract: Steering aligns large language models (LLMs) by injecting a bias into selected activations at inference time, offering a far cheaper alternative to weight-update methods such as supervised fine-tuning or reinforcement learning.… 35 Vercel — AI dev-tools 3d ago Basic build machines are now available on Pro and Enterprise Pro and Enterprise teams can now select Basic build machines. Basic build machines have 2 vCPUs and 8 GB of memory, offering a more cost-efficient option for smaller apps or agents that build with fewer resources. New Pro and Enterprise projects still default to Elastic build… 11 r/LocalLLaMA community 4d ago Confirmed bolting Q8 NGram into IQ4 Qwen no speed degradation This came from another thread or comment. I forgot exactly where, but the basic idea was to replace the 51B N-gram layer in Qwen 3.8 Next with a much higher precision version. Someone running a 5090 replaced the N-gram portion of their Qwen 3.8 UD Q4 model with BF16. Since I'm… 33 r/MachineLearning community 4d ago Where can I find legally usable datasets for advanced audio chord recognition? [D] I’m researching how to build or fine-tune an audio-to-chord-recognition engine comparable in ambition to Song Master Pro / Auralis Sound Prism. The goal is not basic major/minor chord detection. I need reliable recognition of dense harmonic material: jazz, soul, funk, neo-soul,… 26 The Information — AI news-outlet 4d ago Anthropic Leader Praises President Trump’s Data Center Stance at G20 Meeting Anthropic co-founder Tom Brown called on countries to build more data centers at the G20 Innovation Ministerial in North Carolina on Wednesday, arguing that adding more computing power would be the best way for governments to benefit from AI gains. Elon Musk made similar… 32 arXiv — Machine Learning research 4d ago Stochastic complexity of vectors containing cluster structure arXiv:2609.00084v1 Announce Type: new Abstract: This paper studies the problem of computing the stochastic probability (shortest code length) of the encoded vectors containing cluster structure using Normalized Maximum Likelihood (NML) model. This is of great theoretical and… 35 arXiv — Machine Learning research 4d ago Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You arXiv:2609.00374v1 Announce Type: new Abstract: Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption is restrictive for inference-only accelerators, frozen or third-party models, and memory-constrained deployments,… 10 arXiv — Machine Learning research 4d ago DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering arXiv:2609.00647v1 Announce Type: new Abstract: Multiple kernel $k$-means integrates complementary nonlinear similarities by learning a combination of base kernels. Its pointwise optimization, however, is sensitive to noisy and boundary samples and repeatedly operates on… 38 arXiv — Machine Learning research 4d ago Contribution-Aware Bandwidth Allocation for Multimodal Split Learning arXiv:2609.01406v1 Announce Type: new Abstract: Multimodal models are increasingly the default option for perception at the network edge, yet they are trained almost entirely in the datacenter, because a client holding several sensor streams cannot host an encoder per modality.… 24 arXiv — NLP / Computation & Language research 4d ago OUTLETS: Output-Length Prediction from Speculative Decoding Backbones arXiv:2609.01068v1 Announce Type: new Abstract: The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-length prediction can mitigate these issues, existing… 17 r/LocalLLaMA community 4d ago 4 x DGX Sparks vs AMD Epyc 9xx5 system I see a lot of people buy DGX Sparks, and turn them in to clusters to run large models. Wouldn't it be better to invest $16k into an AMD Epyc server with 768GB or even 384GB of 6000Mhz DDR5 ram, and let's say 2x3090s or 5080s, instead of 4 DGX Sparks with 512GB of ram? Epyc's… 17 The Information — AI news-outlet 4d ago Dell Raises Sales Forecast on AI Server Demand Dell said Tuesday that revenue for its quarter ending in July rose 58% from the year-ago quarter to $47 billion, beating the high end of its own forecast. The maker of personal computers and servers for AI data centers increased its outlook for the year ending in January 2027 by… 26 The Information — AI news-outlet 5d ago Elon Musk Calls on Countries to Build Data Centers at North Carolina G20 Meeting Elon Musk lamented a global “crisis of power” on Tuesday as the Tesla and SpaceX CEO spoke virtually at a G20 innovation event hosted in North Carolina by the White House Office of Science and Technology Policy and the Commerce Department. Musk cited a likely 15 gigawatt… 28 The Information — AI news-outlet 5d ago SoftBank’s SB Energy Files to Go Public SoftBank’s affiliate company SB Energy filed to go public, laying out in detail its ambition to become a neocloud. So far, however, as the filing showed, SB Energy has “no data center capacity” currently operating. SB Energy currently generates revenue from the sale of… 14 The Information — AI news-outlet 5d ago SpaceX Shakes Up Data Center Leadership After Aggressive Build-Out Elon Musk has shaken up the SpaceX team building the company’s data centers in recent weeks, replacing several leaders with executives from SpaceX’s rocket and satellite internet businesses, people familiar with the matter said. The data center reshuffle follows some civil… 17 Hugging Face Daily Papers research 5d ago On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Abstract Qwen3.8-Flash-Next is a sparse mixture-of-experts architecture combining hybrid gated delta-net and sparse attention layers, gated residual branches, and off-accelerator n-gram embeddings to improve efficiency, capability, and training stability. Generated by… 5 arXiv — Machine Learning research 5d ago NVE: A Separability and Coverage-Aware Internal Validation Metric for Biclustering arXiv:2608.29045v1 Announce Type: new Abstract: Biclustering, or co-clustering, aims to discover coherent submatrices by grouping rows and columns of a data matrix simultaneously. This local two-dimensional structure makes validation more difficult than in ordinary clustering,… 33 arXiv — Machine Learning research 5d ago Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models arXiv:2608.30081v1 Announce Type: new Abstract: Understanding which training samples influence a generated image is an important problem in generative modeling. In flow matching, training samples influence the generated image through the velocity field along the generation… 7 arXiv — NLP / Computation & Language research 5d ago GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon arXiv:2608.28667v1 Announce Type: new Abstract: The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM… 27 arXiv — NLP / Computation & Language research 5d ago A Hub of Short Rows Inflates Intrinsic Dimension Estimation of Token Embeddings arXiv:2608.29702v1 Announce Type: new Abstract: A token-embedding table holds a hub of short rows near its origin, and we show that this cluster biases what nearest-neighbor intrinsic-dimension (ID) estimators report. Because of the concentration of measure, a token is closer to… 15 The Information — AI news-outlet 5d ago Anthropic Said to Reach $35 Billion Compute Deal With Nvidia-Backed Lambda Anthropic signed a deal to rent $35 billion worth of compute capacity from Nvidia-backed cloud provider Lambda Labs, The Wall Street Journal reported Monday. Nvidia, an investor and supplier to Lambda , will supply its chips for the data center providing the capacity to… 5 The Information — AI news-outlet 5d ago Saudi AI Firm Humain Partners With Together AI, MinIO on Data Centers Humain, the Saudi state-owned AI company founded last year, announced a series of partnerships with U.S. startups on Monday tied to data centers in Riyadh and Dammam. Together AI, which rents out computing power to developers who train and run AI models, will share a portion of… 17 The Information — AI news-outlet 5d ago Trump Condemns Communities Opposed to Data Centers Communities that don’t want data centers will end up being “backwards and poor,” President Donald Trump said on Monday, signaling his backing of the politically-charged industry ahead of the upcoming midterm elections. “If we kill the Golden Goose, you will only have yourselves… 37 Ars Technica — AI news-outlet 7d ago Inside Meta’s push to put robots to work in data centers The company is testing robots on tasks that can performed by technicians. 28 The Information — AI news-outlet 7d ago Musk Responds to The Information’s Report About SpaceX Setting Up a Turbine Blade Factory SpaceX CEO Elon Musk on Saturday responded to The Information’s report that his company was laying the groundwork for a foundry to build turbine vanes and blades to avoid a power shortage for AI data centers. “By doing in-house casting [of vanes and blades] at SpaceX, we can… 14 The Information — AI news-outlet 8d ago Exclusive: SpaceX Lays Groundwork For Turbine-Blade Factory to Solve Data Center Power Crunch Clues are emerging that Elon Musk intends to bypass the power supply chain for AI data centers in a way others assumed was impossible, by making highly complex components himself. It’s a move no one else has been bold enough to attempt. 27 r/LocalLLaMA community 8d ago Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering Exo labs making some very exciting and interesting claims. The headline is bandwidth scales linearly on Mac Studio clusters with their solution. There is a thread over at localllm subreddit ( https://www.reddit.com/r/LocalLLM/s/qEYLOFaYwc ) where one of their employees speaks… 17 TechCrunch — AI news-outlet 8d ago Nvidia’s AI advantage is moving beyond the GPU The new generation of data center systems is increasing efficiency with smarter traffic control instead of just more processor cycles. 9 r/LocalLLaMA community 8d ago Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks Hey all, and hello fellow DGX Spark-ers! Today I managed some pretty crazy numbers: 181 tok/s aggregate on 2× DGX Spark on Qwen3.8-Flash-Next at 512Kcontext (2.8M kvc) I hit 181 tok/s aggregate today across a multi-agent fleet on a 2-node DGX Spark cluster. Single-stream decode… 22 Stratechery (Ben Thompson) community 9d ago 2026.35: Internet Hype and Real World Change The best Stratechery content from the week of August 24, 2026 including the breaker's advantage, the new battle for HDMI1, and how data center discourse ends. 7 arXiv — Machine Learning research 9d ago ClusterAttention: A training-free speedup of bidirectional attention arXiv:2608.26965v1 Announce Type: new Abstract: This paper introduces ClusterAttention, a general training-free speedup of bidirectional attention layers. Existing sparse attention methods either rely on structure in the input, such as order in language or spatial proximity in… 13 arXiv — Machine Learning research 9d ago Inductive Correlation Clustering with Graph Neural Networks arXiv:2608.27153v1 Announce Type: new Abstract: Correlation Clustering (CC) is a natural formulation of clustering in combinatorial optimization, which uses a graph representation of the input and does not require a pre-specified number of clusters. Given $n$ objects and a… 6 OpenAI official-blog 9d ago Supporting Thailand’s next generation of AI startups OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into trusted products. 26 ThursdAI news-outlet 9d ago NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley From CoreWeave - join Alex and ThursdAI co-host, covering the last week of the summer in AI, with 4 Flash models, Datacenter debate & more AI news 11 Ars Technica — AI news-outlet 10d ago AI industry says Trump plans to tax chips in the “single dumbest way imaginable” Tech industry is perplexed by Trump’s plan to win AI race by taxing data centers. 35 Page 1 of 10 · 500 articles Older →