News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow Hacker News — AI on Front Page community 20d ago I've operated petabyte-scale ClickHouse clusters for 5 years Article URL: https://www.tinybird.co/blog/what-i-learned-operating-clickhouse Comments URL: https://news.ycombinator.com/item?id=49601138 Points: 214 # Comments: 78 14 The Information — AI news-outlet 20d ago Desperation to Get Data Centers Online Is Reshaping Companies' Bargaining Power No two contracts for AI data center buildouts look quite the same. It’s a matchmaking process on every level involving a range of companies: those that secure land and power, those that build and operate data centers, the financiers that lend money for the projects, and the… 26 Ars Technica — AI news-outlet 21d ago The complex corporate web behind a $3.2 billion AI data center When multiple companies are behind one project, who bears responsibility for problems? 38 Ars Technica — AI news-outlet 21d ago The complex corporate web behind a $3.2 billion AI data center When multiple companies are behind one project, who bears responsibility for problems? 16 arXiv — Machine Learning research 21d ago Compute-in-Memory Attention: A Time-Domain Analog Softmax Circuit with RC-Tunable Temperature arXiv:2609.04266v1 Announce Type: cross Abstract: Softmax is a key operation in Transformer attention, but its exponentiation and normalization add significant overhead in compute-in-memory (CIM) accelerators, especially when analog attention scores must first be converted to… 32 arXiv — Machine Learning research 21d ago Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters arXiv:2609.04417v1 Announce Type: cross Abstract: Distributed AI training involves recurring rounds of data exchange between multiple pairs of GPU nodes. Slowdown in even one flow due to congestion can cause the entire communication round to slowdown. Current approaches for… 5 r/LocalLLaMA community 22d ago 48 tg/s 440 prefill on my grandma's cluster (2xP40) (sort of) TL;DR: switching KV cache to f16 may give a boost in speed if using MTP and ngrams. I have a self-built "AI mega-cluster" with 2x P40s on a cheap Chinese motherboard and a Xeon CPU (around $1,100 to build, including water cooling for the GPUs). I was normally getting up to 15… 31 r/LocalLLaMA community 22d ago My local LLM demoscene generator can now watch its own output and rewrite it! I've updated my auto_demo_scener project with Ninfer support and a “rewrite based on video” feature that I thought you might find interesting. The project is basically an endless demoscene machine. A local LLM writes Three.js effects (from a library of editable prompts), you… 23 r/LocalLLaMA community 23d ago AMD unveils Threadripper Halo Station AMD Threadripper Halo Station CPU Ryzen Threadripper PRO 9995WX (Zen 5, "Shimada Peak") 96 cores / 192 threads Up to 5.4 GHz boost 384 MB L3 cache 350 W TDP 8-channel DDR5 128 PCIe 5.0 lanes System Memory 2 TB DDR5 (as shown at IFA) Accelerators 2 x Liquid Cooled AMD Instinct… 13 r/LocalLLaMA community 23d ago NVIDIA PAIR — Your Personal AI Cluster That is interesting, I got bunch of old hardware I could connect, wonder what the speed would looks like.   submitted by   /u/SpendLucky1273 [link]   [comments] 22 The Information — AI news-outlet 24d ago DeepSeek Plans Major Huawei Chip Order in New AI Data Center DeepSeek plans to install at least 160,000 Huawei AI chips at a data center in Inner Mongolia, Northern China, Bloomberg reported, citing people familiar with the situation. The project would support China’s push to replace Nvidia silicon in light of U.S. chip restrictions.… 11 arXiv — Machine Learning research 24d ago Selective Hypergraph Refinement for Frozen Graph Clustering arXiv:2609.03265v1 Announce Type: new Abstract: Existing graph-clustering methods typically improve clustering performance by optimizing model parameters and node representations. Effective means of further improving the clustering results of an already trained and frozen model,… 7 arXiv — Machine Learning research 24d ago Geometry-Aware Graph Construction via Adaptive Spectral Bandwidth Control arXiv:2609.03306v1 Announce Type: new Abstract: Kernelized graph methods - spectral clustering, diffusion maps, and sparse kernel -regression graphs - that use Gaussian kernels depend on the choice of Gaussian bandwidth sigma, which governs the spectral character of the local… 28 arXiv — NLP / Computation & Language research 24d ago Fixed Suffix Dependency Ratio: Quantifying the Dual-Track Mechanism of Gender Assignment in Latvian Loanwords arXiv:2609.03930v1 Announce Type: new Abstract: Existing research has repeatedly observed the tendency for English loanwords to cluster in the masculine gender across different recipient languages, yet the origin of this pattern remains difficult to determine, as fixed… 14 TechCrunch — AI news-outlet 24d ago Crusoe reportedly raises $3B at a $30B valuation The round came together after the data center developer reportedly secured a $13 billion contract with Jane Street. 32 r/MachineLearning community 24d ago Grounding LLMs with JEPA-based world models trained in simulation — has this been tried? [D] LLMs describe physics well but don't "understand" it in any grounded sense — they've learned statistical relationships between tokens like "falls" and "gravity", not actual physical intuition. This is basically the Mary's Room problem: Mary knows every physical fact about color… 22 arXiv — Machine Learning research 25d ago Tri-Band Channel Measurement-Enabled Multi-Layer Digital Twin for Terahertz Wireless Data Centers arXiv:2609.01699v1 Announce Type: new Abstract: The rapid growth of AI computing has driven increasing demands for flexible and high-capacity data-center interconnections. Owing to its ultra-wide bandwidth and high spatial reuse capability, terahertz (THz) communication has… 33 arXiv — Machine Learning research 25d ago Toward Explainable and Policy-Aware AI for Carbon Credit Price Prediction: A Research Framework for Emerging Carbon Markets arXiv:2609.01765v1 Announce Type: new Abstract: Carbon markets put a price on emissions, yet that price remains hard to forecast. Work in this area clusters on the EU and Chinese schemes, compresses regulatory text into a sentiment score, and reports accuracy without calibration… 38 arXiv — Machine Learning research 25d ago Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks arXiv:2609.01942v1 Announce Type: new Abstract: Bitcoin's pseudonymous nature makes it challenging to analyze user-level activity, since a single user may control multiple identifiers (addresses). Existing heuristic-based methods attempt to identify addresses belonging to the… 22 arXiv — Machine Learning research 25d ago Differentiable Electricity-Market Clearing for Gradient-Based Planning arXiv:2609.02646v1 Announce Type: new Abstract: Planning a large data center is difficult because a facility big enough to matter changes the electricity prices it will pay. Those prices are set by market clearing, a constrained optimization problem solved anew in every… 21 arXiv — Machine Learning research 25d ago Private Computation Space: Experience with Trusted Multi-Cluster Federated Learning for Agriculture arXiv:2609.01667v1 Announce Type: cross Abstract: Artificial Intelligence has shown to help improve agricultural practices, yet adoption remains limited: 69% of U.S. farmers have privacy concerns with sharing their data, and these concerns must be addressed before adoption is… 5 arXiv — NLP / Computation & Language research 25d ago IDEEA: training-free Input-Dependent stEEring via Activation cluster matching arXiv:2609.02089v1 Announce Type: new Abstract: Steering aligns large language models (LLMs) by injecting a bias into selected activations at inference time, offering a far cheaper alternative to weight-update methods such as supervised fine-tuning or reinforcement learning.… 35 Vercel — AI dev-tools 25d ago Basic build machines are now available on Pro and Enterprise Pro and Enterprise teams can now select Basic build machines. Basic build machines have 2 vCPUs and 8 GB of memory, offering a more cost-efficient option for smaller apps or agents that build with fewer resources. New Pro and Enterprise projects still default to Elastic build… 11 r/LocalLLaMA community 25d ago Confirmed bolting Q8 NGram into IQ4 Qwen no speed degradation This came from another thread or comment. I forgot exactly where, but the basic idea was to replace the 51B N-gram layer in Qwen 3.8 Next with a much higher precision version. Someone running a 5090 replaced the N-gram portion of their Qwen 3.8 UD Q4 model with BF16. Since I'm… 33 r/MachineLearning community 25d ago Where can I find legally usable datasets for advanced audio chord recognition? [D] I’m researching how to build or fine-tune an audio-to-chord-recognition engine comparable in ambition to Song Master Pro / Auralis Sound Prism. The goal is not basic major/minor chord detection. I need reliable recognition of dense harmonic material: jazz, soul, funk, neo-soul,… 26 The Information — AI news-outlet 25d ago Anthropic Leader Praises President Trump’s Data Center Stance at G20 Meeting Anthropic co-founder Tom Brown called on countries to build more data centers at the G20 Innovation Ministerial in North Carolina on Wednesday, arguing that adding more computing power would be the best way for governments to benefit from AI gains. Elon Musk made similar… 32 arXiv — Machine Learning research 26d ago Stochastic complexity of vectors containing cluster structure arXiv:2609.00084v1 Announce Type: new Abstract: This paper studies the problem of computing the stochastic probability (shortest code length) of the encoded vectors containing cluster structure using Normalized Maximum Likelihood (NML) model. This is of great theoretical and… 35 arXiv — Machine Learning research 26d ago Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You arXiv:2609.00374v1 Announce Type: new Abstract: Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption is restrictive for inference-only accelerators, frozen or third-party models, and memory-constrained deployments,… 10 arXiv — Machine Learning research 26d ago DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering arXiv:2609.00647v1 Announce Type: new Abstract: Multiple kernel $k$-means integrates complementary nonlinear similarities by learning a combination of base kernels. Its pointwise optimization, however, is sensitive to noisy and boundary samples and repeatedly operates on… 38 arXiv — Machine Learning research 26d ago Contribution-Aware Bandwidth Allocation for Multimodal Split Learning arXiv:2609.01406v1 Announce Type: new Abstract: Multimodal models are increasingly the default option for perception at the network edge, yet they are trained almost entirely in the datacenter, because a client holding several sensor streams cannot host an encoder per modality.… 24 arXiv — NLP / Computation & Language research 26d ago OUTLETS: Output-Length Prediction from Speculative Decoding Backbones arXiv:2609.01068v1 Announce Type: new Abstract: The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-length prediction can mitigate these issues, existing… 17 r/LocalLLaMA community 26d ago 4 x DGX Sparks vs AMD Epyc 9xx5 system I see a lot of people buy DGX Sparks, and turn them in to clusters to run large models. Wouldn't it be better to invest $16k into an AMD Epyc server with 768GB or even 384GB of 6000Mhz DDR5 ram, and let's say 2x3090s or 5080s, instead of 4 DGX Sparks with 512GB of ram? Epyc's… 17 The Information — AI news-outlet 26d ago Dell Raises Sales Forecast on AI Server Demand Dell said Tuesday that revenue for its quarter ending in July rose 58% from the year-ago quarter to $47 billion, beating the high end of its own forecast. The maker of personal computers and servers for AI data centers increased its outlook for the year ending in January 2027 by… 26 The Information — AI news-outlet 26d ago Elon Musk Calls on Countries to Build Data Centers at North Carolina G20 Meeting Elon Musk lamented a global “crisis of power” on Tuesday as the Tesla and SpaceX CEO spoke virtually at a G20 innovation event hosted in North Carolina by the White House Office of Science and Technology Policy and the Commerce Department. Musk cited a likely 15 gigawatt… 28 The Information — AI news-outlet 26d ago SoftBank’s SB Energy Files to Go Public SoftBank’s affiliate company SB Energy filed to go public, laying out in detail its ambition to become a neocloud. So far, however, as the filing showed, SB Energy has “no data center capacity” currently operating. SB Energy currently generates revenue from the sale of… 14 The Information — AI news-outlet 26d ago SpaceX Shakes Up Data Center Leadership After Aggressive Build-Out Elon Musk has shaken up the SpaceX team building the company’s data centers in recent weeks, replacing several leaders with executives from SpaceX’s rocket and satellite internet businesses, people familiar with the matter said. The data center reshuffle follows some civil… 17 Hugging Face Daily Papers research 27d ago On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Abstract Qwen3.8-Flash-Next is a sparse mixture-of-experts architecture combining hybrid gated delta-net and sparse attention layers, gated residual branches, and off-accelerator n-gram embeddings to improve efficiency, capability, and training stability. Generated by… 5 arXiv — Machine Learning research 27d ago NVE: A Separability and Coverage-Aware Internal Validation Metric for Biclustering arXiv:2608.29045v1 Announce Type: new Abstract: Biclustering, or co-clustering, aims to discover coherent submatrices by grouping rows and columns of a data matrix simultaneously. This local two-dimensional structure makes validation more difficult than in ordinary clustering,… 33 arXiv — Machine Learning research 27d ago Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models arXiv:2608.30081v1 Announce Type: new Abstract: Understanding which training samples influence a generated image is an important problem in generative modeling. In flow matching, training samples influence the generated image through the velocity field along the generation… 7 arXiv — NLP / Computation & Language research 27d ago GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon arXiv:2608.28667v1 Announce Type: new Abstract: The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM… 27 arXiv — NLP / Computation & Language research 27d ago A Hub of Short Rows Inflates Intrinsic Dimension Estimation of Token Embeddings arXiv:2608.29702v1 Announce Type: new Abstract: A token-embedding table holds a hub of short rows near its origin, and we show that this cluster biases what nearest-neighbor intrinsic-dimension (ID) estimators report. Because of the concentration of measure, a token is closer to… 15 The Information — AI news-outlet 27d ago Anthropic Said to Reach $35 Billion Compute Deal With Nvidia-Backed Lambda Anthropic signed a deal to rent $35 billion worth of compute capacity from Nvidia-backed cloud provider Lambda Labs, The Wall Street Journal reported Monday. Nvidia, an investor and supplier to Lambda , will supply its chips for the data center providing the capacity to… 5 The Information — AI news-outlet 27d ago Saudi AI Firm Humain Partners With Together AI, MinIO on Data Centers Humain, the Saudi state-owned AI company founded last year, announced a series of partnerships with U.S. startups on Monday tied to data centers in Riyadh and Dammam. Together AI, which rents out computing power to developers who train and run AI models, will share a portion of… 17 The Information — AI news-outlet 27d ago Trump Condemns Communities Opposed to Data Centers Communities that don’t want data centers will end up being “backwards and poor,” President Donald Trump said on Monday, signaling his backing of the politically-charged industry ahead of the upcoming midterm elections. “If we kill the Golden Goose, you will only have yourselves… 37 Ars Technica — AI news-outlet 29d ago Inside Meta’s push to put robots to work in data centers The company is testing robots on tasks that can performed by technicians. 28 The Information — AI news-outlet 29d ago Musk Responds to The Information’s Report About SpaceX Setting Up a Turbine Blade Factory SpaceX CEO Elon Musk on Saturday responded to The Information’s report that his company was laying the groundwork for a foundry to build turbine vanes and blades to avoid a power shortage for AI data centers. “By doing in-house casting [of vanes and blades] at SpaceX, we can… 14 The Information — AI news-outlet 29d ago Exclusive: SpaceX Lays Groundwork For Turbine-Blade Factory to Solve Data Center Power Crunch Clues are emerging that Elon Musk intends to bypass the power supply chain for AI data centers in a way others assumed was impossible, by making highly complex components himself. It’s a move no one else has been bold enough to attempt. 27 r/LocalLLaMA community 29d ago Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering Exo labs making some very exciting and interesting claims. The headline is bandwidth scales linearly on Mac Studio clusters with their solution. There is a thread over at localllm subreddit ( https://www.reddit.com/r/LocalLLM/s/qEYLOFaYwc ) where one of their employees speaks… 17 TechCrunch — AI news-outlet 29d ago Nvidia’s AI advantage is moving beyond the GPU The new generation of data center systems is increasing efficiency with smarter traffic control instead of just more processor cycles. 9 r/LocalLLaMA community 1mo ago Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks Hey all, and hello fellow DGX Spark-ers! Today I managed some pretty crazy numbers: 181 tok/s aggregate on 2× DGX Spark on Qwen3.8-Flash-Next at 512Kcontext (2.8M kvc) I hit 181 tok/s aggregate today across a multi-agent fleet on a 2-node DGX Spark cluster. Single-stream decode… 22 Page 3 of 10 · 500 articles ← Newer Older →