News / #hardware Tag Hardware 470 articles archived under #hardware · RSS Sign in to follow arXiv — Machine Learning research 3h ago H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities arXiv:2608.12926v1 Announce Type: new Abstract: Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT)… 31 arXiv — Machine Learning research 3h ago TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures arXiv:2608.13212v1 Announce Type: new Abstract: Networked systems, from power grids to traffic networks and cloud clusters, carry loads across nodes with limited capacity. A node whose load exceeds its capacity fails and sheds its load onto its neighbors, which can trigger a… 5 arXiv — NLP / Computation & Language research 3h ago Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model arXiv:2608.13277v1 Announce Type: new Abstract: We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular… 36 r/LocalLLaMA community 11h ago Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation Ive been experimenting with task-aware GGUF quants for months, taking inspiration from TASA and TAQO but pushing the allocation lower to to the tensor level. The basic idea is to generate a custom imatrix from a category-specific corpus, measure where quantization causes damage,… 5 r/MachineLearning community 14h ago UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P] I’m working on performance regression detection using machine learning/anomaly detection. My setup is basically: Healthy runs are used to learn normal behaviour Regression runs are used to see whether the model detects the anomaly For each counter group I only have about 10… 33 arXiv — Machine Learning research 1d ago Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision arXiv:2608.12027v1 Announce Type: new Abstract: Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing… 13 arXiv — Machine Learning research 1d ago Clustered Randomized Smoothing for Stochastic Prediction Functions arXiv:2608.12037v1 Announce Type: new Abstract: Modern stochastic predictors can model rich, multi-modal outcome distributions. However, this expressive power comes with challenges in ensuring robust predictions $-$ a critical requirement in safety-critical domains. Randomized… 8 arXiv — Machine Learning research 2d ago STCAD: Scalable Trajectory Clustering and Anomaly Detection on Terabyte-Scale AIS Data arXiv:2608.10249v1 Announce Type: new Abstract: We present a scalable framework for unsupervised clustering of maritime trajectories derived from terabyte-scale Automatic Identification System (AIS) archives. Variable-length trajectories are encoded with a custom BERT-based… 29 arXiv — Machine Learning research 2d ago CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling arXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and navigational realism. We propose the Continuous Regression Hybrid Transformer… 27 arXiv — Machine Learning research 2d ago Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts arXiv:2608.10605v1 Announce Type: new Abstract: In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and training recipe, optimizing loss under compute… 38 r/LocalLLaMA community 2d ago I will be parting with my 4x Spark Cluster. Laid off then my partner of 10 years said he's leaving, have to move, etc... I will post the r/hardwareswap link when I make it. I'm willing to add some incentive for r/LocalLLaMA folks. I will also add the super node configs and all the cool stuff that may not be apparent that… 26 r/LocalLLaMA community 3d ago I ran Muse Glimmer @ 1M context - All tests passed. Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Muse Glimmer 30B running the day after release — then pushed its context from the… 36 arXiv — Machine Learning research 3d ago When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes arXiv:2608.07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache management an attractive lever: a policy that raised the hit rate would cut… 28 arXiv — Machine Learning research 3d ago Exact Rank-Space KL Projection for Shared-Marginal Low-Rank Factors: Application to Doubly Stochastic Clustering arXiv:2608.08642v1 Announce Type: new Abstract: We study exact Kullback--Leibler (KL) projection for low-rank factorizations whose two nonnegative factors have prescribed row marginals and a shared, learned column marginal. For arbitrary positive row marginals of equal total… 24 r/LocalLLaMA community 3d ago I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter MoonViT image encoder, then trained a 40.1M-parameter connector between them on… 16 Ars Technica — AI news-outlet 3d ago Amazon backs power plant that may become top source of US climate pollution Amazon announces first off-the-grid data center in race to reap AI profits. 27 MIT News — AI research 3d ago With a feel for physics, AI models simulate a wider range of real-world scenarios “GeoPT” helps AI models understand the basics of physics so they can simulate how objects respond to things like wind and water more efficiently and accurately. 8 arXiv — Machine Learning research 4d ago SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction arXiv:2608.06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We present SNI-GNN, a SmartNIC-assisted full-graph training system that reduces… 21 arXiv — Machine Learning research 4d ago Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs arXiv:2608.06990v1 Announce Type: new Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. Among various clustering methods, hierarchical clustering, density-based clustering, and graph clustering stand out as… 10 arXiv — Machine Learning research 4d ago Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU arXiv:2608.07323v1 Announce Type: new Abstract: We test whether decoder-only language-model FFNs require SwiGLU's open positive tail. We introduce MemGLU as a closed-tail comparator derived from a memristive branch geometry. Across paired 9M and 30M pretraining runs with three… 27 arXiv — NLP / Computation & Language research 4d ago Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models arXiv:2608.06429v1 Announce Type: new Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce… 28 r/LocalLLaMA community 4d ago endless-frontier/BigBang-v1 - qwen 3.5 finetunes table bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline number is basically meaningless. "Performance between DeepSeek Flash (old one)… 26 r/LocalLLaMA community 5d ago Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size Firstly a big thanks to the poster "hellohazine", he basically only removed the multi-lingual fat of the model and just kept the English language intact. It is the exact model, and the rest of the model still intact with all of its high intelligence. I think that was a brilliant… 29 TechCrunch — AI news-outlet 5d ago Planned Amazon data center could become the biggest climate polluter in the U.S. As part of a planned Texas data center, Amazon is investing in an on-site power plant that could reportedly become the largest source of climate pollution in the United States. 9 r/LocalLLaMA community 5d ago Showoff Saturday: Local 4x 6000 Pro (multi-year progression) Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological order! With the pricing apocalypse meaning less builds shared here recently,… 38 r/LocalLLaMA community 5d ago My first run of Kimi K3 locally. Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to run. Goal will be to get all the GPUs in one system and without RPC, I should probably see 2-3x faster speed. Running… 7 Simon Willison community 5d ago Now we have a timeline of the OpenAI accidental attack against Hugging Face My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News. I think one of the most interesting details here might be tucked away in that first bulletin point: May 7: OpenAI starts a new training run for an experimental,… 16 Simon Willison community 5d ago Now we have a timeline of the OpenAI accidental attack against Hugging Face My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News. I think one of the most interesting details here might be tucked away in that first bulletin point: May 7: OpenAI starts a new training run for an experimental,… 37 r/LocalLLaMA community 5d ago Claude Code in 9 lines python I was wondering what a minimal coding agent implementation would look like that can be used like Claude Code or Codex Not feature-by-feature of course but basically stripping everything out that is not needed here is what I came up with: 9 lines of python no 3rd party deps… 29 r/LocalLLaMA community 5d ago Has anyone here fiddled with TPUs for inference ? I discovered recently that Google uses their own TPUs, like tiny ASIC cards like the toy ones that existed for bitcoin. And while it sounds inefficient the fact they use thousands of them because...they can...means at scale they aren't so bad. Has no one here given them a try? I… 26 r/LocalLLaMA community 6d ago PSA for anyone with multiple V620's or other gfx1030 cards having problems making llama.cpp tensor split work -- set "-ub 384" and -b to a multiple of that depending on number of GPUs Basically what the title says. For me, it would always crash and burn trying to use tensor split. Apparently, there's some bug where GPU memory gets corrupted with the default microbatch (512) or higher. I will be opening an issue report on the llama.cpp GitHub if there isn't… 6 r/LocalLLaMA community 6d ago Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM? Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the community feedback on what I could do to get more RAM for headroom. The DGX has an… 32 r/LocalLLaMA community 6d ago Got job as Director of AI and Systems development self-taught Hey everyone, I just wanted to share my journey here for some motivation. Three years ago, I saw the sudden spike in AI and realized it was the future of tech. My goal at the time was to be an indie game dev, and seeing that AI could write basic code, I told myself I needed to… 20 arXiv — Machine Learning research 7d ago Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation arXiv:2608.05880v1 Announce Type: new Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability… 12 arXiv — Machine Learning research 7d ago CohortHijack: Robustness of Single Cell Annotation to Companion Cell Removal arXiv:2608.05900v1 Announce Type: new Abstract: Many single-cell annotation tools refine an initial cell label using nearby cells or cluster-level voting. We study whether this refinement can be manipulated without changing the target cell. We introduce CohortHijack, a… 23 r/LocalLLaMA community 8d ago I get that AI labs need to make money, but zero-warning price spikes are a nightmare for production builds Seen a ton of posts today about the DeepSeek API price hike. Half the feed is doom-posting, the other half is explaining basic GPU economics. Honestly, I get the cost side. Sub-cent tokens were never gonna last forever. But what actually sucks is the zero-day notice. Dropping a… 23 arXiv — Machine Learning research 8d ago On Hamming-Lipschitz Type Stability of the Subdominant (Minmax) Ultrametric: Theory and Simple Proofs arXiv:2608.04014v1 Announce Type: new Abstract: The subdominant (minmax) ultrametric is a canonical tree-structured summary of a dissimilarity matrix, arising equivalently as the ultrametric induced by single-linkage clustering. While its classical stability theory is usually… 17 arXiv — Machine Learning research 8d ago Random features for Grassmannian kernel approximation with bounded rank-one projections arXiv:2608.04227v1 Announce Type: new Abstract: We propose a family of random feature maps for scalable kernel machines on low-dimensional subspaces, ie on the Grassmannian manifold. Such representations are useful when data classes or clusters are well described by the span of… 6 Hacker News — AI on Front Page community 8d ago Nashville uses eminent domain to block data center near zoo Article URL: https://www.costar.com/article/970809918/nashville-council-approves-eminent-domain-action-to-halt-data-center-project Comments URL: https://news.ycombinator.com/item?id=49191624 Points: 203 # Comments: 216 9 r/LocalLLaMA community 8d ago Deepseek V4 Flash just hit Colibri, does anyone have numbers? I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older hardware will likely be slow - no Unsloth GGUF supported either. I'd be curious what people are getting with V100s, R9700s,… 8 r/LocalLLaMA community 8d ago I updated my localy run benchmark with DeepSeek V4 Flash 0731 It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartoswski with Dspark at 1K t/s prefill and 90 t/s gen (average). I tried different sampling params, you can check the detail. It's very efficient while scoring the best yet. Too bad… 12 arXiv — Machine Learning research 9d ago Learning and Clustering on Temporal Graphs: Principles, Primitives, and Pooling arXiv:2608.03696v1 Announce Type: new Abstract: This work focuses on the problem of learning on temporal graphs, with particular emphasis on the task of clustering: obtaining coarse-grained representations by aggregating information from nodes, edges, and temporal dynamics - a… 34 arXiv — NLP / Computation & Language research 9d ago Predicting Deep Neural Network Training Outcomes from Early Training Telemetry arXiv:2608.03709v1 Announce Type: new Abstract: Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single training run's own early telemetry - per-epoch loss,… 32 Ars Technica — AI news-outlet 9d ago Texas halts data center connections to power grid amid overwhelming demand Governor who touted Texas as AI “epicenter” pauses data center grid connections. 36 r/LocalLLaMA community 9d ago Kimi K3 full model running on 16x GB10 cluster at 20+tps Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the… 13 TechCrunch — AI news-outlet 9d ago Texas halts new data centers as governor calls for audits Texas Governor Greg Abbott has paused new data center development until an audit has been completed. 4 TechCrunch — AI news-outlet 9d ago Is the future of data centers portable? Runware builds a pod to find out On Tuesday, AI infrastructure company Runware announced the launch of its own modular data center called Sonic Inference Pod. 18 arXiv — Machine Learning research 10d ago Ensemble of Unsupervised Deep Learning for Clustering Imbalanced Tabular Data arXiv:2608.00346v1 Announce Type: new Abstract: Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy. Unsupervised deep clustering can be immune to class… 6 arXiv — Machine Learning research 10d ago RHEA: Reliability-Harmonized Reconstruction and Assignment for Robust Multimodal-Attributed Graph Clustering arXiv:2608.00621v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry heterogeneous attributes such as text and images over a relational structure, have become a fundamental substrate for label-free entity grouping tasks, including community… 23 arXiv — Machine Learning research 10d ago Cluster-Aware Over-the-Air Federated Learning with Energy-Harvesting Devices: From Global Training to Model Personalization arXiv:2608.01426v1 Announce Type: new Abstract: Federated learning (FL) enables distributed optimization and learning across decentralized edge devices while preserving data privacy, but its performance is fundamentally constrained by heterogeneous data distributions, limited… 4 Page 1 of 10 · 470 articles Older →