News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow r/MachineLearning community 5h ago Are there any good research papers around Text clustering using LLMs [R] Hi, same as the title, I am currently trying to begin with some research on clustering using LLMs. So my requirement is as follows: I will be given some 100 document files, the end goal is to have clusters in such a way that documents with similar procedures or content should be… 10 r/LocalLLaMA community 1d ago 85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri I've been experimenting with ways to run MoE models that don't fit comfortably in RAM, and I ended up making Overspill, a disk tier for FreeToken . The basic idea came from looking at how Colibri handles experts across disk/RAM/VRAM so I took inspiration from the general… 11 r/MachineLearning community 1d ago LLMs were told they could lie in Diplomacy. Here's who actually kept their promises. [D] the stats are from the game of diplomacy. diplomacy is basically a strategy game where you negotiate, form alliances, betray, and outmaneuver other players to expand your influence. the games were played in multi-agent simulations, with different LLMs playing against each other… 25 TechCrunch — AI news-outlet 2d ago Crusoe abandons $1.25B plan to use Boom turbines at AI data centers Boom Supersonic CEO Blake Scholl said its new stationary power plants were no longer in Crusoe's near-term plans. 35 TechCrunch — AI news-outlet 2d ago Ahead of US IPO, British AI neocloud Nscale secures $3.36B in convertible financing The funding, which comes from Third Point, Nvidia, and others, will fuel the company's massive AI data center buildout. 9 arXiv — Machine Learning research 3d ago The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality arXiv:2609.29530v1 Announce Type: new Abstract: Validating a model on a time series asks for three things at once: each training run should use most of the sample (sufficiency), the test sets should together cover most of the sample (coverage), and training data should come… 5 arXiv — Machine Learning research 3d ago A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes arXiv:2609.29630v1 Announce Type: new Abstract: Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying… 14 arXiv — Machine Learning research 3d ago From Graphs to Feeders: Constraint-Guided Diffusion for Rule-Compliant Feeder Generation arXiv:2609.29879v1 Announce Type: new Abstract: Generative modeling approaches often focus on recovering broad statistical characteristics from the training data. In the context of graph generation, this may refer to degree distributions, clustering coefficients, or spectral… 28 r/LocalLLaMA community 3d ago I'm new and it's kinda overwhelming to get into Hi, sorry if this doesn't belong here. Getting to the point basically, I've been using online-only AI like GPT/Gemini since 2022, and have been interested in local models but am clueless overall. Yes I'm extremely late. I only use laptop (I'm a student), and I currently own:… 38 Ars Technica — AI news-outlet 3d ago New Jersey fines data center $1.1M after drone pics expose 62 gas generators Million-dollar fines won’t end billionaires’ data center pollution, neighbors fear. 14 TechCrunch — AI news-outlet 3d ago Oracle sends force majeure notice on its New Mexico Stargate data center The notice would allow Oracle to delay payments should the facility miss its 2028 target to come online. 13 The Information — AI news-outlet 3d ago Former OpenAI Data Center Chief Is Now At Nvidia Chris Malone, OpenAI’s former head of data centers, joined Nvidia as the vice president of Nvidia’s DSX Platform this month, according to his LinkedIn profile. DSX is the Nvidia division that helps customers design and build AI data centers according to Nvidia’s specifications.… 4 Ars Technica — AI news-outlet 3d ago Google's first Suncatcher orbital data center test launches October 1 Google's experimental orbital data center will have four TPUs and only run for 15 minutes at a time. 21 The Information — AI news-outlet 3d ago Oracle Invokes Force Majeure in Aim to Protect Itself From Data Center Cost Overruns Oracle, which is set to lease a New Mexico data center on behalf of OpenAI, sent a notice to the project’s developer asserting its contractual rights to withhold payments if project delays persist, according to Bloomberg . Oracle sent a notice that cited “force majeure” to a… 11 The Information — AI news-outlet 3d ago Are GPU Loans Safe From Rising Rates? The Federal Reserve’s recent rate hike, and rising expectations of another soon, put a fresh spotlight on financing arrangements for the wide array of companies borrowing money to pay for AI data centers and chips. Rates are rising just as investors have been growing more… 5 r/LocalLLaMA community 3d ago R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB… Here's my *first* implementation of KVA projectors on QFN (just the uncensored model for now) the highlights are basically as follows for using the projectors at each different layer: Starting at layer 12, prompt processing speeds up 1.85x [1700 t/s -> 3150 t/s] at the tradeoff… 24 arXiv — Machine Learning research 4d ago Efficient Linear Bandits via Cluster-Aware Sketching arXiv:2609.27594v1 Announce Type: new Abstract: We study the problem of computational efficiency for linear bandits in high-dimensional settings with a finite arm set. In linear bandits, the increase in the dimension $d$ of the feature vectors leads to growing computational… 4 NVIDIA Developer Blog official-blog 4d ago Validate GPU Cluster Readiness Before AI Workloads Land A GPU cluster can pass every health check and still fail to run an AI workload. Even when every GPU, network link, and pod reports healthy, a 512-GPU training... 28 r/LocalLLaMA community 4d ago MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam MiMo-V2.6-Pro has an insanely high score of 46 on AA, putting it at the head of the opensource models available. It also costs pennies. Flash is not out on AA yet, but it costs less than half on datacenter and is slightly below on Xiaomi's own benchmarks. It also fits in 192GB,… 12 The Information — AI news-outlet 5d ago Anthropic in Talks to Lease Apollo-Controlled Stream Data Centers Anthropic is in early talks to lease up to 1 gigawatt in compute capacity from Stream Data Centers, a developer majority-owned by Apollo Global Management, which could house tensor processing units designed by Broadcom and Google, The Information reported Tuesday. Anthropic… 29 arXiv — Machine Learning research 5d ago Multi-View Fair Clustering Guided by Cross-View Sensitive Information Discrepancy arXiv:2609.25811v1 Announce Type: new Abstract: Multi-view clustering (MVC) aims to uncover latent cluster structures by exploiting complementary information from multiple views. Despite substantial progress in clustering performance, fairness remains an important concern when… 38 arXiv — NLP / Computation & Language research 5d ago MICRO: Multi-Fidelity Active Search for Severe Error Discovery arXiv:2609.26025v1 Announce Type: cross Abstract: Human feedback can vary in cost and informativeness. Strong feedback can reveal severe errors but is costly, so cheaper quality ratings can help decide which items to annotate. We propose MICRO (Multi-Fidelity Impact Clustered… 14 arXiv — Machine Learning research 5d ago Towards Adaptive Federated Graph Clustering: A Global Community-aware Contrastive Learning-based Approach arXiv:2609.26063v1 Announce Type: new Abstract: Federated graph learning (FGL) enables multiple clients to collaboratively train graph models without sharing their private graph data, providing a promising paradigm for mining knowledge from distributed graph repositories. While… 27 arXiv — Machine Learning research 5d ago Mode Collapse Is Cheap to Detect: A Ground-Truth-Free Pre-Flight Check for Neural Samplers arXiv:2609.26272v1 Announce Type: new Abstract: Neural samplers are trained against an unnormalised target $\tilde\pi=e^{-E}$ with no samples from $\pi$, which leaves the practitioner with no way to tell whether an expensive training run has silently dropped part of the target.… 27 arXiv — NLP / Computation & Language research 5d ago ClusterFewshot: Improving Few-shot Optimization for LLMs workflow arXiv:2609.25939v1 Announce Type: new Abstract: The performance of large language model (LLM) workflows often depends on selecting a small set of in-context demonstrations to guide model behavior on new tasks. Recent methods improve this process by augmenting prompts with… 19 The Information — AI news-outlet 5d ago Anthropic in Talks to Cement Control Over More Data Centers Anthropic is in early talks to lease up to 1 gigawatt in compute capacity from a data center developer majority-owned by Apollo Global Management, part of a major effort by the AI company to reduce its reliance on cloud providers. The maker of Claude has discussed signing on as… 16 TechCrunch — AI news-outlet 5d ago Everyone can find a reason to dislike data center construction Inside two years of fraught AI data center debates in Pennsylvania. 18 TechCrunch — AI news-outlet 5d ago Nscale’s IPO will test Wall Street’s appetite for concentrated AI bets once again The British AI data center developer depends on tech giants Microsoft and Anthropic for most of its revenue. 9 r/LocalLLaMA community 6d ago Did Alibaba abandon 35B A3B? Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an announcement later but ngl I kinda lost hope   submitted by   /u/Akainu_Fan [link]   [comments] 8 The Information — AI news-outlet 6d ago Alibaba Unveils New AI Chip And Data Center Expansion Plan Alibaba on Tuesday unveiled a powerful new AI chip for training and running models, highlighting the rapid progress in China’s domestic semiconductor capabilities. During its annual Apsara tech conference, Alibaba said its new AI chip, Zhenwu V900, delivers three times the… 10 arXiv — Machine Learning research 6d ago Clustering-Based Collective Anomaly Detection in IoT Systems: A Graph Neural Network Approach arXiv:2609.22166v1 Announce Type: new Abstract: The rapid advancement of Internet of Things (IoT) technology has led to the widespread deployment of smart, interconnected devices across a range of domains. However, this expansion has also resulted in a substantial increase in… 21 arXiv — NLP / Computation & Language research 6d ago Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments arXiv:2609.22705v1 Announce Type: new Abstract: Online discourse about urban issues - walkability, cycling infrastructure, public transit, housing density, and street safety - is voluminous but unstructured, and existing city-evaluation tools capture none of it. We present a… 24 r/LocalLLaMA community 6d ago What are normie reactions to your local AI use like? There's quite a lot of public anti-AI sentiment at the moment, but a lot of the most common concerns and objections have to do with the politics, economics and ethics of datacenter-based AI. So, when you explain to people that you're running inference locally, what's the… 12 The Information — AI news-outlet 6d ago SB Energy Faces Investor Skepticism in IPO, Report Says SB Energy’s initial public offering is facing investor skepticism, The New York Times reported Monday. The SoftBank majority-owned company is planning to raise $5 billion to $7 billion in a public offering to help develop a massive data center complex OpenAI plans to lease,… 11 arXiv — Machine Learning research 7d ago Multi-Domain Clustering via Measure Quantization arXiv:2609.21664v1 Announce Type: new Abstract: Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples… 5 arXiv — Machine Learning research 7d ago Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data arXiv:2609.21829v1 Announce Type: new Abstract: Clustering high-dimensional data is a fundamental task in unsupervised machine learning with applications to a variety of domains. In the centralized data scenario, this task is commonly solved using deep clustering methods that… 25 arXiv — NLP / Computation & Language research 7d ago Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection arXiv:2609.21179v1 Announce Type: new Abstract: Neural morphological generation systems often achieve high aggregate accuracy on benchmark datasets, yet such performance can conceal systematic errors clustered in rare morphological subclasses. We present an orthography-aware… 10 arXiv — NLP / Computation & Language research 7d ago Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention arXiv:2609.22056v1 Announce Type: cross Abstract: Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure… 29 r/LocalLLaMA community 7d ago Do NOT trust StepFun's Plan subscriptions., They stole >$100 from me with no warning, and I have not heard back from support at all. I know cloud subscription plans are not exactly the core focus of r/LocalLLaMA , so I want to be clear about why I’m posting this here. I’ve been building my own local text-based RPG/game harness and using external models as testing infrastructure: basically stress-testing the… 33 r/LocalLLaMA community 7d ago Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak. ​ I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster. Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but… 6 The Information — AI news-outlet 7d ago How Nvidia Is Trying to Solve the Data Center Power Bottleneck Nvidia is tracking “every single gigawatt of land, power and shell around the world, literally everything on the planet,” CEO Jensen Huang said at Goldman Sachs’ annual tech conference earlier this month. “We know where everything is.” Why track power so intently? Nvidia sees… 33 r/LocalLLaMA community 7d ago I tested 9 LLMs on the exact same web-dev prompt for ~8 hours — RTX 3060 12GB results (Rate the best!) I’ve spent basically the last 8 hours testing different models on the exact same web-development prompt, and I finally finished. The whole point of this nine-hour test was that which local model matches the frontier-level intelligence at size and could fit easily in an RTX… 32 r/MachineLearning community 7d ago Autograd project [P] Hello, I'm a 3rd year Highschooler interested in machine learning and for the last few weeks have been working on a small project meant to learn the basics of machine learning. I have implemented a simple tensor library and autograd in c++. It's very simple but i want some… 31 The Information — AI news-outlet 7d ago Jane Street-Linked Data Center Debt Sours Secretive Wall Street trading firm Jane Street has become famous in the past couple of years for its sizable AI-related spending, as it has struck big cloud deals to handle its AI computing needs and made investments in neocloud CoreWeave. Yet even a Jane Street–related… 32 The Information — AI news-outlet 9d ago Nscale IPO Files to Go Public, Shows Huge Revenue Jump And Steep Losses Nvidia-backed startup Nscale disclosed Friday that revenue had surged 10 times in the first half of this year from the year-ago period, but losses mounted as the cloud startup spent billions on data centers and the AI chips they’ll house. The two-year-old UK startup, a spinout… 12 r/MachineLearning community 9d ago I posted my embedding migration project here, it got a lot of attention, so I added the features you guys said were missing [R] A little while ago, I posted about embedflow https://github.com/arnsri33/embedflow and it got a lot of attention. The basic idea was pretty simple keep your existing embedding index for candidate retrieval → rerank "k" candidates with the new embedding model → progressively… 6 arXiv — Machine Learning research 10d ago Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices arXiv:2609.19243v1 Announce Type: new Abstract: Spectral co-clustering is a useful tool for discovering latent structure in word-document matrices, but its reliance on singular value decomposition (SVD) can make standard formulations expensive on high-dimensional data. This… 5 arXiv — Machine Learning research 10d ago COMPASS: Ordered Clustered Routing at 100K Scale arXiv:2609.20352v1 Announce Type: new Abstract: Large-scale routing often requires visiting clusters of nodes in a prescribed order, giving rise to the Ordered Clustered Traveling Salesman Problem (OCTSP). Optimizing each cluster independently seems natural, but misses non-local… 20 r/LocalLLaMA community 10d ago dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode Note not my work, but something i found and wanted to share so hopefully more people can push this along even further. https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt I basically run 2 x 7900 xtx on a consumer pc. This repo takes qwen 3.8 Q8 and optimizes it to run on… 5 TechCrunch — AI news-outlet 10d ago Crusoe raises $3.9B to build massive data centers and small modular “AI factories” The round values the data center giant at $30.9 billion. 14 Page 1 of 10 · 500 articles Older →