News / #hardware Tag Hardware 470 articles archived under #hardware · RSS Sign in to follow r/LocalLLaMA community 1mo ago If you use Open Code or other agenting programs you are leaving a lot of t/s if you don't actually use agents in parallel. Benchmark : RTX5090, Qwen3.6 35B loaded via LM studio with parallel tasks set to 8 As many of you know t/s is super important. It's how fast your stuff gets done. I create via open code benchtest and run it. Thanks to it i know that if i don't run at least 4 agents i basically leave HALF of performance. So whatever you do single project in open code that uses… 34 llama.cpp releases dev-tools 1mo ago b9957 server: improve tools, remove apply_diff ( #25498 ) server: improve tools, remove apply_diff improve edit tool add tools_io abstraction add tools_io_basic fix build move utils to class member add const macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI… 9 llama.cpp releases dev-tools 1mo ago b9949 opencl: cluster-parallel decode FA for Adreno ( #25473 ) macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64… 37 arXiv — Machine Learning research 1mo ago Structure Learning on Clustered Data arXiv:2607.08238v1 Announce Type: new Abstract: Recent algorithmic advances have made directed acyclic graph (DAG) structure learning scalable for causal discovery. Yet, the currently available techniques assume a completely homogeneous population, precluding their application… 6 arXiv — Machine Learning research 1mo ago CASL-VAE: Learning Structured Latent Variables from Unpaired Data for Semi-supervised Clustering and Paired Sample Generation arXiv:2607.08254v1 Announce Type: new Abstract: Quantifying variability in a target population relative to a reference population is central to many scientific and clinical problems (e.g., diseased vs. healthy). Yet, without paired data and in the presence of heterogeneous… 12 arXiv — NLP / Computation & Language research 1mo ago A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embedding arXiv:2607.07974v1 Announce Type: new Abstract: Intent detection is a critical task that bridges human intents and system actions in human-machine interaction systems. However, there still exist challenges for detecting out-of-scope (OOS) intents. (i) The traditional methods… 35 r/LocalLLaMA community 1mo ago Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores TL;DR: Quantization has a marked impact on agentic performance but little effect on knowledge. I manage a small HPC cluster at a university, and we have recently begun running common benchmarks to help our users understand the effects of quantization. We have just completed the… 34 r/LocalLLaMA community 1mo ago Exploring FlashAttention-3/4 optimizations on RTX GPUs I was curious whether any of the FA-3/4 optimizations transfer to RTX GPUs. vLLM/SGLang attention falls back to FA-2 on consumer cards (FA-3 and FA-4 are datacenter-only), so I wanted to know if there's any performance left on the table, and I rebuilt the attention kernels from… 21 Hacker News — AI on Front Page community 1mo ago No leap second will be introduced at the end of December 2026 Article URL: https://datacenter.iers.org/data/latestVersion/bulletinC.txt Comments URL: https://news.ycombinator.com/item?id=48846281 Points: 209 # Comments: 167 18 arXiv — Machine Learning research 1mo ago Converge to Surprise: Evolutionary Self-supervised Image Clustering arXiv:2607.06887v1 Announce Type: new Abstract: Most self-supervised image clustering models, actually almost all deep learning approaches, are based on gradient descent: In order to calculate the loss, every optimization step requires a clearly defined target, whether a… 17 arXiv — Machine Learning research 1mo ago Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery arXiv:2607.06930v1 Announce Type: new Abstract: Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis. Real-world datasets often exhibit complex latent structures composed of multiple subgroups… 18 arXiv — Machine Learning research 1mo ago FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series arXiv:2607.07258v1 Announce Type: new Abstract: In many realistic scenarios, large volumes of time series data are generated with limited or expensive annotations. This limitation makes supervised learning methods difficult to apply and leads to the use of unsupervised… 25 r/LocalLLaMA community 1mo ago What GUI-first coding tool tool are you pairing your local LLMs with? Opencode isn't it for me. I've grown very frustrated with OpenCode. The web GUI and desktop app ideas are good, but the execution not so much. The GUI is lacking so many basic features. It's clear that the TUI is more important to the devs. Is there anything free that provides a more feature-rich GUI?… 14 r/MachineLearning community 1mo ago Why does the same H100 cost 5x more depending on where you rent it? [D] I kept finding wildly different prices for the same GPU across providers and data centers, so I built a OS CLI that searches live GPU capacity and shows the cheapest available routes npx gpu-price-finder Supports RTX 4090, RTX 5090, L40S, A100, H100 and lets you filter by… 20 arXiv — Machine Learning research 1mo ago Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes arXiv:2607.05464v1 Announce Type: new Abstract: The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects. However, most of the existing clustering methods treat the two categorical… 19 arXiv — Machine Learning research 1mo ago Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy arXiv:2607.05469v1 Announce Type: new Abstract: Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks. Although Graph Contrastive Learning has demonstrated promising performance, existing methods often suffer… 30 arXiv — Machine Learning research 1mo ago Modeling Normal Is All You Need: Joint Latent Clustering for Anomaly Detection in Multimodal Cyber-Physical Systems arXiv:2607.06094v1 Announce Type: new Abstract: Faults on a cyber-physical system (CPS) are too rare and unrepresentative to characterise, or even to select a model on, so detection must instead model normal behaviour; the standard point-adjusted evaluation, however, rewards… 26 arXiv — Machine Learning research 1mo ago Performance Optimization and Comparative Analysis of Generative AI Models on Advanced Accelerators arXiv:2607.05400v1 Announce Type: cross Abstract: Generative AI models, such as Large Language Models (LLMs) and diffusion models, have demonstrated impressive performance across a wide range of tasks. Despite these advances, deployment remains challenging due to substantial… 25 Ars Technica — AI news-outlet 1mo ago Data centers’ energy demand threatens Trump’s “Made in America” plan Squeeze on Rust Belt electricity bills threatens Trump’s manufacturing plan. 32 Hugging Face Daily Papers research 1mo ago Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Abstract Temporal aggregation methods for speech-based depression detection show inconsistent performance across different backbones and training runs, highlighting the need for robust benchmarking criteria. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Speech-based depression… 6 arXiv — Machine Learning research 1mo ago Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation arXiv:2607.03007v1 Announce Type: new Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains often come without reliable structural grounding. In particular, existing approaches… 12 arXiv — Machine Learning research 1mo ago Heterogeneous Graph Condensation via Role-Aware Clustering arXiv:2607.03097v1 Announce Type: new Abstract: Heterogeneous Graph Neural Networks (HGNNs) have exhibited remarkable efficacy in modeling complex systems with multiple types of nodes and relations, yet their training on large-scale heterogeneous graphs remains computationally… 37 Hugging Face Daily Papers research 1mo ago Measuring the Gap Between Human and LLM Research Ideas Abstract Large language models generate research ideas that cluster around specific opportunity patterns and paradigms, diverging systematically from the broader and more diverse distributions found in human research papers. Generated by Qwen/Qwen2.5-Coder-32B-Instruct LLMs are… 21 TechCrunch — AI news-outlet 1mo ago Station F ramps up as a launchpad for Europe’s hottest AI startups Station F, a Paris-based startup hub founded by French billionaire Xavier Niel, is gearing up for a new edition of its F/ai accelerator program in a bid to strengthen its positioning as a stepping stone for promising AI startups. 20 r/LocalLLaMA community 1mo ago GLM 5.2 FP8 with FP8 KV - Terminal-Bench 2.1 = 79.8 (with one time-out that I didnt re-run) I wanted to test the official results vs fp8 + fp8 kv. basic sglang setup on H200. If anyone wants one of the official tests do ping me. I didnt rerun the one so it might go up a bit :) TERMINAL-BENCH 2.1 — FINAL RESULTS (mymodel via mini-swe-agent) TOTAL: 89 tasks PASSED: 71… 17 r/LocalLLaMA community 1mo ago Supra Reasoning Summarizer — a tiny model to summarize thinking traces from coding agents Hi, r/LocalLLaMA ! SupraLabs just released a model called: SupraLabs/reasoning-summarizer-800m-pre-gguf It is a thought trace summarizer. Basically, the user/dev can send a reasoning + tool calls (if you want), and then the model will generate a JSON which looks liek this: {… 4 r/LocalLLaMA community 1mo ago Learning to write AI harness old fashioned way. Need help with attention drift and ignoring tool call results! I've been writing a no-compile node.js based AI Harness for llama.cpp as a learning exercise and can really use some help. I'm basing my code off https://github.com/av/mi and https://pi.dev/ with really basic agentic loops. It basically loop until there are no more tool calls… 7 r/LocalLLaMA community 1mo ago [RELEASE] Supra-Router-51M - a tiny prompt routing model/orchestrator Hey r/LocalLLaMA ! SupraLabs is back with a new model. Supra-Router-51M. Basically, this is a model which is made for routing requests to smaller or bigger models - based on the user prompt/request. This is so cool, because it has only 51M parameters and so it can be used in… 34 r/LocalLLaMA community 1mo ago Considering Buying Another RTX 3090 - Benefits? Currently using dual RTX 3090s, and am happy with it. But never satisfied lol :) I know I've basically maxed out my single stream TPS. (140+ on standard benchmarks now). But I only have 48GB VRAM, So I can only do two concurrent requests @ 256k Context Length, anymore and my… 6 r/LocalLLaMA community 1mo ago Best choice of model 40B+ Parameters currently using Qwen3.6 35B as my main assistant model + coding agent but I think sometimes it misses basical general knowledge things, and it is more like executioner that assistant. That's why I though should I go with bigger models, But I don't want to lose speed I am on… 38 Hacker News — AI on Front Page community 1mo ago GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance Article URL: https://github.com/openai/codex/issues/30364 Comments URL: https://news.ycombinator.com/item?id=48789428 Points: 213 # Comments: 75 28 r/LocalLLaMA community 1mo ago Gemma 4 12B - MLX Kernel I've mentioned this kernel project I was working on in a few posts and figured I would just open the project code for anyone curious: MLX Gemma 12B The main constraints for this on my end is an M5 16GB Macbook Pro. I usually do a model development on clusters in the cloud but… 26 arXiv — Machine Learning research 1mo ago Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander arXiv:2607.01736v1 Announce Type: new Abstract: We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choosing the right checkpoint from a world-model training run is difficult: validation loss and… 17 arXiv — NLP / Computation & Language research 1mo ago ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators arXiv:2607.01602v1 Announce Type: new Abstract: SRAM-based FPGAs provide an attractive platform for energy- and latency-constrained CNN inference at the network edge, yet transient faults can lead to silent errors that compromise reliability. Always-on redundancy (e.g., full… 17 arXiv — NLP / Computation & Language research 1mo ago Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts arXiv:2607.01852v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information outside their parameters. We evaluate if cluster-based semantic chunking improves… 29 arXiv — NLP / Computation & Language research 1mo ago Introduction to Transformers: an NLP Perspective arXiv:2311.17633v2 Announce Type: replace Abstract: Transformers have dominated empirical machine learning models of natural language processing. In this paper, we introduce basic concepts of Transformers and present key techniques that form the recent advances of these models.… 22 r/MachineLearning community 1mo ago Improving machine-translated novels via style transfer — looking for advice on the faithfulness/fluency tradeoff [P] Hey all. I recently started working on a project to improve machine-translated webnovels via style transfer. The basic idea is to take the clunky translated prose and rewrite it to something that reads like it was written by a professional author, while remaining as faithful as… 22 r/LocalLLaMA community 1mo ago openlumara, my manually coded super-token-efficient harness, now works across any UI that can connect to an openAI endpoint! koboldlite, openwebui, you name it. basically, openAI bridge. yay! this was a long time coming, but it's finally here! you can now basically supercharge whichever UI you're already using with the power of openlumara . click that link for more information about openlumara itself. TL;DR: super token efficient framework built from the ground up… 25 Ars Technica — AI news-outlet 1mo ago Google’s AI buildout drove 37% increase in electricity use in 2025 Google tries balancing AI data center emissions with clean energy efforts. 34 arXiv — NLP / Computation & Language research 1mo ago Controllable Narrative Rendering for Enhanced Assisted Writing arXiv:2607.00009v1 Announce Type: new Abstract: Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary failure. This issue manifests as an oscillation… 13 arXiv — NLP / Computation & Language research 1mo ago Structural Pattern Mining in Inka Khipus: Unsupervised Clustering, Provenance Classification, and a Computational Validation of the Santa Valley Match arXiv:2607.00185v1 Announce Type: new Abstract: Khipus--knotted cord devices--were the primary recording medium of the Inka Empire (c. 1400-1532 CE), yet their system remains undeciphered. We present a reproducible machine-learning pipeline applied to the Open Khipu Repository… 29 Hugging Face Daily Papers research 1mo ago Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models Abstract Act2Answer protocol evaluates embodied vision-language-action models by having agents answer questions through physical actions, revealing knowledge retention and generalization patterns across different semantic categories. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 35 r/LocalLLaMA community 1mo ago LokalBot - fully local macOS app: meetings, autocomplete, and day tracking that all run on your machine with a user friendly UI Been lurking here a while, this sub is basically why LokalBot exists. It's a Mac app that records + summarizes your meetings, autocompletes your typing in any app, and tracks where your day went, with every model running on-device . No cloud, no account, no API keys. Most of the… 15 arXiv — Machine Learning research 1mo ago TabPATE: Differentially Private Tabular In-Context Learning Without Public Data arXiv:2606.31474v1 Announce Type: new Abstract: Tabular foundation models enable accurate in-context learning (ICL) from small labeled datasets, but the private records placed in context can leak through model predictions. We first show that even basic membership inference… 38 Hacker News — AI on Front Page community 1mo ago County with 37 Data Centers Asks Schools to 'Conserve Electricity' Article URL: https://www.404media.co/henrico-virginia-datacenter-energy-cost-email/ Comments URL: https://news.ycombinator.com/item?id=48734699 Points: 209 # Comments: 105 19 arXiv — Machine Learning research 1mo ago scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering arXiv:2606.28459v1 Announce Type: new Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical noise hinder robust expression representation and cell graph construction.… 27 arXiv — Machine Learning research 1mo ago Improving Patient Subtyping on Longitudinal Data using Representations from Mamba-based Architecture arXiv:2606.28623v1 Announce Type: new Abstract: Effective sub-typing (also known as grouping or clustering) of patients using their electronic health record (EHR) data can greatly inform precision medicine efforts. However, subtyping temporal EHR datasets is known to be… 37 arXiv — Machine Learning research 1mo ago Nonlinear mixture model motivated subspace clustering arXiv:2606.29261v1 Announce Type: new Abstract: We derive the linear union-of-subspaces (UoS) model for subspace clustering (SC) from the nonlinear mixture model (NMM) used in blind source separation (BSS) to represent a D-dimensional observation vector as an unknown… 7 r/LocalLLaMA community 1mo ago Instead of decentralized training effort we should build the “One dataset” There are many threads here calling for united LLM training run of a new open model. Mainly, after govt. stunt of banning commercial frontier models. And also due to the lack of small-medium open-weight models releases lately. I genuinelly believe at some point we’ll have “SETI… 38 Import AI (Jack Clark) community 1mo ago Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era What eras bookend our interregnum? 36 Page 4 of 10 · 470 articles ← Newer Older →