arXiv — Machine Learning
500 articles archived · Visit source ↗ · RSS
-
arXiv — Machine Learning research 4d ago
Sharding Prevents LLM Oversight Failures and Adversarial Exploitation
arXiv:2608.06422v1 Announce Type: new Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or…
13 -
arXiv — Machine Learning research 4d ago
Adversarial Causal Intervention Falsification
arXiv:2608.06427v1 Announce Type: new Abstract: Generative models can reproduce an observational distribution while encoding an incorrect causal structure. We study a sequential game in which a structural causal generator proposes observational and interventional distributions,…
27 -
arXiv — Machine Learning research 4d ago
Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces
arXiv:2608.06428v1 Announce Type: new Abstract: Deep Operator Networks (DeepONets; arXiv:1910.03193) typically encode an input function through point values on a fixed discretization. Building on the Topological DeepONet framework of Ismailov (arXiv:2603.11972), we replace point…
10 -
arXiv — Machine Learning research 4d ago
MiGHT-EHR: A Multi-task Graph Transformer for Heterogeneous Temporal Electronic Health Records
arXiv:2608.06430v1 Announce Type: new Abstract: Learning from Electronic Health Records (EHRs) has gained significant attention due to its potential to improve clinical prediction. However, effective learning remains challenging because EHRs encode heterogeneous, temporally…
8 -
arXiv — Machine Learning research 4d ago
SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction
arXiv:2608.06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We present SNI-GNN, a SmartNIC-assisted full-graph training system that reduces…
21 -
arXiv — Machine Learning research 4d ago
ED-CSP: Crystal Structure Prediction from Electron Diffraction
arXiv:2608.06448v1 Announce Type: new Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels,…
36 -
arXiv — Machine Learning research 4d ago
Beyond Attention: Signed Integrated Gradients Attribution in a BiomeGPT-Style Microbiome Transformer
arXiv:2608.06486v1 Announce Type: new Abstract: In a feature-tokenized transformer (arXiv:2106.11959) such as BiomeGPT (doi:10.64898/2026.01.05.697599), each input token is built by fusing a fixed identity with a sample-specific measurement: a fixed species and a variable…
13 -
arXiv — Machine Learning research 4d ago
Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability
arXiv:2608.06503v1 Announce Type: new Abstract: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood. In this preliminary empirical study, we show that compression can weaken the influence of recent…
19 -
arXiv — Machine Learning research 4d ago
Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning
arXiv:2608.06511v1 Announce Type: new Abstract: Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly…
35 -
arXiv — Machine Learning research 4d ago
Target-Weighted Neyman Allocation: Experimental Design for Heterogeneous Treatment Effects under Population Shift
arXiv:2608.06512v1 Announce Type: new Abstract: Randomized experiments are often run in one population to guide decisions in another. Allocating by experimental proportions wastes budget on groups that rarely appear in deployment, whereas allocating by deployment proportions…
24 -
arXiv — Machine Learning research 4d ago
CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions
arXiv:2608.06516v1 Announce Type: new Abstract: Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the level of task decisions. A connected route can expand cross-modal reach while changing an…
18 -
arXiv — Machine Learning research 4d ago
Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks
arXiv:2608.06520v1 Announce Type: new Abstract: We study online cooperative control of a multi-agent system under Byzantine attacks. Namely, an unknown, fixed subset of agents are Byzantine comprised and can stealthily overwrite its own coordinates of the team's planned joint…
33 -
arXiv — Machine Learning research 4d ago
Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions
arXiv:2608.06545v1 Announce Type: new Abstract: Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $\varepsilon$-optimal…
12 -
arXiv — Machine Learning research 4d ago
Newton-Schulz Retraction-Based Inference Enables Hidden Quantum Markov Models to Outperform Classical HMMs
arXiv:2608.06554v1 Announce Type: new Abstract: Hidden Markov models (HMMs) are widely used probabilistic models for discrete sequential data but can be limited when hidden dynamics are complex. Hidden quantum Markov models (HQMMs) generalize HMMs by replacing probability…
31 -
arXiv — Machine Learning research 4d ago
Bootstrap-Conditioned Action Selection with Tabular Foundation Models
arXiv:2608.06559v1 Announce Type: new Abstract: Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates, and severe cold starts. We study…
14 -
arXiv — Machine Learning research 4d ago
Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization
arXiv:2608.06563v1 Announce Type: new Abstract: Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications. Modern large-scale training relies on classical optimization…
11 -
arXiv — Machine Learning research 4d ago
Quantization Damage Is Multiplicative, Not Additive
arXiv:2608.06564v1 Announce Type: new Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's decisions will change at a given bit-width. The damage is silent: a compressed…
38 -
arXiv — Machine Learning research 4d ago
CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure Prediction
arXiv:2608.06582v1 Announce Type: new Abstract: Flow-based generative models can efficiently produce candidate structures for crystal structure prediction (CSP), but their pretrained objectives do not directly optimize downstream target recovery. Reinforcement-learning…
33 -
arXiv — Machine Learning research 4d ago
Flowing Through States: Neural ODE Regularization for Reinforcement Learning
arXiv:2608.06595v1 Announce Type: new Abstract: Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. While environment dynamics dictate how semantic states evolve, the corresponding latent transitions are…
38 -
arXiv — Machine Learning research 4d ago
Retrofitting Linear Attention into Diffusion Language Models
arXiv:2608.06628v1 Announce Type: new Abstract: Diffusion language models (dLLMs) offer a promising alternative to autoregressive models by accelerating inference through parallel decoding. Recent dLLMs commonly use blockwise semi-autoregressive decoding, generating blocks…
37 -
arXiv — Machine Learning research 4d ago
The Sparsity Whisperer
arXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly…
13 -
arXiv — Machine Learning research 4d ago
Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation
arXiv:2608.06631v1 Announce Type: new Abstract: Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and for a Transformer's final projection matrix. These methods do not recover the bias-free Gated…
25 -
arXiv — Machine Learning research 4d ago
Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning
arXiv:2608.06637v1 Announce Type: new Abstract: Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggregation rules, such as Krum and Multi-Krum, select updates that are closest to the majority…
10 -
arXiv — Machine Learning research 4d ago
Dirichlet Follow-the-Leader Closes the Gap in Simultaneous Multiclass U-Calibration
arXiv:2608.06656v1 Announce Type: new Abstract: Can one forecaster attain the optimal regret rate for every bounded proper loss and also adapt to every smooth proper loss? Recent work answered this up to a dimension gap. Its self-concordant perturbation gives roughly…
37 -
arXiv — Machine Learning research 4d ago
EpiFlow: A framework for improving the utility of wastewater signals for disease forecasting
arXiv:2608.06671v1 Announce Type: new Abstract: Wastewater-based surveillance is an effective tool for disease monitoring and can provide early warning of outbreaks. Although wastewater viral loads (WVL) correlate with disease burden, their utility for improving real-time…
38 -
arXiv — Machine Learning research 4d ago
A Transferable Autologistic Model for Predicting Rare Failures in Heterogeneous Equipment
arXiv:2608.06695v1 Announce Type: new Abstract: Predicting failures before they occur remains a major challenge in predictive maintenance, particularly when failures are rare, when equipment of the same family differ in sensor configurations, and when the goal is anticipation…
23 -
arXiv — Machine Learning research 4d ago
Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection
arXiv:2608.06706v1 Announce Type: new Abstract: Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go action-blind: predictions for different actions become indistinguishable even as the…
19 -
-
arXiv — Machine Learning research 4d ago
Solver-Guided Reasoning for Mixed-Equilibrium Strategies
arXiv:2608.06741v1 Announce Type: new Abstract: Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. In fact,…
22 -
arXiv — Machine Learning research 4d ago
KReF: Training-Free Retrieval for Long-Term Time-Series Forecasting and Predictive Uncertainty
arXiv:2608.06748v1 Announce Type: new Abstract: Probabilistic long-term time-series forecasting commonly relies on trained models. Training-free conformal methods typically construct intervals around a pre-existing point forecaster and do not natively represent a complete…
8 -
-
arXiv — Machine Learning research 4d ago
CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights
arXiv:2608.06763v1 Announce Type: new Abstract: Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU execution. Uniform integers constrain each group to a linear grid. Low-bit…
9 -
arXiv — Machine Learning research 4d ago
Hidden Gauge Controls Feature Specialization in ReLU Networks
arXiv:2608.06766v1 Announce Type: new Abstract: Training changes a network's predictions while allocating task-relevant structure across its internal units. In an overparameterized ReLU network, several neurons can begin with exactly the same functional role, yet one may acquire…
12 -
arXiv — Machine Learning research 4d ago
ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling
arXiv:2608.06772v1 Announce Type: new Abstract: Accurate estimation of building energy use is essential for achieving carbon neutral and sustainable buildings. To better understand the influence of design decisions on building energy use and calibrate machine learning models…
38 -
arXiv — Machine Learning research 4d ago
Faster Query-Key Learning Sharpens Attention in Self-Attention Models
arXiv:2608.06776v1 Announce Type: new Abstract: A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended representations to predictions. Collapsed and factorized…
8 -
arXiv — Machine Learning research 4d ago
Understanding Differentiable Embeddings Through Differential and Integral Geometry
arXiv:2608.06809v1 Announce Type: new Abstract: How can an analyst decide whether a nonlinear dimensionality reduction embedding can be trusted? Existing diagnostics provide only partial answers: projection glyphs characterize local sensitivity, map-continuity scores measure…
34 -
arXiv — Machine Learning research 4d ago
Multiscale Reward Hedging from Correct Demonstrations
arXiv:2608.06825v1 Announce Type: new Abstract: Learning from correct demonstrations is harder than supervised learning when many answers are correct: after predicting, the learner sees one valid answer but not whether its own answer was valid, nor any reward. Existing…
6 -
arXiv — Machine Learning research 4d ago
Graph Machine: Exploring Edge Mechanisms as an Inductive Bias
arXiv:2608.06834v1 Announce Type: new Abstract: Transformers provide a powerful architecture for global content-based matching, but reasoning problems may benefit from a stronger inductive bias toward iterative traversal of latent relations. We introduce Graph Machine, an…
38 -
arXiv — Machine Learning research 4d ago
Mathematical Principles and Experimental Discoveries of the Emergence of Symbolic Patterns in Artificial Neural Networks
arXiv:2608.06839v1 Announce Type: new Abstract: Artificial Neural networks (ANNs) are often treated as black-box models, making explainability a central challenge in deep learning. Many engineering methods have been proposed to approximately explain the ANN from various…
19 -
arXiv — Machine Learning research 4d ago
Bridging the Gap Between Hyperdimensional Computing and Kernel Methods via the Nystr\"om Method
arXiv:2608.06860v1 Announce Type: new Abstract: Hyperdimensional computing (HDC) is an approach from the cognitive science literature for solving information processing tasks using data represented as high-dimensional random vectors. The technique has a rigorous mathematical…
28 -
arXiv — Machine Learning research 4d ago
SkillAligner: Treating Retrieved Skills as Adaptable Drafts at Execution Time
arXiv:2608.06880v1 Announce Type: new Abstract: General-purpose skills promise reusable procedural knowledge for language agents, yet semantic relevance does not guarantee execution utility: a retrieved skill may encode assumptions that conflict with the current task, execution…
23 -
arXiv — Machine Learning research 4d ago
PRISM: Principled Reference Identification for Schrodinger Bridge Model
arXiv:2608.06893v1 Announce Type: new Abstract: Schr\"odinger bridge models restore a clean signal from a degraded observation by following the conditional bridges of a reference process, yet this reference is chosen heuristically, typically white noise with a hand-tuned…
29 -
-
arXiv — Machine Learning research 4d ago
MiCoPro: End-to-End Mixed Precision HW/SW Co-design with HW-aware Proxy Model
arXiv:2608.06916v1 Announce Type: new Abstract: Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision…
31 -
arXiv — Machine Learning research 4d ago
Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
arXiv:2608.06934v1 Announce Type: new Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences. Existing studies, however, often reduce these diverse judgements to…
17 -
arXiv — Machine Learning research 4d ago
ELMZip: Onboard Satellite Image Compression via Extreme Learning Machines for Efficient Downlink
arXiv:2608.06942v1 Announce Type: new Abstract: The acquisition of multispectral imagery via small satellites (e.g., CubeSats) presents significant data downlink challenges due to high data volumes and restricted communication windows. While onboard image compression is critical…
27 -
arXiv — Machine Learning research 4d ago
A Rate Separation for Agnostic Direct Sums
arXiv:2608.06951v1 Announce Type: new Abstract: Hanneke, Moran, and Waknine \cite{HannekeMoranWaknine2024} asked how the agnostic PAC learning curve of the direct sum $C^r$ depends on the single-instance learning curve $\epsagn(n\mid C)$ and on $r$. We show that the…
30 -
arXiv — Machine Learning research 4d ago
How Molecular Generative Models Organize Molecular Identity
arXiv:2608.06956v1 Announce Type: new Abstract: Generative models for matter are often evaluated as samplers over output representations, and their latent spaces are commonly used as proxies for navigating chemical space. Much less is known about how these models internally…
6 -
arXiv — Machine Learning research 4d ago
Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs
arXiv:2608.06990v1 Announce Type: new Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. Among various clustering methods, hierarchical clustering, density-based clustering, and graph clustering stand out as…
10 -