arXiv — Machine Learning
500 articles archived · Visit source ↗ · RSS
-
arXiv — Machine Learning research 2d ago
CurveFP: Rational-Radix Logarithmic Datatypes with Closed Products for Language Models
arXiv:2608.10010v1 Announce Type: new Abstract: Low-precision datatypes reduce language-model cost, but most formats optimize scalar fidelity while leaving the arithmetic induced by their products unchanged. We introduce CurveFP, a closed-product codebook family that distributes…
22 -
arXiv — Machine Learning research 2d ago
Sheaf-Based Federated Representation Learning
arXiv:2608.10016v1 Announce Type: new Abstract: Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model architectures, latent dimensionalities, and local learning…
6 -
arXiv — Machine Learning research 2d ago
DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents
arXiv:2608.10037v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the…
38 -
arXiv — Machine Learning research 2d ago
FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows
arXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures. However,…
20 -
arXiv — Machine Learning research 2d ago
UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs
arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark…
29 -
arXiv — Machine Learning research 2d ago
Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons
arXiv:2608.10045v1 Announce Type: new Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. In this problem, the goal is to…
7 -
arXiv — Machine Learning research 2d ago
Detecting Soft Skills in ML Engineering Roles CVs
arXiv:2608.10046v1 Announce Type: new Abstract: Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. Job advertisements, surveys, and…
12 -
arXiv — Machine Learning research 2d ago
Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review
arXiv:2608.10047v1 Announce Type: new Abstract: In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learning (ML) has largely driven advancements in diagnostics and prognostics, yet purely…
10 -
arXiv — Machine Learning research 2d ago
Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs
arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as…
23 -
arXiv — Machine Learning research 2d ago
ChronoSSM: Training for Temporally Aware Representations in Autoregressive State Space Models
arXiv:2608.10120v1 Announce Type: new Abstract: Modern sequence models, from Transformers to State Space Models, have enabled powerful generative modeling across diverse domains, yet they are typically trained to predict what happens while treating when it happens as a secondary…
13 -
arXiv — Machine Learning research 2d ago
Procedural Fairness Failures in RLHF from Preference Averaging
arXiv:2608.10126v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural fairness…
13 -
arXiv — Machine Learning research 2d ago
SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks
arXiv:2608.10144v1 Announce Type: new Abstract: We consider federated parameter efficient fine-tuning of large neural networks with low-rank adaptation (LoRA,~Hu et al.\ 2022). Combining LoRA with federated PEFT introduces challenges absent from either setting alone: clients may…
22 -
arXiv — Machine Learning research 2d ago
The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom
arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. We reproduce that result by…
18 -
arXiv — Machine Learning research 2d ago
REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting
arXiv:2608.10149v1 Announce Type: new Abstract: Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed…
38 -
arXiv — Machine Learning research 2d ago
Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability
arXiv:2608.10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the…
28 -
arXiv — Machine Learning research 2d ago
From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation
arXiv:2608.10182v1 Announce Type: new Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When the business goal is incremental impact, as in marketing campaigns, incentives, and…
9 -
arXiv — Machine Learning research 2d ago
ELMER: Evolutionary Language Model that Explores and Refines
arXiv:2608.10196v1 Announce Type: new Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a…
7 -
arXiv — Machine Learning research 2d ago
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning
arXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies that whenever the constraint is active at…
32 -
arXiv — Machine Learning research 2d ago
A matched-integrator evaluation of Hamiltonian neural networks on pendulum and Kepler dynamics
arXiv:2608.10235v1 Announce Type: new Abstract: Hamiltonian Neural Networks (HNNs) parameterize conservative dynamics through a learned scalar Hamiltonian, providing an architectural prior that is absent from generic vector-field neural networks. We evaluate this prior under a…
34 -
arXiv — Machine Learning research 2d ago
STCAD: Scalable Trajectory Clustering and Anomaly Detection on Terabyte-Scale AIS Data
arXiv:2608.10249v1 Announce Type: new Abstract: We present a scalable framework for unsupervised clustering of maritime trajectories derived from terabyte-scale Automatic Identification System (AIS) archives. Variable-length trajectories are encoded with a custom BERT-based…
29 -
arXiv — Machine Learning research 2d ago
CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling
arXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and navigational realism. We propose the Continuous Regression Hybrid Transformer…
27 -
arXiv — Machine Learning research 2d ago
Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this…
14 -
arXiv — Machine Learning research 2d ago
Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
arXiv:2608.10288v1 Announce Type: new Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated…
28 -
arXiv — Machine Learning research 2d ago
MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale
arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods…
34 -
arXiv — Machine Learning research 2d ago
Accelerated Learning of High Dimensional Functions with a Tensor-Featured Training Network
arXiv:2608.10351v1 Announce Type: new Abstract: In this work we present a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN). This optimization procedure introduces contextual features into the first layer of a DNN. The…
4 -
arXiv — Machine Learning research 2d ago
Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks
arXiv:2608.10357v1 Announce Type: new Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy…
20 -
arXiv — Machine Learning research 2d ago
Invertible Logits Transformation for Accuracy-Preserving Post-Hoc Uncertainty Calibration
arXiv:2608.10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining. An ideal calibrator should correct nonlinear miscalibration, scale gracefully to large label spaces, and preserve the…
25 -
arXiv — Machine Learning research 2d ago
Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher Geometry
arXiv:2608.10374v1 Announce Type: new Abstract: Training neural networks to jointly predict mean and uncertainty estimates from noisy observations can be unstable, prompting a series of independent stabilization efforts. We argue that these interventions highlight a common…
16 -
arXiv — Machine Learning research 2d ago
Generator-Guided Inverse Sampling for L\'evy-Driven Generative Models
arXiv:2608.10384v1 Announce Type: new Abstract: This paper studies inverse sampling for L\'evy-driven generative models from the perspective of Markov generators. Unlike conventional diffusion models, L\'evy-driven dynamics involve infinite jump activities, which makes their…
19 -
arXiv — Machine Learning research 2d ago
Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving
arXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world models reduce the reliance on costly environment interactions, policy optimization…
37 -
arXiv — Machine Learning research 2d ago
Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation
arXiv:2608.10392v1 Announce Type: new Abstract: Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared-expert designs preserve reusable knowledge, fine-grained methods vary computation within experts, and dynamic routers…
13 -
arXiv — Machine Learning research 2d ago
ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation
arXiv:2608.10398v1 Announce Type: new Abstract: Variational autoencoders generate samples from probabilistic latent representations but do not distinguish uncertainty about the latent location from variability around it. We formulate ELVAE, an evidential learning-based VAE in…
13 -
arXiv — Machine Learning research 2d ago
Do Judges Behave Like Algorithms?
arXiv:2608.10400v1 Announce Type: new Abstract: What if judges already behave like algorithms? As artificial intelligence and algorithms are deployed in many settings, including the judicial system, many have debated whether judges should be allowed to rely on them. Instead, we…
4 -
arXiv — Machine Learning research 2d ago
TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling
arXiv:2608.10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times.…
16 -
arXiv — Machine Learning research 2d ago
Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique
arXiv:2608.10430v1 Announce Type: new Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. Existing detection…
14 -
arXiv — Machine Learning research 2d ago
Do Time-Series Forecasters Use the Right History: Recoverability, Recovery, and Functional Use of Temporal Delays
arXiv:2608.10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structure: can the true delay be recovered from the observed data, does the model…
37 -
arXiv — Machine Learning research 2d ago
Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents
arXiv:2608.10441v1 Announce Type: new Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using.…
30 -
arXiv — Machine Learning research 2d ago
A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes
arXiv:2608.10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute $S$ requires a representation $Z$ that is statistically independent of $S$. Existing criteria, including generalized demographic parity, the expectation of integral…
37 -
arXiv — Machine Learning research 2d ago
Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning
arXiv:2608.10473v1 Announce Type: new Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online…
19 -
arXiv — Machine Learning research 2d ago
Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation
arXiv:2608.10499v1 Announce Type: new Abstract: Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's…
8 -
arXiv — Machine Learning research 2d ago
Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits
arXiv:2608.10526v1 Announce Type: new Abstract: Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. We consider three information structures: (A)~unobserved actions with…
31 -
arXiv — Machine Learning research 2d ago
Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry
arXiv:2608.10529v1 Announce Type: new Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. However, real-world applications often involve heavy-tailed reward distributions and…
6 -
arXiv — Machine Learning research 2d ago
Retrieval-Corrected Conformal Prediction for Time Series
arXiv:2608.10553v1 Announce Type: new Abstract: Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and…
12 -
arXiv — Machine Learning research 2d ago
MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction
arXiv:2608.10562v1 Announce Type: new Abstract: Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), yet treats every click as the same event. In reality, users provide a free,…
33 -
arXiv — Machine Learning research 2d ago
BREAD: Baseline-Referenced Explanations for Anomaly Diagnosis
arXiv:2608.10587v1 Announce Type: new Abstract: Artificial Intelligence (AI)-based prospective anomaly detection methods are increasingly deployed in high-dimensional and nonlinear settings. Among these approaches, AI-based statistical process monitoring (SPM) is widely used,…
20 -
arXiv — Machine Learning research 2d ago
$\beta$-VAEs as Effective Theories: Tolerance-Dependent Dimension
arXiv:2608.10599v1 Announce Type: new Abstract: In a $\beta$-VAE, increasing the regularization strength acts as a spectral cutoff by collapsing low-utility latent coordinates. In the linear Gaussian VAE, the collapse order matches the ranking of reconstruction utilities…
11 -
arXiv — Machine Learning research 2d ago
Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts
arXiv:2608.10605v1 Announce Type: new Abstract: In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and training recipe, optimizing loss under compute…
38 -
arXiv — Machine Learning research 2d ago
Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment
arXiv:2608.10619v1 Announce Type: new Abstract: Message-passing neural networks (MPNNs) often struggle when task-relevant information is distributed across distant regions of a graph, since local propagation must compress remote signals through limited structural interfaces.…
34 -
arXiv — Machine Learning research 2d ago
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete…
37 -
arXiv — Machine Learning research 2d ago
IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning
arXiv:2608.10634v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making. Numerous methods have been developed to improve dynamics…
13