arXiv — Machine Learning
500 articles archived · Visit source ↗ · RSS
-
arXiv — Machine Learning research 1d ago
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
arXiv:2608.11698v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond…
8 -
arXiv — Machine Learning research 1d ago
Consolidator: Learning Persistent Routed Memory Across Context Boundaries
arXiv:2608.11701v1 Announce Type: new Abstract: Copying short-term memory (STM) into a slower store can preserve state across a context boundary, but persistence alone does not ensure that the retained state influences subsequent memory access. We test this distinction in a…
4 -
-
arXiv — Machine Learning research 1d ago
High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions
arXiv:2608.11713v1 Announce Type: new Abstract: Multi-objective Bayesian optimization (MOBO) is effective in identifying the Pareto fronts for expensive black-box problems. However, most current MOBO approaches are limited to low-dimensional decision space due to its exponential…
35 -
arXiv — Machine Learning research 1d ago
Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity
arXiv:2608.11716v1 Announce Type: new Abstract: Chain of Thought (CoT) lifts the expressive ceiling of bounded-depth Transformers, with characterizations tying the number of CoT steps to circuit complexity classes. What remains largely missing are concrete instantiations with…
6 -
arXiv — Machine Learning research 1d ago
Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization
arXiv:2608.11746v1 Announce Type: new Abstract: Modern systems are increasingly expected to transfer across tasks not specified during training. What data facilitates generalization in these new, unanticipated settings? One hypothesis is that data with more structural…
35 -
arXiv — Machine Learning research 1d ago
MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
arXiv:2608.11749v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and…
8 -
arXiv — Machine Learning research 1d ago
TradingMoE: Routing the Right Experts in Evolving Markets
arXiv:2608.11785v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary across assets, decision fields, and market…
36 -
arXiv — Machine Learning research 1d ago
High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving
arXiv:2608.11790v1 Announce Type: new Abstract: Accurate Global Navigation Satellite System (GNSS)-based localization is essential for safe and reliable autonomous driving. However, spoofing attacks can manipulate vehicle position estimates. Continuous and subtle attacks are…
5 -
arXiv — Machine Learning research 1d ago
Orientation, not magnitude: the causal structure of task-vector interference in merged language models
arXiv:2608.11797v1 Announce Type: new Abstract: Model merging by task arithmetic works until it doesn't, and the field diagnoses why with magnitudes: layerwise representation bias, deviations from cross-task linearity, parameter overlap. Tracking the exact layerwise cross-term…
9 -
arXiv — Machine Learning research 1d ago
JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series
arXiv:2608.11801v1 Announce Type: new Abstract: Multivariate time-series anomaly prediction aims to identify whether and when anomalies will occur over a future horizon from historical observations. Existing methods primarily characterize anomalies as deviations in future…
18 -
arXiv — Machine Learning research 1d ago
Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks
arXiv:2608.11815v1 Announce Type: new Abstract: Transfer-based adversarial attacks craft adversarial examples using surrogate models to mislead black-box victim models. Beyond perturbation generation, transferability is fundamentally governed by the coupling of initialization,…
18 -
arXiv — Machine Learning research 1d ago
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
arXiv:2608.11829v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledge from a stronger teacher model, thereby expanding…
29 -
arXiv — Machine Learning research 1d ago
Kernel Methods for Learning Operators with Multiple Inputs and Outputs
arXiv:2608.11831v1 Announce Type: new Abstract: Learning mappings between infinite-dimensional objects is a central challenge in scientific machine learning. We introduce a general kernel-based encoder-decoder framework for operator learning that separates observation,…
12 -
arXiv — Machine Learning research 1d ago
Air Quality Station Simulation via LSTM and Attention-Based Modelling
arXiv:2608.11839v1 Announce Type: new Abstract: Poor air quality in urban areas is driven by a complex chain of processes and presents a significant public health concern. To better understand and control the mechanisms that determine air quality, cities deploy networks of…
24 -
arXiv — Machine Learning research 1d ago
Small-Scale Experiments: Are We There Yet?
arXiv:2608.11859v1 Announce Type: new Abstract: Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot…
28 -
-
arXiv — Machine Learning research 1d ago
DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks
arXiv:2608.11873v1 Announce Type: new Abstract: In this work, we extend the Dependent Click Model (DCM) Bandits to a multiplayer information-asymmetric setting, where multiple agents interact with a shared ranked list and may observe multiple clicks per session, introducing new…
30 -
arXiv — Machine Learning research 1d ago
Disentangling the Expressivity of RoPE
arXiv:2608.11909v1 Announce Type: new Abstract: Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize…
16 -
arXiv — Machine Learning research 1d ago
A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression
arXiv:2608.11917v1 Announce Type: new Abstract: Multi-output Gaussian process regression scales cubically in the number of observations times outputs, and dense kernel-matrix methods need bespoke handling whenever different outputs are observed at different inputs. We express…
36 -
arXiv — Machine Learning research 1d ago
Distillation of Foundation Models for Time-dependent PDEs
arXiv:2608.11937v1 Announce Type: new Abstract: Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively to new downstream tasks. After fine-tuning on only a few…
25 -
arXiv — Machine Learning research 1d ago
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
arXiv:2608.11951v1 Announce Type: new Abstract: Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. Such events are rare in historical records,…
38 -
arXiv — Machine Learning research 1d ago
LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
arXiv:2608.11967v1 Announce Type: new Abstract: Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress,…
31 -
arXiv — Machine Learning research 1d ago
TESLA: Taylor Expansion of Sinusoidal Learnable Activations
arXiv:2608.11970v1 Announce Type: new Abstract: The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an…
15 -
arXiv — Machine Learning research 1d ago
Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh
arXiv:2608.12001v1 Announce Type: new Abstract: Rapid urbanization in Dhaka District, Bangladesh has triggered substantial alterations in land use and environmental conditions, necessitating systematic monitoring for informed urban planning and ecological sustainability. This…
27 -
-
arXiv — Machine Learning research 1d ago
Reducing Symmetry Increase in Equivariant Neural Networks
arXiv:2608.12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric…
11 -
arXiv — Machine Learning research 1d ago
SoftWater: Class-Aware Rate Allocation for Softmax Quantization
arXiv:2608.12026v1 Announce Type: new Abstract: Post-training quantization pipelines routinely leave the softmax output layer in high precision. Yet in small LLMs with modern vocabularies, the head holds 15--30\% of all parameters, so a nominal ``2-bit'' model with an fp16 head…
27 -
arXiv — Machine Learning research 1d ago
Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision
arXiv:2608.12027v1 Announce Type: new Abstract: Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing…
13 -
arXiv — Machine Learning research 1d ago
Clustered Randomized Smoothing for Stochastic Prediction Functions
arXiv:2608.12037v1 Announce Type: new Abstract: Modern stochastic predictors can model rich, multi-modal outcome distributions. However, this expressive power comes with challenges in ensuring robust predictions $-$ a critical requirement in safety-critical domains. Randomized…
8 -
arXiv — Machine Learning research 1d ago
Towards Truly Unsupervised Evaluation of Feature Selection
arXiv:2608.12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method. Most of the methods…
9 -
arXiv — Machine Learning research 1d ago
Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion
arXiv:2608.12083v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic explanation of their predictions. This…
23 -
arXiv — Machine Learning research 1d ago
NAE: Normalizing AutoEncoder
arXiv:2608.12084v1 Announce Type: new Abstract: We consider the setting of Normalizing flows with approximate inverses, an established paradigm spanning both full-dimensional ($d=D$) and bottleneck ($d<D$) settings, and group these models under the term flow autoencoders. We…
9 -
arXiv — Machine Learning research 1d ago
Task- and dataset-specific information in protein language models
arXiv:2608.12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid…
17 -
arXiv — Machine Learning research 1d ago
Confidence Calibration of Deep Learning Systems
arXiv:2608.12100v1 Announce Type: new Abstract: In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. Confidence calibration ensures that predicted probabilities reflect the likelihood of correctness, making it essential for…
36 -
arXiv — Machine Learning research 1d ago
Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning
arXiv:2608.12108v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative model training across distributed clients while keeping data local. A central challenge is determining which client updates are beneficial for aggregation with respect to each client's…
6 -
arXiv — Machine Learning research 1d ago
Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification
arXiv:2608.12117v1 Announce Type: new Abstract: Arterial pulse waveform morphology evolves with age, reflecting structural and functional changes in the cardiovascular system. Thus, vascular age is a valuable surrogate marker of cardiovascular health, and premature vascular…
22 -
-
arXiv — Machine Learning research 1d ago
HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
arXiv:2608.12194v1 Announce Type: new Abstract: Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions. However, assigning an independent function to every connection results in substantial…
4 -
arXiv — Machine Learning research 1d ago
ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening
arXiv:2608.12219v1 Announce Type: new Abstract: Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive,…
6 -
arXiv — Machine Learning research 1d ago
An Efficient Near-Optimal Algorithm for Adversarial $m$-Set Bandits
arXiv:2608.12231v1 Announce Type: new Abstract: We study adversarial combinatorial bandits with $m$-set actions, where at each round the learner selects $m$ out of $d$ items and observes only the aggregate loss of the selected items. The resulting action set contains…
31 -
arXiv — Machine Learning research 1d ago
Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting
arXiv:2608.12259v1 Announce Type: new Abstract: Financial forecasting models are typically developed in full precision, yet production deployment often requires low-precision inference to reduce memory and computational cost. Post-training quantization (PTQ) enables such…
38 -
arXiv — Machine Learning research 1d ago
Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling
arXiv:2608.12271v1 Announce Type: new Abstract: Global weather reanalyses and forecasts resolve the evolving atmospheric state on coarse grids, but site-specific applications require predictions at arbitrary locations where near-surface conditions also depend on unresolved…
21 -
arXiv — Machine Learning research 1d ago
A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions
arXiv:2608.12302v1 Announce Type: new Abstract: We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural…
20 -
arXiv — Machine Learning research 1d ago
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
arXiv:2608.12306v1 Announce Type: new Abstract: Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We…
29 -
arXiv — Machine Learning research 1d ago
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
arXiv:2608.12307v1 Announce Type: new Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we…
33 -
-
-
arXiv — Machine Learning research 2d ago
Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory
arXiv:2608.09997v1 Announce Type: new Abstract: Transformers have had a profound impact on the world of language processing and computer vision. As efforts to answer the million-dollar question of ``How does a Transformer learn?" have been increasing, existing interpretability…
6 -
arXiv — Machine Learning research 2d ago
Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification
arXiv:2608.10007v1 Announce Type: new Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and…
30