arXiv — Machine Learning
500 articles archived · Visit source ↗ · RSS
-
arXiv — Machine Learning research 1d ago
FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting
arXiv:2608.11254v1 Announce Type: new Abstract: Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids. All-sky imagers (ASI) provide high-resolution observations of clouds, making them well suited for…
37 -
arXiv — Machine Learning research 1d ago
Why AI Detection Fails for Academic Integrity
arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains;…
15 -
arXiv — Machine Learning research 1d ago
Basin: Efficient and Extensible Numerical Optimization in Rust
arXiv:2608.11279v1 Announce Type: new Abstract: Basin is a numerical optimization library for the Rust programming language. Numerical optimization is the task of finding the inputs that minimize a function, and it is a fundamental element across the sciences: fitting a model to…
30 -
arXiv — Machine Learning research 1d ago
Federated Learning for Distributed CNC Tool Wear Prediction
arXiv:2608.11281v1 Announce Type: new Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability. Machine learning methods have shown potential for this task, but their use in…
7 -
arXiv — Machine Learning research 1d ago
Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction
arXiv:2608.11318v1 Announce Type: new Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. We develop a decision-resource view of terminal symmetry: process evidence supplies…
28 -
-
arXiv — Machine Learning research 1d ago
Long-Horizon Forecasting of Complete Financial Statements with Forma
arXiv:2608.11327v1 Announce Type: new Abstract: Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm…
6 -
arXiv — Machine Learning research 1d ago
Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport
arXiv:2608.11342v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining, its…
35 -
arXiv — Machine Learning research 1d ago
Dynamics Models for Offline Hyperparameter Selection in Real-World RL
arXiv:2608.11349v1 Announce Type: new Abstract: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly. Prior work has proposed calibration models…
31 -
-
arXiv — Machine Learning research 1d ago
Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter
arXiv:2608.11361v1 Announce Type: new Abstract: Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis. We show that the cost-optimal…
26 -
arXiv — Machine Learning research 1d ago
PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR
arXiv:2608.11368v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this cost by assigning budgets to prompts, rollouts, or tokens according to…
26 -
arXiv — Machine Learning research 1d ago
Towards an approach to multivariate outlier detection for District Heating System data
arXiv:2608.11375v1 Announce Type: new Abstract: In this paper, we test different methods for multivariate detection of outliers in the data of transmitted heat energy in the selected substation of local District Heating System, by also considering outside ambient temperature,…
37 -
arXiv — Machine Learning research 1d ago
Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints
arXiv:2608.11383v1 Announce Type: new Abstract: We study new algorithms for Contextual Bandits with Knapsack. In these problems, there are finitely many types of customers, products, and resources. Each product is made from a fixed combination of resources, and resources have…
7 -
arXiv — Machine Learning research 1d ago
Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes
arXiv:2608.11390v1 Announce Type: new Abstract: Generative engines are reshaping the web ecosystem by making citations a key mechanism for allocating attention, attribution, and downstream value. This creates a strategic tension: content providers are incentivized to optimize…
10 -
-
arXiv — Machine Learning research 1d ago
Diffusion-Based Data-Driven Assortment Optimization
arXiv:2608.11419v1 Announce Type: new Abstract: Assortment optimization is a fundamental problem in revenue management, typically addressed using parametric choice models such as the multinomial logit (MNL) and its variants. While these models enable tractable formulations,…
36 -
-
arXiv — Machine Learning research 1d ago
Click2Poly: A VLM for vector mapping buildings and walls
arXiv:2608.11424v1 Announce Type: new Abstract: Accurate vector mapping of buildings and walls is critical for geospatial applications but remains a labor-intensive process. While recent deep learning methods have improved automatic extraction, in order to meet cartographic…
36 -
arXiv — Machine Learning research 1d ago
Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention
arXiv:2608.11427v1 Announce Type: new Abstract: Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch. We show that this distinction becomes exponential at the first context length containing two competing…
29 -
arXiv — Machine Learning research 1d ago
AutoGrable: What Is a Good Graph for a Table?
arXiv:2608.11431v1 Announce Type: new Abstract: Graph learning presupposes a graph, and tables and relational databases do not come with one. Applying a GNN to them requires deciding which entities become nodes, which of them to connect, and through which relations---a decision…
18 -
arXiv — Machine Learning research 1d ago
Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates
arXiv:2608.11435v1 Announce Type: new Abstract: Forward and inverse modeling of parametric dynamical systems requires surrogate models that are not only accurate for state prediction, but also informative for parameter calibration. However, a systematic end-to-end differentiable…
17 -
arXiv — Machine Learning research 1d ago
XGBoost "is all you need": the case of forecasting transmitted heat energy in District Heating Systems
arXiv:2608.11446v1 Announce Type: new Abstract: This paper presents a comparative study of two distinct approaches, XGBoost and Long-Short Term Memory (LSTM), for forecasting transmitted heat energy in District Heating Systems (DHS). The objective is to explore scenarios in…
23 -
arXiv — Machine Learning research 1d ago
PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition
arXiv:2608.11465v1 Announce Type: new Abstract: PAC-Bayes theory provides generalization guarantees by controlling the Kullback--Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation. However, predictive risk depends only on…
17 -
arXiv — Machine Learning research 1d ago
Dual-Primal Graph VAEs for Noisy Label Aggregation
arXiv:2608.11473v1 Announce Type: new Abstract: Inferring the ground-truth from noisy crowdsourced labels is an important theoretical and practical problem. Neural network-based methods offer an alternative to classical Bayesian models which require specifying a family of…
36 -
arXiv — Machine Learning research 1d ago
Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
arXiv:2608.11479v1 Announce Type: new Abstract: We establish convergence guarantees of gradient descent for general feedforward neural networks of arbitrary width or depth, with no special requirements on the initialization or dataset. We only assume that the activation…
14 -
arXiv — Machine Learning research 1d ago
Defending against Model Extraction for GNNs with Model Reprogramming
arXiv:2608.11495v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box deployment exposes them to Model Extraction (ME) attacks, in which adversaries steal…
24 -
arXiv — Machine Learning research 1d ago
HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging
arXiv:2608.11499v1 Announce Type: new Abstract: Task vectors enable model merging without joint retraining. In practice, the subset of task vectors to be merged may vary, but many existing methods use scalar tuning for a particular subset, requiring repeated tuning across…
32 -
arXiv — Machine Learning research 1d ago
RelShap: Relationally Consistent Shapley Explanations
arXiv:2608.11508v1 Announce Type: new Abstract: Machine learning pipelines commonly flatten relational data into single-table representations, discarding structural constraints. Widely used Shapley value-based feature attributions then rely on feature independence, evaluating…
30 -
arXiv — Machine Learning research 1d ago
Let it Cook: Learning to Wait in Sequential Decision Making
arXiv:2608.11511v1 Announce Type: new Abstract: In sequential decision making, an agent typically observes its environment and acts at every timestep. However, such active participation may not always be necessary; tasks such as brewing coffee include periods that are served…
20 -
arXiv — Machine Learning research 1d ago
FLARE++: Low-rank attention with dynamic attention routing
arXiv:2608.11519v1 Announce Type: new Abstract: Full self-attention is a strong token mixer for PDE surrogates on irregular domains, but its quadratic cost limits its use on high-resolution problems. Efficient latent-attention models such as the Fast Low-rank Attention Routing…
32 -
arXiv — Machine Learning research 1d ago
Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks
arXiv:2608.11532v1 Announce Type: new Abstract: In recent research on the Digital Twin-based Vehicular Ad hoc Network(DT-VANET), Federated Learning (FL) has shown its ability to provide data privacy. However, Federated learning struggles to adequately train a global model when…
30 -
arXiv — Machine Learning research 1d ago
Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency
arXiv:2608.11541v1 Announce Type: new Abstract: Machine learning models should be robust, in the sense of remaining predictively consistent under permissible variations. A model's predictions should ideally remain unchanged when it is replaced by a functionally equivalent one,…
6 -
-
-
arXiv — Machine Learning research 1d ago
Sparse and robust geometric twin support vector machine via asymmetric RoBoSS loss function
arXiv:2608.11567v1 Announce Type: new Abstract: In real-world scenarios, the training data usually contains redundant features, label noise and feature noise, which provide severe challenges for the efficiency of machine learning methods. Since standard support vector machine…
31 -
arXiv — Machine Learning research 1d ago
RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers
arXiv:2608.11572v1 Announce Type: new Abstract: Coarse-grid numerical solvers can substantially reduce the computational cost of time-dependent PDE simulation, but under-resolution often degrades both the trajectory and the spatial fidelity of the solution. We introduce RECAST…
7 -
arXiv — Machine Learning research 1d ago
Dion3: Full-Stack Orthogonal Updates
arXiv:2608.11612v1 Announce Type: new Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in…
23 -
arXiv — Machine Learning research 1d ago
A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields
arXiv:2608.11613v1 Announce Type: new Abstract: In this paper, we propose a local Sinkhorn divergence framework for conditional distribution reconstruction of multidimensional random fields. By utilizing the debiased Sinkhorn divergence, our proposed approach develops a…
32 -
arXiv — Machine Learning research 1d ago
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting
arXiv:2608.11623v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational…
23 -
arXiv — Machine Learning research 1d ago
Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration
arXiv:2608.11638v1 Announce Type: new Abstract: Spatially continuous quantification of forest above-ground biomass (AGB) is what makes carbon accounting credible and mitigation strategies actionable. While field inventories provide high localized accuracy, they are spatially…
17 -
arXiv — Machine Learning research 1d ago
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem
arXiv:2608.11654v1 Announce Type: new Abstract: Despite the wide deployment of memory in large-model agents, there is no unified formal account of what a memory is or when it is optimal. This paper takes a first step toward this account. The central idea is that memory is a…
10 -
arXiv — Machine Learning research 1d ago
Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models
arXiv:2608.11656v1 Announce Type: new Abstract: Recent advances in EEG foundation models have demonstrated the potential of large-scale pretraining to enable generalizable neural decoding across subjects, recording environments, and datasets. However, dominant pretraining…
18 -
-
arXiv — Machine Learning research 1d ago
Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
arXiv:2608.11661v1 Announce Type: new Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed independently in operator learning, bipartite matching,…
26 -
arXiv — Machine Learning research 1d ago
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
arXiv:2608.11669v1 Announce Type: new Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The rubric, however, is a fixed proxy for quality,…
37 -
arXiv — Machine Learning research 1d ago
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
arXiv:2608.11674v1 Announce Type: new Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior…
34 -
arXiv — Machine Learning research 1d ago
FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation
arXiv:2608.11675v1 Announce Type: new Abstract: Coupon campaigns seek to lift both conversion and revenue, but gross merchandise value (GMV) follows a deterministic funnel from conversion to conditional order value and is zero-inflated and heavy-tailed. We propose…
29 -
arXiv — Machine Learning research 1d ago
Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning
arXiv:2608.11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting. Yet its generalization…
15 -
arXiv — Machine Learning research 1d ago
LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection
arXiv:2608.11691v1 Announce Type: new Abstract: Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a…
18