News / #training Tag Training 500 articles archived under #training · RSS Sign in to follow arXiv — Machine Learning research 7d ago Do Tabular Foundation Models Agree with Themselves? arXiv:2608.06004v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) are currently the best approach to tabular prediction problems. They are constructed as transformers that approximate the Bayesian posterior predictive distribution based on a pre-training prior.… 23 arXiv — Machine Learning research 7d ago Dynamic Graph Prompting via Topology-Routed Mixed-Curvature Experts arXiv:2608.06031v1 Announce Type: new Abstract: Dynamic graph prompting freezes a pre-trained temporal backbone and adapts it to label-scarce downstream tasks using lightweight prompts. However, existing methods operate within a single, fixed embedding space. In this work, we… 4 arXiv — Machine Learning research 7d ago Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations arXiv:2608.06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models. However,… 6 arXiv — Machine Learning research 7d ago A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance arXiv:2608.06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning,… 15 arXiv — NLP / Computation & Language research 7d ago Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness arXiv:2608.05510v1 Announce Type: new Abstract: Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates… 21 arXiv — NLP / Computation & Language research 7d ago Training-Free Token-Level Steering for LLM Personalized Co-Writing arXiv:2608.06069v1 Announce Type: new Abstract: While Large Language Models (LLMs) show great promise for personalization, they often lack specialized domain knowledge. Conventional solutions like fine-tuning struggle with high computational costs and rapid data updates, while… 12 Vercel — AI dev-tools 7d ago Audit Log Drains now support Datadog, Splunk, and Panther Audit Log Drains now stream your team's audit events into Datadog , Splunk , and Panther , joining the existing custom HTTPS endpoint and Amazon S3 destinations. An Audit Log Drain forwards every event from your team's Activity Log, plus additional audit metadata, to the… 10 r/MachineLearning community 8d ago Observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting? [D] I've been running informal experiments on RLHF-aligned LLMs and consistently observing something I can't fully explain. Posting here to get feedback and find out if this is a known phenomenon or if my methodology is flawed. The observation Inserting a long, thematically coherent… 38 arXiv — Machine Learning research 8d ago Geometry-Informed Parameter-Efficient Fine-Tuning of Pre-trained Molecular GNNs for Blood-Brain Barrier Permeability Prediction arXiv:2608.04257v1 Announce Type: new Abstract: Blood-brain barrier permeability (BBBP) prediction is a critical screening task in central nervous system drug discovery, where candidate molecules must be assessed for whether they can cross, or should be prevented from crossing,… 5 arXiv — Machine Learning research 8d ago Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that… 35 arXiv — Machine Learning research 8d ago BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning arXiv:2608.05104v1 Announce Type: new Abstract: Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges. Achieving state-of-the-art performance in natural language processing with a large pre-trained model such as… 26 arXiv — NLP / Computation & Language research 8d ago DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific fine-tuning can also weaken the safety guardrails of aligned LLMs. A widely… 27 arXiv — NLP / Computation & Language research 8d ago MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages arXiv:2608.04433v1 Announce Type: new Abstract: We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes MERaLiON-SpeechEncoder-2, a large conformer based… 29 arXiv — NLP / Computation & Language research 8d ago Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs arXiv:2608.04488v1 Announce Type: new Abstract: Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high VRAM, time, and energy costs. Parameter-Efficient Fine-Tuning (PEFT) of Small… 8 arXiv — NLP / Computation & Language research 8d ago State2State: Environment-Derived Mid-Training for LLM Agents arXiv:2608.04934v1 Announce Type: new Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers. Though effective, both remain bottlenecked by externally… 30 arXiv — NLP / Computation & Language research 8d ago Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training arXiv:2608.05148v1 Announce Type: new Abstract: Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion-supervised fine-tuning. We introduce Reasoning Core, a collection of 50 generators spanning… 37 arXiv — NLP / Computation & Language research 8d ago EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment arXiv:2608.04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic… 18 Hugging Face Daily Papers research 8d ago BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation Abstract Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution… 4 llama.cpp releases dev-tools 8d ago b10282 server: Adding spec-decode counters to /metrics endpoint ( #26389 ) server: add spec-decode counters to /metrics endpoint server: fixed review comments and now aligned param names exactly with vLLM. Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple… 28 Vercel — AI dev-tools 9d ago Export AI Gateway traces with Vercel Drains AI Gateway now produces an OpenTelemetry trace for every request. Pro and Enterprise teams can send these traces through Vercel Drains to any OTLP/HTTP-compatible endpoint, including native integrations for Braintrust, Dash0, Kubiks, Sentry, and Statsig. Each trace shows the… 33 arXiv — Machine Learning research 9d ago SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling arXiv:2608.02951v1 Announce Type: new Abstract: Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-free methods are either restricted to bandits or deterministic MDPs, such as DPO or P3O, or… 22 arXiv — Machine Learning research 9d ago SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation arXiv:2608.03092v1 Announce Type: new Abstract: We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization Policy Optimization (GDPO) has mitigated the issue of reward signals masking one another… 22 arXiv — Machine Learning research 9d ago Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models arXiv:2608.03277v1 Announce Type: new Abstract: Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. Existing aggregation-based DP-ZO methods reconstruct model updates at a… 23 arXiv — Machine Learning research 9d ago Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language arXiv:2608.03855v1 Announce Type: new Abstract: Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. However, domain-adaptive pre-training often… 14 arXiv — Machine Learning research 9d ago Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning arXiv:2608.03875v1 Announce Type: new Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass… 8 arXiv — Machine Learning research 9d ago Omega-S: A Functional Resilience Index for LLM Fine-Tuning arXiv:2608.03887v1 Announce Type: new Abstract: Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of… 23 arXiv — NLP / Computation & Language research 9d ago MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow… 15 arXiv — NLP / Computation & Language research 9d ago Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study arXiv:2608.03480v1 Announce Type: new Abstract: The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational consumption due to overly large vocabularies and embedding layers. Although… 20 arXiv — NLP / Computation & Language research 9d ago Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension arXiv:2608.03494v1 Announce Type: new Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency. We… 11 arXiv — NLP / Computation & Language research 9d ago SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs arXiv:2608.03573v1 Announce Type: new Abstract: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT… 9 Hugging Face Daily Papers research 9d ago Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge Abstract Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data… 24 arXiv — Machine Learning research 10d ago EulerLoRA: Rank-Driven Jump Dynamics for Calibrated Parameter-Efficient Fine-Tuning arXiv:2608.01142v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning, but standard LoRA produces a single deterministic model and does not directly support predictive uncertainty estimation. We introduce EulerLoRA, a stochastic… 19 arXiv — Machine Learning research 10d ago FedChronos: Federated Fine-Tuning of Time-Series Foundation Models for Privacy-Preserving Commodity Price Forecasting arXiv:2608.01290v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) such as Chronos have demonstrated strong forecasting capabilities across domains, yet adapting them to institutionally fragmented settings, where data cannot be centralized due to regulatory,… 18 arXiv — Machine Learning research 10d ago Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics arXiv:2608.01522v1 Announce Type: new Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often exhibit an apparent ceiling… 19 arXiv — NLP / Computation & Language research 10d ago A Heuristic Perspective on Debiasing Language Models arXiv:2608.00622v1 Announce Type: new Abstract: Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm. Existing methods often rely on counterfactual augmentation or representation projection.… 22 arXiv — NLP / Computation & Language research 10d ago Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning arXiv:2608.01471v1 Announce Type: new Abstract: Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data is scarce. In this work, we present SentiBanglaBERT, a two-stage Bengali… 28 arXiv — Machine Learning research 11d ago Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients arXiv:2607.29071v1 Announce Type: new Abstract: Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. Existing heterogeneous federated… 27 arXiv — Machine Learning research 11d ago The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs arXiv:2607.29601v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interference when adapting heterogeneous task sequences,… 31 arXiv — NLP / Computation & Language research 11d ago Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation arXiv:2607.28658v1 Announce Type: new Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in… 33 arXiv — Machine Learning research 11d ago DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs arXiv:2607.28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak. We present DeltaServe, a host-agnostic co-serving design that converts this… 20 arXiv — NLP / Computation & Language research 11d ago FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale arXiv:2601.22146v3 Announce Type: replace Abstract: Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised "predict the next word" objective on a vast amount of unstructured text data. To make the resulting model… 22 arXiv — NLP / Computation & Language research 11d ago Disentangling Similarity and Relatedness in Topic Models arXiv:2603.10619v3 Announce Type: replace Abstract: The recent success of large pre-trained language models (PLMs) has motivated their integration into topic modeling. However, PLM-augmented topic models differ from classical co-occurrence models such as Latent Dirichlet… 33 arXiv — Machine Learning research 14d ago A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation arXiv:2607.27501v1 Announce Type: new Abstract: We present a lightweight approach to foundation modeling (\textbf{NEXUS}) that leverages pre-trained learning from collider physics data towards out-of-domain tasks in other scientific datasets, using a fully connected autoencoder… 33 arXiv — NLP / Computation & Language research 14d ago Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection arXiv:2607.27680v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become the standard mechanism for fine-tuning large pretrained models, yet its statistical properties remain only partially understood. Existing generalization results provide upper bounds of the… 22 arXiv — Machine Learning research 14d ago VESTIGE: A Knowledge-Guided Masking Strategy for Corruption-Aware Fine-Tuning of Genomic Transformers, Validated on Ancient DNA Reconstruction arXiv:2607.27712v1 Announce Type: new Abstract: Standard masked-language-model fine-tuning applies a uniform masking probability across every token position, assuming reconstruction difficulty is position-agnostic. When the degradation process is characterised and concentrated… 4 arXiv — Machine Learning research 14d ago Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control arXiv:2607.27914v1 Announce Type: new Abstract: Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model predictive control and reinforcement learning are widely studied, but… 31 arXiv — Machine Learning research 14d ago Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting arXiv:2607.27928v1 Announce Type: new Abstract: The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data. While manual heuristics were prevalent in early models, they increasingly fail to… 21 arXiv — NLP / Computation & Language research 14d ago TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement arXiv:2607.27940v1 Announce Type: cross Abstract: Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint [1] (arXiv:2606.20553), demonstrates that a malicious parameter server can… 31 arXiv — Machine Learning research 14d ago HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks arXiv:2607.28301v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race detection and benchmark question answering. However, knowledge alone does not… 23 arXiv — NLP / Computation & Language research 14d ago CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising arXiv:2607.28236v1 Announce Type: cross Abstract: Pre-trained language models have significantly improved sentence representation learning, yet their embedding remain sensitive to semantic preserving textual perturbations such as synonym substitution, masking and word dropout.… 19 Page 2 of 10 · 500 articles ← Newer Older →