News / #training Tag Training 500 articles archived under #training · RSS Sign in to follow arXiv — NLP / Computation & Language research 5h ago I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization arXiv:2608.12957v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this gap with… 33 arXiv — Machine Learning research 5h ago Into the ORBIT for Time Series: Training Regimes for Foundation Models arXiv:2608.13262v1 Announce Type: new Abstract: Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored. As a result, pre-training distributions are often… 11 arXiv — Machine Learning research 5h ago Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology arXiv:2608.13518v1 Announce Type: new Abstract: Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint. However, recovery after a procedure often unfolds as an irregular trajectory: clinical… 17 arXiv — NLP / Computation & Language research 5h ago Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition arXiv:2608.12327v1 Announce Type: new Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium,… 34 arXiv — NLP / Computation & Language research 5h ago Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation? arXiv:2608.12332v1 Announce Type: new Abstract: In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters. In this work, we uncover… 24 arXiv — NLP / Computation & Language research 5h ago LoRA-Diffusion: Parameter-Efficient Fine-Tuning via Low-Rank Trajectory Decomposition arXiv:2608.12328v1 Announce Type: new Abstract: Parameter-efficient fine-tuning methods such as LoRA have transformed the adaptation of large autoregressive language models, enabling task-specific customization with substantially fewer trainable parameters. However, these… 11 arXiv — NLP / Computation & Language research 5h ago Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring arXiv:2608.12330v1 Announce Type: new Abstract: The detection of online sexism remains an open problem. Sexism detection is inherently subjective, yet most existing systems reduce multi-annotator labels to a single majority decision and treat all instances uniformly. This… 6 arXiv — NLP / Computation & Language research 5h ago Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model arXiv:2608.13277v1 Announce Type: new Abstract: We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular… 36 arXiv — NLP / Computation & Language research 1d ago Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport arXiv:2608.11342v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining,… 35 arXiv — Machine Learning research 1d ago Distillation of Foundation Models for Time-dependent PDEs arXiv:2608.11937v1 Announce Type: new Abstract: Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively to new downstream tasks. After fine-tuning on only a few… 25 arXiv — NLP / Computation & Language research 1d ago Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs arXiv:2608.11573v1 Announce Type: new Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement… 35 arXiv — NLP / Computation & Language research 1d ago AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention arXiv:2608.11758v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic… 34 arXiv — NLP / Computation & Language research 1d ago TELLME: Test-Enhanced Learning for Language Model Enrichment arXiv:2608.11788v1 Announce Type: new Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale… 16 arXiv — NLP / Computation & Language research 1d ago Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existing approaches to building SLMs typically… 34 arXiv — NLP / Computation & Language research 1d ago DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation arXiv:2608.11889v1 Announce Type: cross Abstract: Prompting-based (\textit{i}.\textit{e}., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not changed for the task, face three problems: (\textit{i})~relying on coarse-grained schema… 38 Hugging Face Daily Papers research 1d ago MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Abstract Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only… 21 r/LocalLLaMA community 1d ago CohereLabs/North-Micro-Vision-Instruct · Hugging Face North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation for prototyping, task-specific fine-tuning, and specialized multimodal… 8 arXiv — Machine Learning research 2d ago Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons arXiv:2608.10045v1 Announce Type: new Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. In this problem, the goal is to… 7 arXiv — NLP / Computation & Language research 2d ago Procedural Fairness Failures in RLHF from Preference Averaging arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural… 13 arXiv — Machine Learning research 2d ago SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks arXiv:2608.10144v1 Announce Type: new Abstract: We consider federated parameter efficient fine-tuning of large neural networks with low-rank adaptation (LoRA,~Hu et al.\ 2022). Combining LoRA with federated PEFT introduces challenges absent from either setting alone: clients may… 22 arXiv — Machine Learning research 2d ago Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning arXiv:2608.10473v1 Announce Type: new Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online… 19 arXiv — Machine Learning research 2d ago Diffract: Spectral View of LLM Domain Adaptation arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight… 6 arXiv — NLP / Computation & Language research 2d ago ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization arXiv:2608.11045v1 Announce Type: cross Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization… 30 arXiv — NLP / Computation & Language research 2d ago Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies arXiv:2608.10273v1 Announce Type: new Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic… 4 arXiv — NLP / Computation & Language research 2d ago Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation arXiv:2608.10812v1 Announce Type: new Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a… 30 arXiv — NLP / Computation & Language research 2d ago REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a budget of at most 32B parameters and no model fine-tuning. Our system combines… 14 arXiv — NLP / Computation & Language research 2d ago Data Attribution of Emergent Misalignment with Persona Features arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions… 29 arXiv — NLP / Computation & Language research 2d ago myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality… 14 r/MachineLearning community 2d ago Context-Induced Activation Drift: Long benign context passively decouples RLHF alignment without adversarial prompts (Mechanistic Interpretability + Ablation) [D] TL;DR: We observed that feeding a long, benign, thematically coherent context prefix ($L \in [100, 3000]$ tokens) into google/gemma-3-1b-it causes a massive passive shift in internal activations ($\Delta h_2 \approx 3434$) at deep layers ($\sim 85%$ depth). This leads to a logit… 7 r/MachineLearning community 2d ago Research direction: Intelligent Model Weight transfer between LLMs [R] Few days ago I feel like I need to get started with researching about LLMs. One thing which strikes the most in my mind , how we can reduce the time required for pre-training an LLM model to just few minutes. Right now the most efficient method that we have is knowledge… 38 r/LocalLLaMA community 2d ago Local Benchmark : Muse Glimmer 30B vs Qwen 3.6 27B vs Gemma4 31B (and many other models and finetunes) Needs a lot of requests compared to Qwen (almost twice) and Gemma (almost x3). Final score is fine, even though it is "not a coding model" https://wonderrico.github.io/local_llm_benchmark/benchmark-main.html more details on… 31 Hugging Face Daily Papers research 2d ago Omega-S: A Functional Resilience Index for LLM Fine-Tuning Abstract Omega-S is a lightweight, data-free regularization penalty for low-rank fine-tuning that improves retention of original model capabilities by penalizing variance in weight-matrix node degrees. Generated by thinkingmachines/Inkling-Small Fine-tuning a large language… 36 Hugging Face Daily Papers research 3d ago What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems Abstract A three-stage multimodal framework improves follow-up edit recommendations in image-creation conversations by combining supervised fine-tuning, multi-objective reinforcement learning, and visual verification. Generated by thinkingmachines/Inkling-Small Conversational… 22 arXiv — Machine Learning research 3d ago Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models arXiv:2608.07890v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recent theory shows that pruning experts with the smallest router-norm changes during fine-tuning… 28 arXiv — Machine Learning research 3d ago Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention arXiv:2608.07921v1 Announce Type: new Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outliers. We validate this decomposition causally:… 20 arXiv — Machine Learning research 3d ago ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling arXiv:2608.07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. Although existing studies proposed pipeline parallelism to address the limited memory and computing resources… 10 arXiv — Machine Learning research 3d ago Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training arXiv:2608.08224v1 Announce Type: new Abstract: Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits finite capability across tasks. A representative… 32 arXiv — NLP / Computation & Language research 3d ago Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation arXiv:2608.07629v1 Announce Type: new Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. When fine-tuning these models for an unseen language,… 18 arXiv — NLP / Computation & Language research 3d ago STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs arXiv:2608.08164v1 Announce Type: new Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller… 29 arXiv — NLP / Computation & Language research 3d ago Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs arXiv:2608.08744v1 Announce Type: new Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds the one-time cost of fine-tuning. Yet most efficiency interventions target either… 8 arXiv — NLP / Computation & Language research 3d ago The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora arXiv:2608.09093v1 Announce Type: new Abstract: How a document's arrangement is written down, its notation, is a training variable that no dataset card records. The field has established that text-extraction choices change model behaviour, and has never once measured the… 35 arXiv — NLP / Computation & Language research 3d ago Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization arXiv:2608.09568v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing equally to the preference signal. However, the contribution of individual… 30 Hugging Face Daily Papers research 4d ago SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Abstract Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under… 17 arXiv — Machine Learning research 4d ago FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching arXiv:2608.07294v1 Announce Type: new Abstract: Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves… 26 Hugging Face Daily Papers research 4d ago YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family Abstract Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous operators and detection-specific components impose placement constraints absent from regular Transformer stacks. We… 17 r/LocalLLaMA community 4d ago endless-frontier/BigBang-v1 - qwen 3.5 finetunes table bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline number is basically meaningless. "Performance between DeepSeek Flash (old one)… 26 arXiv — Machine Learning research 7d ago When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters arXiv:2608.05207v1 Announce Type: new Abstract: Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discovery: mining interpretable features of a frozen forecaster's residual to drive a… 33 arXiv — Machine Learning research 7d ago Beyond Full-Model Rollback: AuroSFT for Adapter-State Multi-Task Fine-Tuning arXiv:2608.05250v1 Announce Type: new Abstract: Multi-task supervised fine-tuning (SFT) often casts a heterogeneous data mixture as a single optimization problem, even though different tasks may reach their best generalization at different times. msft exposes this mismatch… 31 arXiv — Machine Learning research 7d ago Beyond Rotations: AuroOFT for Expressive Quantized Orthogonal Fine-Tuning arXiv:2608.05253v1 Announce Type: new Abstract: Quantized orthogonal fine-tuning (qoft) enables parameter-efficient adaptation of low-bit language models by learning structured activation rotations before frozen quantized weights. However, its task-specific updates remain… 36 arXiv — Machine Learning research 7d ago Align-RAG: Alignment Is All You Need for TSFM In-Context Learning arXiv:2608.05571v1 Announce Type: new Abstract: Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge… 11 Page 1 of 10 · 500 articles Older →