News / #training Tag Training 500 articles archived under #training · RSS Sign in to follow arXiv — NLP / Computation & Language research 10d ago Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models arXiv:2609.19384v1 Announce Type: cross Abstract: Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude… 5 arXiv — NLP / Computation & Language research 10d ago AutoData: Agentic Search for Pre-training Data Selection arXiv:2609.19754v1 Announce Type: cross Abstract: LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame… 20 r/LocalLLaMA community 10d ago Thank you :) Swift Qwen 3.8 27B now has 100k+ downloads, is #1 finetune and #9 model on HuggingFace Trending Hey everyone, Jovan from UkisAI here, a small lab building the tech to make tiny frontier LLMs possible (and doing it open-source!) The purpose of this post is simply to thank the community for all the amazing finetunes, quantizations and overall improvements over our original… 38 Hugging Face Daily Papers research 11d ago CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents Abstract Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent… 20 arXiv — Machine Learning research 11d ago Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds arXiv:2609.17560v1 Announce Type: new Abstract: Every production model is updated, by retraining, fine-tuning, quantization, or a silent vendor swap, and each update risks being worse than what it replaced. We formalize update promotion as certified paired risk-difference… 26 arXiv — Machine Learning research 11d ago Hybrid coupling with numerics-informed neural networks and the overlapping Schwarz alternating method arXiv:2609.17841v1 Announce Type: new Abstract: We develop a hybrid modeling framework for coupling pre-trained numerics-informed neural networks (NINNs) with classical full order models (FOMs) using the overlapping Schwarz alternating method. We consider the two-dimensional… 34 arXiv — Machine Learning research 11d ago Dataset-Dependent Effects of Cross-Depth Aggregation and Soft-Routed Experts in EEG Foundation Model Fine-Tuning arXiv:2609.17886v1 Announce Type: new Abstract: EEG decoding tasks can rely on different temporal dynamics and cross-channel relationships. We test whether specialized modules improve a fully fine-tuned EEG foundation model by augmenting CBraMod with cross-depth Attention… 9 arXiv — Machine Learning research 11d ago Spatially Adaptive Noise Injection arXiv:2609.18466v1 Announce Type: new Abstract: Diffusion samplers reverse a learned noising process using either stochastic (DDPM) or deterministic (DDIM) updates, which represent endpoints of a single family controlled by a scalar noise-injection variance that is applied… 5 arXiv — NLP / Computation & Language research 11d ago SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale arXiv:2609.17848v1 Announce Type: new Abstract: Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents. We evaluate supervised fine-tuning (SFT) with LoRA, reinforcement… 34 arXiv — NLP / Computation & Language research 11d ago Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning arXiv:2609.18417v1 Announce Type: new Abstract: Multi-turn agent trajectories often contain redundant rounds (failed tool calls, parallel sub-queries, verification-only steps) that inflate both training and inference cost. We propose viewing each trajectory as a… 15 arXiv — NLP / Computation & Language research 11d ago WordPolo: Evaluating Language Models Through Iterative Semantic Feedback arXiv:2609.19006v1 Announce Type: new Abstract: Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into the quality or faithfulness of their reasoning processes. We… 23 arXiv — NLP / Computation & Language research 11d ago Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs arXiv:2609.17981v1 Announce Type: cross Abstract: Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR)… 32 arXiv — NLP / Computation & Language research 12d ago POSPAN: Position-Constrained Span Masking for Language Model Pre-training arXiv:2609.16061v1 Announce Type: cross Abstract: Span-level masked language modeling (MLM) has shown to be advantageous to pre-trained language models over the original single-token MLM, as entities/phrases and their dependencies are critical to language understanding. Previous… 17 arXiv — Machine Learning research 12d ago Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning arXiv:2609.16255v1 Announce Type: new Abstract: We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Our approach fine-tunes a 2B-parameter model using only $\sim$900… 21 arXiv — NLP / Computation & Language research 12d ago TAME: Token Attribution and Masking for Emergent misalignment arXiv:2609.16754v1 Announce Type: cross Abstract: Fine-tuning an aligned language model on narrow, flawed data can induce harmful behavior far outside the training domain, known as emergent misalignment (EM). Prior work has localized EM in model weights, activations, and… 6 arXiv — NLP / Computation & Language research 12d ago Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT) arXiv:2609.16010v1 Announce Type: new Abstract: The complexity of legal language and limited accessibility to legal information pose significant challenges to justice delivery in Nepal. Traditional legal services remain inaccessible to many citizens due to language barriers,… 25 arXiv — NLP / Computation & Language research 12d ago Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents arXiv:2609.16053v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents operating over extended interactions. Existing memory systems primarily update memory when new information arrives, treating retrieval as the endpoint of memory access rather than… 6 arXiv — NLP / Computation & Language research 12d ago Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation arXiv:2609.16059v1 Announce Type: new Abstract: Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fine-Tuning (SFT), which often leads to surface-level pattern matching and degrades… 22 arXiv — NLP / Computation & Language research 12d ago ReMova: Fine-tuning LLMs for English to Belarusian translation arXiv:2609.16427v1 Announce Type: new Abstract: This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation. Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the… 4 arXiv — NLP / Computation & Language research 12d ago Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data arXiv:2609.16532v1 Announce Type: new Abstract: Continued pretraining (CPT) with data augmentation such as paraphrasing can store inside a large language model (LLM) the knowledge of a small source corpus. The stored knowledge, however, is not always retrieved correctly. We… 10 arXiv — NLP / Computation & Language research 12d ago DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization arXiv:2609.16661v1 Announce Type: new Abstract: Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for… 18 arXiv — NLP / Computation & Language research 12d ago Where Should a Document Live: Context, Representations, or Parameters? arXiv:2609.17346v1 Announce Type: new Abstract: To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as… 21 arXiv — NLP / Computation & Language research 12d ago LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture arXiv:2609.16730v1 Announce Type: cross Abstract: Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation… 33 OpenAI Python SDK releases dev-tools 12d ago v3.14.1 3.14.1 (2026-09-15) Bug Fixes client: validate retry limits and preserve application errors ( #3867 ) ( f86c721 ) correct typo "th" to "the" in StreamAlreadyConsumed error message ( #3022 ) ( 7186203 ) examples: correct Azure endpoint hostname ( #3298 ) ( 543516c ) examples:… 16 arXiv — Machine Learning research 13d ago Task-Aware Federated Fine-Tuning for MoE-based Large Language Models arXiv:2609.13395v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) has become a widely adopted architecture for Large Language Models (LLMs), as it improves model capacity while limiting computational overhead through sparse expert activation. This property makes MoE-based… 26 arXiv — Machine Learning research 13d ago Adaptive Phase-Switching for Communication-Efficient Federated LoRA Fine-Tuning arXiv:2609.13512v1 Announce Type: new Abstract: Federated fine-tuning of large language models with low-rank adaptation reduces per-client trainable parameters, but client-to-server communication remains the dominant cost. Existing accounting for federated LoRA protocols omits… 10 arXiv — Machine Learning research 13d ago Pre-training with Graph Transformers arXiv:2609.13844v1 Announce Type: new Abstract: This article investigates pre-training strategies for graph transformers in the biochemistry domain. By conducting comprehensive experiments, the study reveals that supervised pre-training using computed properties as labels… 9 arXiv — Machine Learning research 13d ago A Multi-Resolution Multi-Domain Pre-Training Framework for Universal Traffic Forecasting arXiv:2609.13878v1 Announce Type: new Abstract: Spatio-temporal traffic data are central to intelligent transportation systems, yet their heterogeneity poses significant challenges for large-scale modeling. Existing pre-trained models often rely on a homogeneous modeling… 24 arXiv — Machine Learning research 13d ago Retrieval-Guided Fine-Tuning as Noisy Estimation: Risk bounds and Architectural Analysis arXiv:2609.14485v1 Announce Type: new Abstract: Retrieval-Guided Fine-Tuning (RAG-FT) incorporates retrieved data directly into the training objective, but the statistical consequences of noisy retrieval during training remain theoretically undercharacterized. We study this… 23 arXiv — NLP / Computation & Language research 13d ago Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning arXiv:2609.13151v1 Announce Type: new Abstract: Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this inefficiency by… 38 arXiv — NLP / Computation & Language research 13d ago LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference arXiv:2609.13682v1 Announce Type: new Abstract: We introduce LayerRoute, a parameter-efficient method for adaptive transformer layer-skipping that combines per-layer hard-gated routing (trained via a straight-through estimator) with joint LoRA fine-tuning. LayerRoute augments… 11 arXiv — NLP / Computation & Language research 13d ago Learning to Refer from Estimated Listener Gaze arXiv:2609.14207v1 Announce Type: new Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals.… 29 arXiv — NLP / Computation & Language research 13d ago Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports arXiv:2609.14226v1 Announce Type: new Abstract: Automated radiology report generation has advanced rapidly in diagnostic accuracy, yet generated reports frequently diverge from the stylistic conventions of authentic radiologist writing in structure, diction, and uncertainty… 11 arXiv — NLP / Computation & Language research 13d ago Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection arXiv:2609.14825v1 Announce Type: new Abstract: Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency to hallucinate. Retrieval augmentation, fine-tuning, and external verifiers require new infrastructure that clinical… 32 Hacker News — AI on Front Page community 14d ago EuroBirdPortal – Live bird movements across Europe Article URL: https://www.eurobirdportal.org/ebp/en/ Comments URL: https://news.ycombinator.com/item?id=49693610 Points: 201 # Comments: 61 10 arXiv — Machine Learning research 14d ago Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning arXiv:2609.12584v1 Announce Type: new Abstract: Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly includes overrepresented sample groups while weakly… 30 arXiv — Machine Learning research 14d ago Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery arXiv:2609.12591v1 Announce Type: new Abstract: Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities. Understanding the internal computations that support these adaptations is… 25 arXiv — Machine Learning research 14d ago Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner arXiv:2609.12651v1 Announce Type: new Abstract: While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent… 33 arXiv — NLP / Computation & Language research 14d ago HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge arXiv:2609.12260v1 Announce Type: new Abstract: Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major… 37 arXiv — NLP / Computation & Language research 14d ago All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing arXiv:2410.17355v4 Announce Type: replace Abstract: Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity typing tasks where the space of labels is extremely large. In this work, we… 38 arXiv — NLP / Computation & Language research 14d ago ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning arXiv:2604.19254v2 Announce Type: replace Abstract: Popular low-rank parameter-efficient fine-tuning (PEFT) methods represent adaptation as separate updates to selected backbone weights, without maintaining an explicit task-specific state that is updated and reused across depth.… 35 Hugging Face Daily Papers research 14d ago PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization Abstract Posterior Label Correction DPO improves preference optimization by routing noisy pairwise labels into clean, flipped, or tied cases using calibrated policy-reference margins. Generated by thinkingmachines/Inkling-Small Direct Preference Optimization (DPO) simplifies… 24 r/LocalLLaMA community 15d ago internlm/Intern-S2 · Hugging Face from internlm: We introduce Intern-S2-397B , our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent… 6 r/LocalLLaMA community 15d ago I built a serverless hosting platform for LoRA adapters with vLLM It’s always bothered me that after fine-tuning a model for a project, there isn’t a particularly easy way to host it without either running it locally and keeping a GPU on 24/7 or paying for an entire GPU server. There are managed options for LoRA serving on top of vLLM (AWS),… 32 Simon Willison community 16d ago So you want to use OpenRouter? So you want to use OpenRouter? One of OpenRouter's selling points is that it "handles fallbacks automatically and picks the most cost-effective option for each request", so you can call a single API endpoint for a model and get routed to the best available backend provider.… 33 r/LocalLLaMA community 16d ago Is a ZIMA Board 2 + RTX 2000 ADA the cheapest path to a decent Qwen-3.8 27b self-contained endpoint? I just watched a YouTube from Luke’s Dev Lab where he literally just plugged a RTX 2000 ADA Into the side of the Zima Board 2’s PCIE socket and it just friggin worked and had great token speed despite running on shitty Ollama. Ran off the Zima’s power supply and everything.… 26 r/LocalLLaMA community 16d ago CodeFinetuner: Fine-tune a local code autocomplete model on your own codebase Hi everyone, I was interested in learning LoRA fine-tuning, and ended up building CodeFinetuner over the past few months, a full pipeline that fine-tunes a small code autocomplete model (e.g. Qwen2.5-Coder-3B) specific to a codebase. You can then use the resulting GGUF model via… 33 Hugging Face Daily Papers research 16d ago Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning Abstract Optimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language. Generated by thinkingmachines/Inkling-Small Reasoning… 36 r/LocalLLaMA community 16d ago Fine-tuning Qwen 3 4B Base on 100 zebra puzzles yielded +31% on MATH-500. 6.5-min (Single H100/H200) reproduction notebook included.   submitted by   /u/TGSCrust [link]   [comments] 36 arXiv — Machine Learning research 17d ago Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning arXiv:2609.10656v1 Announce Type: cross Abstract: Selecting LoRA rank for diffusion fine-tuning requires balancing quality and compute cost. We present a controlled study on CIFAR-10 using a DDPM U-Net with ranks {2,4,8,16,32}, fixed optimization settings, and a reproducible… 6 Page 2 of 10 · 500 articles ← Newer Older →