News / #training Tag Training 500 articles archived under #training · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models arXiv:2609.30292v1 Announce Type: new Abstract: Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust… 28 arXiv — NLP / Computation & Language research 7h ago Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment arXiv:2609.30802v1 Announce Type: new Abstract: Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during… 35 arXiv — NLP / Computation & Language research 7h ago Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models arXiv:2609.30935v1 Announce Type: new Abstract: Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks… 19 arXiv — NLP / Computation & Language research 7h ago Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge arXiv:2609.31511v1 Announce Type: new Abstract: We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a… 9 arXiv — NLP / Computation & Language research 7h ago Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models arXiv:2609.30784v1 Announce Type: cross Abstract: This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes… 23 r/LocalLLaMA community 16h ago Public MCP server for Canadian privacy law data (free, no auth) - works with any client that speaks Streamable HTTP I built this, so the disclosure goes up front. It's a free, public MCP server plus a REST API with Canadian privacy law data. The MCP endpoint is at https://movahedi.ca/mcp and it uses Streamable HTTP, so any client that speaks that transport can connect. No signup, no API key,… 36 r/LocalLLaMA community 2d ago Jev vs. Kev: open-source Jev alternative tested side by side We hosted Kev 4B (Jared Palmer's Apache-2.0 fine-tune of Qwen3.5-4B) and ran it side by side with Jev on the same endpoint to see how it compares. We built a fresh set of 362 items published after both models shipped (new arXiv papers, Stack Exchange questions, GitHub issues),… 9 arXiv — Machine Learning research 3d ago Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization arXiv:2609.28998v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) has become a popular parameter-efficient fine-tuning method for large language models. A key challenge in LoRA is how to determine the rank of each adaptation matrix, as rank directly controls its… 29 arXiv — NLP / Computation & Language research 3d ago Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus arXiv:2609.28747v1 Announce Type: new Abstract: A language model's confidence in an answer is often read as a proxy for how well it knows the corresponding fact. This manual documents an open toolkit built to test that reading directly: a small causal language model is… 27 arXiv — NLP / Computation & Language research 3d ago Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition arXiv:2609.29800v1 Announce Type: new Abstract: Adapting multilingual speech foundation models to low-resource languages remains difficult, especially for languages that are poorly represented during pre-training. While parameter-efficient fine-tuning (PEFT) reduces the cost of… 29 arXiv — NLP / Computation & Language research 3d ago The Fellowship of the Query: Learning Retrieval Actions arXiv:2609.28653v1 Announce Type: cross Abstract: Retrieval-augmented question answering requires control decisions about when to decompose a question, search, reformulate, extract evidence, synthesize facts, verify progress, and stop. We study whether trajectory fine-tuning can… 16 r/LocalLLaMA community 3d ago FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face from FreedomIntelligence: HuatuoGPT-3-27B is a medical LLM built on Qwen3.8-27B with One-stage Policy Optimization (OnePO) . OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher… 12 arXiv — Machine Learning research 4d ago WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps arXiv:2609.27033v1 Announce Type: new Abstract: Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution,… 27 arXiv — Machine Learning research 4d ago Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders arXiv:2609.27248v1 Announce Type: new Abstract: Next-token prediction has enabled highly fluent autoregressive language models, but it represents global structure only indirectly through sequential factorization. In contrast, high-fidelity autoencoders have become a standard… 22 arXiv — NLP / Computation & Language research 4d ago Does Step Law Transfer to Small-Scale Language Models? An Empirical Recalibration Below 59M Parameters arXiv:2609.27581v1 Announce Type: cross Abstract: Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre-training language models. It was calibrated on models between 59M and 1B parameters; the small-model regime N < 59M was never… 31 arXiv — NLP / Computation & Language research 4d ago Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs arXiv:2609.27262v1 Announce Type: new Abstract: Production customer-support systems often require LLMs to support multiple skills, such as intent classification, question answering, summarization, or tool-use decisions. A central deployment question is whether these skills… 15 arXiv — NLP / Computation & Language research 4d ago Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery arXiv:2609.27980v1 Announce Type: new Abstract: Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo}… 32 arXiv — NLP / Computation & Language research 4d ago Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following arXiv:2609.28395v1 Announce Type: new Abstract: Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings… 34 arXiv — NLP / Computation & Language research 4d ago Contrastive Learning for Authorship Verification arXiv:2609.28471v1 Announce Type: new Abstract: Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context… 9 arXiv — NLP / Computation & Language research 4d ago Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions arXiv:2609.27749v1 Announce Type: cross Abstract: The surge in AI-assisted generation of educational materials has outpaced our capacity to validate their pedagogical quality. Automated evaluation using Bloom Classifier models is a promising approach to assess educational… 13 arXiv — Machine Learning research 5d ago PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics arXiv:2609.25438v1 Announce Type: new Abstract: Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine-tuning on downstream tasks. While much of the excitement in AI for math has… 6 arXiv — Machine Learning research 5d ago From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs arXiv:2609.25655v1 Announce Type: new Abstract: As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models. This shift raises a key question… 22 arXiv — Machine Learning research 5d ago Signed Graph Pre-Training and Prompt Learning arXiv:2609.25722v1 Announce Type: new Abstract: Signed graphs arise in trust--distrust networks, financial correlation systems, biological interaction graphs, and many other domains in which edges can be positive or negative and may also be directed. While signed graph neural… 16 arXiv — Machine Learning research 5d ago Block-Level Weight-Space Structure Persists Under Post-Training: An Empirical Study Across LLM Families arXiv:2609.26147v1 Announce Type: new Abstract: Modern LLMs are deployed as families of post-trained variants (base, instruct, chat, code) derived from a shared set of pre-trained weights. We present an empirical study of how post-training transforms weight-space geometry,… 29 arXiv — NLP / Computation & Language research 5d ago ChainDoRA: Tensor-Train Factorized Weight-Decomposed Low-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning arXiv:2609.25058v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) adapts large language models (LLMs) to downstream tasks while updating only a small fraction of their pretrained parameters. Low-Rank Adaptation (LoRA) uses two trainable low-rank matrices,… 10 arXiv — NLP / Computation & Language research 5d ago TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling arXiv:2609.26347v1 Announce Type: new Abstract: The scarcity of non-English language data in specialized domains significantly limits the development of effective Natural Language Processing (NLP) tools. We present TransBERT, a novel framework for pre-training language models… 26 arXiv — NLP / Computation & Language research 5d ago Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation arXiv:2609.26536v1 Announce Type: new Abstract: In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this,… 24 r/MachineLearning community 5d ago Simulating fault tolerance with stage skipping in pipeline-parallel training [R] Our most recent work at Templar explores fault tolerance in Crucible, our distributed pre-training platform. The goal is to keep healthy workers training when another pipeline stage goes offline. Crucible combines data-parallel replicas with pipeline parallelism. Each replica… 27 The Information — AI news-outlet 5d ago Entertainment-Booking Startup Ande Raises $52 Million From Lightspeed, Redpoint Personal AI agents like Instinct and Muse have promised to take away the hassle of booking hard-to-get restaurant reservations, Broadway tickets and sporting events. Now, a startup is coming out of stealth to do the same for corporate entertainment—events, group dining,… 27 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B https://preview.redd.it/ybwqqbrst0rh1.png?width=811&format=png&auto=webp&s=1e40c5b53e304caf2a10efa2baa6f998b9a9a0bb MiMo-V2.6-Distill-Qwen-9B is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data. Made by Xiaomi… 38 arXiv — Machine Learning research 6d ago GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training arXiv:2609.22146v1 Announce Type: new Abstract: We study how post-training changes the weights of Large Language Models (LLMs) relative to their pretrained weights. Across 12 post-training chains with supervised fine-tuning (SFT) and reinforcement learning (RL), we express each… 30 arXiv — Machine Learning research 6d ago A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents arXiv:2609.22194v1 Announce Type: new Abstract: Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT)… 18 arXiv — Machine Learning research 6d ago CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models arXiv:2609.22253v1 Announce Type: new Abstract: Model merging has emerged as a promising paradigm for integrating multiple task-specific capabilities into a single large language model. However, existing methods predominantly focus on post-hoc processing of independently… 8 arXiv — Machine Learning research 6d ago Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning arXiv:2609.22257v1 Announce Type: new Abstract: Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However,… 7 arXiv — Machine Learning research 6d ago Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization arXiv:2609.22879v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, but fine-tuning them with reinforcement learning (RL) remains constrained by the cost of real-world robot interaction. Model-based… 38 arXiv — NLP / Computation & Language research 6d ago Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare for African Settings arXiv:2609.22110v1 Announce Type: new Abstract: Background: Large language models (LLMs) can improve healthcare information delivery in low-resource settings but may produce inaccurate or culturally inappropriate advice. This study evaluated domain-specific fine-tuning for… 20 arXiv — NLP / Computation & Language research 6d ago Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs arXiv:2609.22144v1 Announce Type: new Abstract: Preserving safety alignment during large language models fine-tuning is critical, however, recent studies have demonstrated that even benign fine-tuning data may contain safety-degrading samples that silently undermine safety… 14 arXiv — NLP / Computation & Language research 6d ago Team DArgk at the 2026 ELOQUENT lab for evaluating generative language model quality: Residuals of Humanity: AI Detection Evasion via GRPO Fine-Tuning arXiv:2609.22221v1 Announce Type: new Abstract: Large language models (LLMs) can generate fluent and coherent text that is increasingly difficult to distinguish from human writing, motivating the development of automatic AI-generated text detectors. However, the robustness of… 32 r/LocalLLaMA community 6d ago yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base It's not a Qwen3 finetune, it's actually its own fully custom architecture. No Llama.cpp support yet sadly (Also note that this model is NOT post-trained like Qwen3.5/3.6)   submitted by   /u/Iwaku_Real [link]… 31 r/LocalLLaMA community 7d ago Hemmingway-1: An AI that writes like a person (Qwen3.8-27B finetune) https://preview.redd.it/d7brwofkhuqh1.png?width=1400&format=png&auto=webp&s=9a0079ebc2710faff9a73378dac58c6dd67e3cec link: https://huggingface.co/Altworld/Hemmingway-1   submitted by   /u/paf1138 [link]   [comments] 5 arXiv — NLP / Computation & Language research 7d ago Prediction Dynamics in Depth-Recurrent Language Models arXiv:2609.21383v1 Announce Type: new Abstract: Depth-recurrent language models refine predictions through repeated latent updates. Why can intermediate answers agree with the endpoint while their scores continue to change? We derive a sharp margin characterization that… 4 arXiv — NLP / Computation & Language research 7d ago Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation arXiv:2609.20875v1 Announce Type: cross Abstract: Parkinson's disease (PD) often manifests through speech impairments, facilitating accessible, non-invasive, and cost-effective severity assessment for early diagnosis and progression tracking. Despite advances in speech… 31 arXiv — NLP / Computation & Language research 7d ago Cultural Alignment in Large Language Models Using Soft Prompt Tuning arXiv:2503.16094v2 Announce Type: replace Abstract: Large Language Model (LLM) alignment is commonly achieved through supervised fine-tuning or reinforcement learning, both of which require labeled or preference data and update model weights. Without targeted cultural… 4 r/LocalLLaMA community 8d ago I gave Jev, Laya, finetuned ModernCE and Qwen3.5 the controls to Doom I gave Jev, Laya, a finetuned ModernCE-base-nli and a finetuned Qwen3.5-4B the controls to Doom. Thanks to TypeSafe AI for Jev access. The video lines up the starts of four separate games with the same seed. After that, each model's actions change its own game and what it sees… 30 OpenAI Python SDK releases dev-tools 9d ago v3.16.0 3.16.0 (2026-09-18) Features api: add webhook endpoint management ( #3892 ) ( 9a11f6e ) Chores api: deprecate MCP connector_id ( #3894 ) ( 1ccaf07 ) 26 arXiv — Machine Learning research 10d ago CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling arXiv:2609.19970v1 Announce Type: new Abstract: Predicting cellular responses to perturbations supports the study of gene function, disease mechanisms, and therapeutic strategies. Despite advances in single-cell perturbation modeling, existing models typically optimize surrogate… 27 arXiv — Machine Learning research 10d ago Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification arXiv:2609.19985v1 Announce Type: new Abstract: Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive… 28 arXiv — NLP / Computation & Language research 10d ago Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry arXiv:2609.19154v1 Announce Type: new Abstract: While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training… 4 arXiv — NLP / Computation & Language research 10d ago Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes arXiv:2609.19156v1 Announce Type: new Abstract: Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning… 30 arXiv — NLP / Computation & Language research 10d ago Fine-Tuning Models for Biomedical Relation Extraction arXiv:2609.20169v1 Announce Type: new Abstract: Next-Generation Sequencing has revolutionized the study of genetic mutations, enabling large-scale investigations into their roles in disease development. However, extracting meaningful insights from the vast amount of biomedical… 7 Page 1 of 10 · 500 articles Older →