News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago Recursive Self-Improvement via On-Policy Distillation for Reasoning arXiv:2609.30652v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the… 15 arXiv — NLP / Computation & Language research 7h ago From annotation to reasoning: Culture in language models arXiv:2609.30897v1 Announce Type: new Abstract: How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave… 30 arXiv — NLP / Computation & Language research 7h ago LocUS: Head Selection and Subspace Projection for Targeted Activation Steering arXiv:2609.31122v1 Announce Type: new Abstract: Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the… 16 arXiv — NLP / Computation & Language research 7h ago When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess arXiv:2609.30328v1 Announce Type: cross Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for.… 11 arXiv — NLP / Computation & Language research 7h ago Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency arXiv:2609.31619v1 Announce Type: cross Abstract: Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly… 19 r/LocalLLaMA community 1d ago Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_nonogram_logic_puzzle/ . That thread shaped v1.2: Reasoning effort is explicit per run Every prompt and output is public. All current top ranking private and open weight… 7 Vercel — AI dev-tools 1d ago Ember-1 from Fireworks now available on AI Gateway Ember-1 from Fireworks is now available on AI Gateway . Ember-1 is a research preview reasoning model built on Kimi K3 for coding and agentic workflows. Fireworks reports approximately 40% fewer generated tokens than Kimi K3 at comparable quality across its evaluations. For… 13 r/LocalLLaMA community 1d ago Improved and fixed template for GPT-OSS (again). Includes preserve_thinking and fix for Unsloth-induced bug I posted an updated GPT-OSS template a couple of months ago , which was based on Unsloth's version . It turns out that both Unsloth's version (and, thus, mine) contain a very serious bug that can degrade the model when chat history is replayed and contains previous reasoning… 14 r/LocalLLaMA community 2d ago Swift1.5-Qwen3.8-Flash-Next is phenomenal vs. base 3.8-Flash! TL;DR - Swift Flash is a killer model that massively reduces excess reasoning. Try it out! If you haven't seen from my previous comparison posts , I'm a huge fan of the Swift Qwen3.8 models. I've been using 27B since it dropped, and I'm really impressed with the performance and… 15 arXiv — Machine Learning research 3d ago RLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory arXiv:2609.28625v1 Announce Type: new Abstract: Despite the importance of reinforcement learning with verifiable rewards (RLVR), the extent to which it can learn new reasoning capabilities remains debated. Here we study the optimization landscape of RLVR on algorithmic tasks,… 36 arXiv — Machine Learning research 3d ago Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models arXiv:2609.28682v1 Announce Type: new Abstract: Post-training hybrid reasoning models in NoThink mode has attracted growing interest as a way to improve performance while keeping inference fast. However, these gains may draw on thinking behavior already accessible through the… 4 arXiv — Machine Learning research 3d ago Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference arXiv:2609.29499v1 Announce Type: new Abstract: Static pruning imposes one sparse structure on every prompt, even though reasoning, retrieval, generation, coding, and translation can depend on different parts of a language model. We introduce Task-Aware Spectral Pruning (TASP),… 11 arXiv — Machine Learning research 3d ago CataOPD: Catalytic On-Policy Distillation for Large Language Model Reasoning arXiv:2609.29518v1 Announce Type: new Abstract: Reinforcement learning (RL) and on-policy distillation (OPD) are two representative paradigms for improving large language model reasoning. However, when no correct trajectory is sampled, RL lacks a positive correctness signal,… 12 arXiv — Machine Learning research 3d ago Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning arXiv:2609.29548v1 Announce Type: new Abstract: Language-model advice can accelerate reinforcement learning, but calls are costly and returned actions may be stale or wrong. We formulate advice acquisition as a response-contingent metareasoning problem: before querying, the… 4 arXiv — NLP / Computation & Language research 3d ago ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks arXiv:2609.29102v1 Announce Type: new Abstract: Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning… 38 arXiv — NLP / Computation & Language research 3d ago Reasoning Instructions Can Break Answer Decoding in Vision--Language Models arXiv:2609.29278v1 Announce Type: new Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On… 7 arXiv — NLP / Computation & Language research 3d ago Rufus-Air: An Open LLM Post-Training Recipe arXiv:2609.29421v1 Announce Type: new Abstract: Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search… 31 arXiv — NLP / Computation & Language research 3d ago Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions arXiv:2609.29496v1 Announce Type: new Abstract: Natural language explanation generation serves as a key mechanism for exposing and evaluating vision-language reasoning. Prior work on explanation-driven vision-language models predominantly follows a post-hoc (answer-first)… 11 arXiv — NLP / Computation & Language research 3d ago When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing arXiv:2609.28475v1 Announce Type: cross Abstract: Forecasting agents increasingly combine language-model reasoning, retrieval, ensembling, and calibration, but it remains unclear when each behavior should be trusted. We study this question on ForecastBench-style binary… 11 r/LocalLLaMA community 3d ago How are you guys thinking about context now, and building around it? Not asking for anyone’s secrets of the trade, I’m more curious how people are thinking about context now that newer models chew through huge amounts of it for reasoning. the TLDR: I’m starting to think of context less as working memory and more as a temp scratchpad to start each… 13 r/LocalLLaMA community 3d ago UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy Hey everyone, Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens related to pathological overthinking patterns and restoring accuracy via RL (GSPO) and OPD . After amazing feedback and 350k+… 26 r/LocalLLaMA community 3d ago MiniMax M3.1 (Space Bunny Alpha) thinks in caveman mode The CoT of MinimMax M3.1, currently available in openrouter and opencode under the guise of "Space Bunny Alpha", has the familiar look of caveman mode in order to save tokens. This has no impact on the final output. (note: in the first screenshot, pi-caveman is set to off; in… 30 r/LocalLLaMA community 4d ago model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp Model Overview Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a… 31 arXiv — Machine Learning research 4d ago Signal2Symbol: Neuro-Symbolic Temporal Reasoning for Explainable Physiological Time-Series Anomaly Detection arXiv:2609.26820v1 Announce Type: new Abstract: Physiological time series such as electrocardiograms (ECG) and electroencephalograms (EEG) exhibit complex temporal structure, substantial acquisition variability, and a strong need for transparent decision-making. Although deep… 4 arXiv — Machine Learning research 4d ago Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models arXiv:2609.27166v1 Announce Type: new Abstract: Capability and efficiency are two key dimensions of reasoning in large language models (LLMs). Capability refers to the ability to solve a given problem correctly, whereas efficiency refers to the ability to do so with limited… 29 arXiv — Machine Learning research 4d ago DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment arXiv:2609.27572v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a key paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing reward systems, such as rule-based and reward-model-based, often exhibit issues… 10 arXiv — Machine Learning research 4d ago From Reasoning Strings to Partial Orders: Verifier-Certified Rule Transport through Quotient Policy Optimization arXiv:2609.27833v1 Announce Type: new Abstract: Many computations admit several valid execution orders because independent subgoals or disjoint state updates can commute. Reinforcement learning with verifiable rewards usually treats each successful trace as a separate token… 25 arXiv — Machine Learning research 4d ago RL Starts before RL: On Policy Distillation for Better Reinforcement Learning arXiv:2609.28145v1 Announce Type: new Abstract: Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond… 24 arXiv — NLP / Computation & Language research 4d ago Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court arXiv:2609.26945v1 Announce Type: new Abstract: Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in… 19 arXiv — NLP / Computation & Language research 4d ago LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning arXiv:2609.27009v1 Announce Type: new Abstract: Large language models are increasingly applied to high-risk domains such as law, yet complex legal reasoning remains limited by two structural challenges. First, existing RAG and GraphRAG methods emphasize lexical or semantic… 10 arXiv — NLP / Computation & Language research 4d ago Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning arXiv:2609.27156v1 Announce Type: new Abstract: Large reasoning models can produce correct yet unnecessarily long reasoning traces. Existing methods improve reasoning efficiency with trajectory-level objectives or local token- and step-level signals, but rarely model inter-step… 11 arXiv — NLP / Computation & Language research 4d ago Realize What Matters: Principled Context Representation for Large-Scale Reasoning arXiv:2609.27173v1 Announce Type: new Abstract: Solving complex tasks in domains such as science, medicine, law, and finance often requires assembling interdependent information scattered across vast, heterogeneous sources far beyond model context limits. Existing approaches… 5 arXiv — NLP / Computation & Language research 4d ago LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models arXiv:2609.27220v1 Announce Type: new Abstract: Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer… 32 arXiv — NLP / Computation & Language research 4d ago Planned Test-Time Scaling with Coordinated Reasoning Paths arXiv:2609.27374v1 Announce Type: new Abstract: Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce… 11 arXiv — NLP / Computation & Language research 4d ago The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA arXiv:2609.27669v1 Announce Type: new Abstract: Small language models (SLMs) are increasingly paired with knowledge graphs (KGs), yet end-to-end KG question answering conflates graph access, search, navigation, reasoning, and answer generation. This coupling makes it difficult… 7 arXiv — NLP / Computation & Language research 4d ago Improving LLM-based Autonomous Web Agents with Filtering arXiv:2609.27770v1 Announce Type: new Abstract: Autonomous web agents, powered by Large Language Models (LLMs), have garnered significant attention for automating various web-based tasks with multi-step reasoning and decision-making capabilities. An open research question in the… 29 arXiv — NLP / Computation & Language research 4d ago LabourCrew: A Multi-Agent RAG Framework for Trustworthy Adversarial Deliberation and Statutory Reasoning over Labour Law arXiv:2609.27814v1 Announce Type: new Abstract: In statutory question answering, every claim must be traceable to evidence, not merely relevant, since unverifiable labour-rights answers carry serious legal consequences. Current systems fall short: single-pass RAG cannot detect… 30 arXiv — NLP / Computation & Language research 4d ago Towards Efficient Reasoning: Learning Causal Shortcuts for Diffusion Language Models arXiv:2609.28272v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) have attracted significant attention for their strong reasoning ability. However, under a bidirectional attention mechanism, DLMs operate over an exponentially large exploration space compared to… 6 arXiv — NLP / Computation & Language research 4d ago Large Knowledge Model: From Papers to a Scientific Reasoning Landscape arXiv:2609.27297v1 Announce Type: cross Abstract: Accumulated scientific knowledge advances inquiry when prior findings help researchers choose new questions, design investigations, and interpret results. Realizing this value at scale requires access to the reasoning that… 9 r/LocalLLaMA community 4d ago New DeepThink built on GLM Source: https://rogueon.ai/blog/rogue-deep-think-preview/tecnica Could this kinda stuff be the way for open models to prevail? Since open models are so much cheaper than, say, Astra and Fable, maybe just cranking up reasoning to the max can get us a bit closer.   submitted… 19 NVIDIA Developer Blog official-blog 4d ago Introducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought Reasoning Radiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and... 17 Latent.Space news-outlet 4d ago 🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics) Radical Numerics is using biological chain-of-thought and multimodal perception to keep up with the bio-defense arms race, design new genomes and gain insights into biology itself. 12 arXiv — Machine Learning research 5d ago xWhyL: Causal Interactive Learning arXiv:2609.26037v1 Announce Type: new Abstract: Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for learning about causality. Despite this, learning from those abductive signals is… 18 arXiv — Machine Learning research 5d ago CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning arXiv:2609.26094v1 Announce Type: new Abstract: Multi-step causal reasoning requires chaining inferences where each step constrains the next. An early error propagates silently, and a correct answer reached via flawed logic evades outcome-level detection. In specialized domains,… 27 arXiv — Machine Learning research 5d ago Spectral Tail Interventions in Decoder-Only Language Models: Reasoning-Sensitive Weight Structure from Controlled Surgery arXiv:2609.26165v1 Announce Type: new Abstract: Weight-space structure often correlates with language-model behavior, but correlation alone does not establish computational involvement. We study concentrated upper spectral tails in decoder-only transformers through controlled… 34 arXiv — Machine Learning research 5d ago Beyond Imitation: Auditing the Recoverability of Reasoning in Distilled Models arXiv:2609.26216v1 Announce Type: new Abstract: A correct teacher solution becomes useful supervision when the receiving student can continue its reasoning. We measure this compatibility with prefix recovery: after revealing 25%, 50%, or 75% of a verified solution, we test… 19 arXiv — NLP / Computation & Language research 5d ago ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains arXiv:2609.25055v1 Announce Type: new Abstract: In this report we present results of the ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains. This competition aimed to advance research in document understanding through the task of Visual Question… 14 arXiv — NLP / Computation & Language research 5d ago TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks arXiv:2609.25356v1 Announce Type: new Abstract: Large language models (LLMs) offer great potential to automate a broad range of telecom engineering tasks by reasoning over standards, network configurations, mathematical models, source code, and operational logs. However,… 8 arXiv — NLP / Computation & Language research 5d ago Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation arXiv:2609.25755v1 Announce Type: new Abstract: Applying large language models to Traditional Chinese Medicine (TCM) prescription generation reveals three clinically critical gaps: models produce end-to-end mappings without auditable reasoning following the li-fa-fang-yao… 25 arXiv — NLP / Computation & Language research 5d ago Differentiable Fuzzy Inference Layer: A Monotone, Compositional Ordinal Reasoning Head for Large Language Models arXiv:2609.26113v1 Announce Type: new Abstract: A state-of-the-art language model asked to interpret "most of most students passed" typically answers "most," though composing two instances of "most" yields a proportion closer to "some." We trace this failure to an architectural… 36 Page 1 of 10 · 500 articles Older →