News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow arXiv — NLP / Computation & Language research 18d ago LeAct: Learning to Reason from Expert Actions arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines,… 17 arXiv — Machine Learning research 18d ago MA-DAR: Manifold-Aligned Dynamic Adaptive Routing for Continual Temporal Knowledge Graph Reasoning arXiv:2607.21949v1 Announce Type: new Abstract: Continual temporal knowledge graph (TKG) reasoning aims to continuously incorporate newly emerging facts while preserving previously acquired knowledge. Replay-based continual learning has achieved promising performance by… 36 arXiv — Machine Learning research 18d ago Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs arXiv:2607.22314v1 Announce Type: new Abstract: Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random… 8 arXiv — NLP / Computation & Language research 18d ago J-CoT: Chain-of-Thought in J-Space arXiv:2607.21981v1 Announce Type: new Abstract: Chain-of-thought prompting improves language-model reasoning by carrying intermediate states across successive computation steps. However, relying on natural language as the only recurrent interface is overly restrictive, since… 4 arXiv — NLP / Computation & Language research 18d ago Scaling Native Multimodal Pre-Training From Scratch arXiv:2607.22043v1 Announce Type: new Abstract: Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this… 31 arXiv — NLP / Computation & Language research 18d ago Dynamic Commonsense Coordination for Empathetic Response Generation arXiv:2607.22136v1 Announce Type: new Abstract: Empathetic Response Generation (ERG) requires models to recognize users' emotions and generate empathetic responses. Commonsense knowledge has been shown to support such reasoning, yet existing approaches typically reuse fixed… 38 arXiv — NLP / Computation & Language research 18d ago Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level… 15 arXiv — NLP / Computation & Language research 18d ago Learning to Reason for Factuality arXiv:2508.05618v2 Announce Type: replace Abstract: Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substantially more hallucinations than their non-reasoning counterparts on long-form… 26 arXiv — NLP / Computation & Language research 18d ago Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs… 12 arXiv — NLP / Computation & Language research 18d ago Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models arXiv:2605.17770v4 Announce Type: replace-cross Abstract: The advancement of Large Reasoning Models (LRMs) has catalyzed a paradigm shift from reactive ``fast thinking'' text generation to systematic, step-by-step ``slow thinking'' reasoning, unlocking state-of-the-art… 17 arXiv — NLP / Computation & Language research 18d ago OpenForgeRL: Train Harness-native Agents in Any Environment arXiv:2607.21557v2 Announce Type: replace-cross Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make… 34 r/LocalLLaMA community 19d ago [Paper] RecGPT-V3 Technical Report Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it… 26 r/LocalLLaMA community 19d ago Best chat model that fits in 128gb I'm looking for a model to chat with, reasoning, maybe get some career or life coaching. I don't care at all about multimodal or coding ability Just it's intelligence in remembering context in a conversation or a specific topic, thinking out of the box, etc. Must fit in 128gb,… 23 Hugging Face Daily Papers research 20d ago OpenForgeRL: Train Harness-native Agents in Any Environment Abstract Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open… 10 r/LocalLLaMA community 21d ago Laguna s.2.1 updated 2 hours ago. A post to show appreciation for the work they are doing. I'm downloading it again now. So far, the model hasn't performed well with reasoning tasks, but I really appreciate the work being done to fix this.   submitted by   /u/LegacyRemaster [link]   [comments] 30 Hugging Face Daily Papers research 21d ago FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents Abstract Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant… 12 r/LocalLLaMA community 21d ago [BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think! https://preview.redd.it/b7ybs7nqx5fh1.png?width=3440&format=png&auto=webp&s=e6aaaa15cbe59debaae1ebb7fcd708167e86dc35 Hey r/LocalLLaMA ! We are back and we have something really amazing today. Our big 5M samples Reasoning Corpus dataset. This dataset features 5 million rows of: -… 33 r/LocalLLaMA community 21d ago If you're running Laguna S 2.1 and it feels "stupid" or isn't reasoning properly, are you using quantization worse than Q8? I can fit the whole thing in RAM in Q8, and it seems to be outperforming qwen 3.5 122B-A10B Q8 for some things. I've been seeing reports for the last couple days of people saying it feels stupid, but it doesn't seem that way to me.   submitted by   /u/burritoresearch… 35 arXiv — Machine Learning research 21d ago Multimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs arXiv:2607.20517v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and the need for multi-hop reasoning across dispersed evidence. We present… 14 arXiv — Machine Learning research 21d ago When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion arXiv:2607.20543v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study this pass@k inversion: after training, the policy may solve fewer distinct problems… 30 arXiv — Machine Learning research 21d ago Monkey King Bang: A Unified Scientific Multimodal Foundation Model arXiv:2607.20557v1 Announce Type: new Abstract: Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition. Existing systems are either specialised for individual domains or unify scientific… 30 arXiv — Machine Learning research 21d ago CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging arXiv:2607.20561v1 Announce Type: new Abstract: LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter storage and task selection at inference time. Model merging addresses this issue… 29 arXiv — Machine Learning research 21d ago Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation arXiv:2607.20908v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code generation. However, existing RLVR approaches primarily rely on outcome-based… 7 arXiv — NLP / Computation & Language research 21d ago The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning arXiv:2607.20952v1 Announce Type: cross Abstract: Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assumed to function as an internal scratchpad the model actively consults during… 34 arXiv — Machine Learning research 21d ago Training Large Language Models for Self-Explanation Faithfulness arXiv:2607.21090v1 Announce Type: new Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects its internal decision-making process. While existing… 25 arXiv — NLP / Computation & Language research 21d ago What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces arXiv:2607.20425v1 Announce Type: new Abstract: What makes writing "good" remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality. In Study 1, we construct a… 11 arXiv — NLP / Computation & Language research 21d ago Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought arXiv:2607.20427v1 Announce Type: new Abstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box. In this paper, we uncover a fundamental governing principle: MoE routing is not merely selection, but a… 15 arXiv — NLP / Computation & Language research 21d ago Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility arXiv:2607.20440v1 Announce Type: new Abstract: Proprietary large language models (LLMs) entail substantial intellectual and financial investment, making them valuable intellectual property (IP). However, even when deployed via black-box APIs, these models remain vulnerable to… 33 arXiv — NLP / Computation & Language research 21d ago Domyn-Small: A European 10B Reasoning Language Model arXiv:2607.20448v1 Announce Type: new Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an initial pre-training phase on 9 trillion tokens multilingual data, followed by a… 32 arXiv — NLP / Computation & Language research 21d ago THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA arXiv:2607.20459v1 Announce Type: new Abstract: Multi-hop question answering requires retrieving and integrating evidence from multiple contexts. Despite the rapid progress of current research, multi-hop reasoning remains constrained by two persistent limitations: attention… 37 arXiv — NLP / Computation & Language research 21d ago REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning arXiv:2607.20833v1 Announce Type: new Abstract: Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when evidence is sparse, noisy, or in conflict with parametric knowledge.… 16 arXiv — NLP / Computation & Language research 21d ago Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs arXiv:2607.21291v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost. Existing acceleration methods often rely on task-specific fine-tuning or training from… 10 arXiv — NLP / Computation & Language research 21d ago Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models arXiv:2607.21433v1 Announce Type: new Abstract: Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhaust it without reaching a conclusion… 5 arXiv — NLP / Computation & Language research 21d ago WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms arXiv:2607.20638v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal reasoning over digital waveform data remains largely unexplored. Although reasoning over… 18 arXiv — NLP / Computation & Language research 21d ago Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad arXiv:2607.20935v1 Announce Type: cross Abstract: Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and… 35 r/LocalLLaMA community 21d ago Benchmarks: AntLing-3.0-flash a hybrid-reasoning MoE model built for production-scale agents. Now live on OpenRouter, and free to use through August 3, 2026. Hoping they will going openweight soon~   submitted by   /u/niacolhealth [link]   [comments] 28 llama.cpp releases dev-tools 21d ago b10093 Fix DeepSeek4 crafted template ( #25414 ) chat: fix DS4 template to explicitly follow reference behavior Support DeepSeekv4 flag ( drop_reasoning ). fix: hook DS3.2 parser for DS4 as well fix: add tool result reordering fix: post-merge Website: https://llama.app macOS/iOS: macOS… 20 r/LocalLLaMA community 22d ago AI9Stars released G9v3-3B AI9Stars has released G9v3-3B an open weights language model designed to deliver strong reasoning capabilities within a lightweight 3 billion parameter size. It is released under the Apache 2.0 license making it fully open for personal and commercial use The best use case is a… 17 Hugging Face Daily Papers research 22d ago SLPO: Scaling Latent Reasoning via a Surrogate Policy Abstract Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language… 22 Hugging Face Daily Papers research 22d ago SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments Abstract Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to… 37 Hugging Face Daily Papers research 22d ago Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Abstract Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible.… 36 Hugging Face Daily Papers research 22d ago Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization Abstract Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identify the essential… 31 arXiv — Machine Learning research 22d ago Reproducing Recurrent Transformers: The CoTFormer arXiv:2607.19405v1 Announce Type: new Abstract: The CoTFormer architecture formalizes Chain-of-Thought as a form of recurrent latent computation, preserving intermediate states as attendable representations to mimic explicit reasoning traces. In this work, we evaluate CoTFormer… 32 arXiv — Machine Learning research 22d ago REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning arXiv:2607.19450v1 Announce Type: new Abstract: Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast… 16 arXiv — Machine Learning research 22d ago Anatomy of a Sound Neural Reasoner: One-Shot Amortization, First-Pass Poisoning, and Search Inertness in Clue-Rich Completion arXiv:2607.19635v1 Announce Type: new Abstract: Neural solvers are built to deduce, branch, and revise intermediate states. The Lattice Deduction Transformer (LDT) appears to do exactly that. In clue-rich Sudoku, it does not: one forward pass commits essentially the entire grid… 22 arXiv — Machine Learning research 22d ago OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization arXiv:2607.19806v1 Announce Type: new Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors… 4 arXiv — NLP / Computation & Language research 22d ago Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models arXiv:2607.19847v1 Announce Type: cross Abstract: Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show great promise in predicting missing values in tables, by reasoning holistically across rows… 27 arXiv — NLP / Computation & Language research 22d ago When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play arXiv:2607.19523v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect on behavioral diversity in sequential decision-making remains under-explored. We study this question in a controlled… 37 arXiv — NLP / Computation & Language research 22d ago Reference-Free Evaluation of Reasoning in Open-Ended Question Answering arXiv:2607.19678v1 Announce Type: new Abstract: AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for… 37 arXiv — NLP / Computation & Language research 22d ago SLPO: Scaling Latent Reasoning via a Surrogate Policy arXiv:2607.19691v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate… 17 Page 8 of 10 · 500 articles ← Newer Older →