News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow arXiv — NLP / Computation & Language research 22d ago TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis arXiv:2607.19794v1 Announce Type: new Abstract: Production LLM-based financial sentiment analysis faces a structural cost trap: most queries are trivially classifiable, yet expensive cloud reasoners process them all, and the bill scales linearly with user count. We present… 25 arXiv — NLP / Computation & Language research 22d ago Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models arXiv:2607.19932v1 Announce Type: new Abstract: Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One… 36 arXiv — NLP / Computation & Language research 22d ago Language-Specific versus Cross-Lingual Knowledge Graphs for Implicit Aspect Identification in Arabic: A Comparative Study of Reasoning and Adaptation Strategies arXiv:2607.20056v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspects that are never named in the text. Implicit identification typically relies on an auxiliary knowledge source (e.g., a… 4 arXiv — NLP / Computation & Language research 22d ago PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference arXiv:2607.20327v1 Announce Type: new Abstract: Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware… 21 arXiv — NLP / Computation & Language research 22d ago HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions arXiv:2607.19398v1 Announce Type: cross Abstract: Multi-entity compositional questions pose significant challenges to existing retrieval-augmented language models. Conventional methods fall into a dilemma: standard RAG lacks dynamic reasoning, traditional Graph-RAG is limited by… 15 arXiv — NLP / Computation & Language research 22d ago Twin Agent: Context Residual Compression for Privilege Separated Agents arXiv:2607.19595v1 Announce Type: cross Abstract: Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untrusted context that manipulate downstream reasoning and tool use. Existing secure-by-design approaches mitigate this… 30 arXiv — NLP / Computation & Language research 22d ago PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity arXiv:2607.20268v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves… 9 arXiv — NLP / Computation & Language research 22d ago NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Knowledge Graph Question Answering arXiv:2602.15353v3 Announce Type: replace Abstract: Large pretrained language models and neural reasoning systems have advanced many natural language tasks, yet they remain challenged by knowledge-intensive queries that require precise, structured multi-hop inference. Knowledge… 31 arXiv — NLP / Computation & Language research 22d ago In-the-Flow Agentic System Optimization for Effective Planning and Tool Use arXiv:2510.05592v2 Announce Type: replace-cross Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full… 21 arXiv — NLP / Computation & Language research 22d ago Prompt Programming for Cultural Bias and Alignment of Large Language Models arXiv:2603.16827v2 Announce Type: replace-cross Abstract: Culture shapes reasoning, values, prioritization, and strategic decision-making, yet large language models (LLMs) often exhibit cultural biases that misalign with target populations. As LLMs are increasingly used for… 11 r/LocalLLaMA community 22d ago Laguna S 2.1 Thinking mode If many people have noticed that there's no reasoning phase in Laguna S 2.1. I noticed the Poolside development team updated the chat template twice in the last 24 hours. There was a bug where, if preserve_thinking was disabled, reasoning wouldn't start at all. However, I don't… 19 r/LocalLLaMA community 22d ago MindControl - llama.cpp fork to guide the reasoning process via injection during sampling The primary driver of this project is that I'd become frustrated with the reasoning behavior of smaller local models such as Qwen3.6-27B (i believe particularly at lower temperatures, and where system prompts are highly specific), their reasoning process is highly unreliable and… 7 r/LocalLLaMA community 23d ago Upstage 'Solar open2' release. performance on par with DeepSeek V4 Flash. Benchmark Solar Open 2 250B-A15B Solar Open 100B 102B-A12B Command A+ 218B-A25B Mistral Medium 3.5 128B dense, high MiMo-V2.5 310B-A15B DeepSeek-V4-Flash 284B-A13B, max Know. & Reasoning MMLU-Pro 86.2 80.4 79.0 81.2 84.6 85.9 GPQA-Diamond 86.3 66.2 75.6 77.5 83.0 88.9 HLE (w/o… 21 Hugging Face Daily Papers research 23d ago H^2SD: Hybrid Hindsight Self-Distillation Abstract Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar outcome reward to an entire… 7 Hugging Face Daily Papers research 23d ago ISO: An RLVR-Native Optimization Stack Abstract Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et… 36 arXiv — Machine Learning research 23d ago PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects arXiv:2607.18777v1 Announce Type: new Abstract: Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PertReason, a knowledge-grounded benchmark and framework suite for… 9 arXiv — NLP / Computation & Language research 23d ago H$^2$SD: Hybrid Hindsight Self-Distillation arXiv:2607.18955v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a… 24 arXiv — NLP / Computation & Language research 23d ago Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains arXiv:2607.18438v1 Announce Type: new Abstract: Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The… 37 arXiv — NLP / Computation & Language research 23d ago Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives arXiv:2607.18443v1 Announce Type: new Abstract: Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics… 25 arXiv — NLP / Computation & Language research 23d ago Reasoning Fine-Tuning Induces Persistent Latent Policy States arXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning… 15 arXiv — NLP / Computation & Language research 23d ago LatentMT: Machine Translation with Latent Reasoning arXiv:2607.18618v1 Announce Type: new Abstract: Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT): instead of increasing parameter count or emitting explicit chain-of-thought tokens, they spend additional recurrent… 24 arXiv — NLP / Computation & Language research 23d ago CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness arXiv:2607.18820v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large language models (LLMs), yet the generated reasoning may not faithfully support the final answer. We study this problem… 22 arXiv — NLP / Computation & Language research 23d ago Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM arXiv:2607.18915v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex problems via long, multi-step reasoning. However, as… 32 arXiv — NLP / Computation & Language research 23d ago DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning arXiv:2607.19088v1 Announce Type: new Abstract: Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We… 19 arXiv — NLP / Computation & Language research 23d ago Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning arXiv:2607.19181v1 Announce Type: new Abstract: Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity of legal language and the high level of precision it requires. The recent emergence of… 24 arXiv — NLP / Computation & Language research 23d ago MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams arXiv:2607.19201v1 Announce Type: new Abstract: Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent,… 13 arXiv — NLP / Computation & Language research 23d ago The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation arXiv:2607.19226v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the… 9 arXiv — NLP / Computation & Language research 23d ago MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings arXiv:2607.19235v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in… 8 arXiv — NLP / Computation & Language research 23d ago Selective State-Space Adaptation and Retrieval for Language Model Reasoning arXiv:2607.19326v1 Announce Type: new Abstract: Low-rank adaptation introduces a static learned update applied identically to every input. The update provides task-level adaptation but does not explicitly represent token-level or instance-level state variation. A family of… 25 arXiv — NLP / Computation & Language research 23d ago Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning arXiv:2607.19345v1 Announce Type: new Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical… 16 arXiv — NLP / Computation & Language research 23d ago MUX: Continuous Reasoning via Multiplexed Tokens arXiv:2607.18264v1 Announce Type: cross Abstract: Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is computationally bottlenecked: each reasoning step conveys only a single subword, and many… 24 arXiv — NLP / Computation & Language research 23d ago Supra Cognitive Modes: A Routed Architecture for Agent Memory arXiv:2607.19096v1 Announce Type: cross Abstract: Agent-memory workloads mix direct factual lookup, relation-chain and current-state reasoning, and broad synthesis over long histories. We describe Supra Cognitive Modes (SCM), an architecture that maps explicit or automatically… 4 arXiv — NLP / Computation & Language research 23d ago Agents in the Wild: Where Research Meets Deployment arXiv:2607.19336v1 Announce Type: cross Abstract: Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale… 19 arXiv — NLP / Computation & Language research 23d ago TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models arXiv:2506.18421v3 Announce Type: replace Abstract: The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden… 4 arXiv — NLP / Computation & Language research 23d ago Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation arXiv:2602.14469v4 Announce Type: replace Abstract: Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but answer-visible generation can justify a pre-committed answer rather than derive it. This post-hoc rationalization creates a… 8 Hugging Face Daily Papers research 23d ago ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning Abstract Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely… 27 r/LocalLLaMA community 23d ago Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency) Kind of unexpected. Happy for Gemma-4/Google, big win for us, LocalLLMers. Yet Qwen3.6 still does better in Hermes than Gemma-4 somehow. We need Gemma-4.1 fine-tuned on Agentic-tasks. That would be killer.   submitted by   /u/JLeonsarmiento [link]   [comments] 16 LangChain releases dev-tools 23d ago langchain-anthropic==1.5.0 Changes since langchain-anthropic==1.4.8 release(anthropic): 1.5.0 ( #38985 ) feat(core): add reasoning_effort as a standard chat model parameter ( #38887 ) fix(anthropic): add advisor_ prefix to builtin tool recognition ( #38686 ) chore(deps): refresh lockfiles ( #38746 )… 24 r/LocalLLaMA community 24d ago New Model: Nanbeige4.2-3B (Looped Transformer, outperforms 4x size) https://huggingface.co/Nanbeige/Nanbeige4.2-3B Nanbeige4.2-3B is a compact agentic model built on Nanbeige4.2-3B-Base , designed to combine strong agentic behavior with broad reasoning and alignment capabilities. Its Looped Transformer architecture reuses the transformer layers… 33 LangChain releases dev-tools 24d ago langchain-fireworks==1.5.0 Changes since langchain-fireworks==1.4.4 release(fireworks): 1.5.0 ( #38986 ) feat(core): add reasoning_effort as a standard chat model parameter ( #38887 ) chore: bump langsmith from 0.10.2 to 0.10.6 in /libs/partners/fireworks ( #38914 ) fix: patch Dependabot dependency… 23 LangChain releases dev-tools 24d ago langchain-xai==1.3.0 Changes since langchain-xai==1.2.2 release(xai): 1.3.0 ( #38984 ) feat(core): add reasoning_effort as a standard chat model parameter ( #38887 ) fix: patch Dependabot dependency vulnerabilities ( #38853 ) chore: bump langsmith from 0.9.5 to 0.10.2 in /libs/partners/xai ( #38832… 29 LangChain releases dev-tools 24d ago langchain-openai==1.4.0 Changes since langchain-openai==1.3.5 release(openai): 1.4.0 ( #38983 ) chore: bump pillow from 12.2.0 to 12.3.0 in /libs/partners/openai ( #38999 ) feat(core): add reasoning_effort as a standard chat model parameter ( #38887 ) chore(model-profiles): refresh model profile data (… 22 arXiv — Machine Learning research 24d ago Quantizing Recursive Reasoning Models arXiv:2607.16237v1 Announce Type: new Abstract: Recursive reasoning models solve hard puzzles by applying compact, weight-tied blocks over many refinement steps. Because these blocks are reused many times, quantizing them creates a unique dynamical problem: the quantization… 31 arXiv — Machine Learning research 24d ago Autonomous mechanistic discovery of colorectal cancer vulnerabilities via multi-scale AI swarms arXiv:2607.16262v1 Announce Type: new Abstract: The acceleration of automated scientific discovery has been fundamentally bottlenecked by the epistemic gap between the semantic reasoning of large language models (LLMs) and the deterministic physics of mammalian biology. While… 32 arXiv — Machine Learning research 24d ago CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation arXiv:2607.16955v1 Announce Type: new Abstract: On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to… 16 arXiv — Machine Learning research 24d ago Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning arXiv:2607.17047v1 Announce Type: new Abstract: LLM constraint reasoners are often evaluated near the random-SAT phase transition, confounding density and solver hardness. We test instance-level transfer while near-matching clause density. At aligned size bins, with near-matched… 19 arXiv — Machine Learning research 24d ago Distilled Reinforcement Learning for LLM Post-training arXiv:2607.17247v1 Announce Type: new Abstract: Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL… 21 arXiv — NLP / Computation & Language research 24d ago Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM arXiv:2607.16451v1 Announce Type: new Abstract: Chat models sometimes commit to an answer and then produce reasoning that justifies it rather than deriving it -- even when the answer contradicts a task premise. We study a minimal probe: "I want to wash my car. The car wash is… 12 arXiv — NLP / Computation & Language research 24d ago NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning arXiv:2607.16603v1 Announce Type: new Abstract: This paper presents the methodologies and results of the NOWJ team's participation across all five tasks of the COLIEE 2026 competition. For Task 1 (Legal Case Retrieval), we propose a four-stage pipeline comprising candidate… 19 arXiv — NLP / Computation & Language research 24d ago Trace-Based On-Policy Distillation for Masked Diffusion Language Models arXiv:2607.16872v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but… 14 Page 9 of 10 · 500 articles ← Newer Older →