News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow arXiv — NLP / Computation & Language research 5h ago Position: Reasoning is a Learnable Rule-Based Process arXiv:2608.12325v1 Announce Type: cross Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite… 22 arXiv — NLP / Computation & Language research 5h ago LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning arXiv:2608.12321v1 Announce Type: new Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional… 10 arXiv — NLP / Computation & Language research 5h ago What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting arXiv:2608.12322v1 Announce Type: new Abstract: Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence… 6 arXiv — NLP / Computation & Language research 5h ago On Measuring Semantic Preservation in Legal Ontology Learning arXiv:2608.12326v1 Announce Type: new Abstract: Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on… 26 arXiv — NLP / Computation & Language research 5h ago Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching arXiv:2608.12331v1 Announce Type: new Abstract: Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat… 25 arXiv — NLP / Computation & Language research 5h ago Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark arXiv:2608.12343v1 Announce Type: new Abstract: AI chatbots are widely used by students as knowledge sources, yet LLM benchmarks rarely assess interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exams (Matura) in history - three… 12 arXiv — NLP / Computation & Language research 5h ago Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities arXiv:2608.12374v1 Announce Type: new Abstract: Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation… 28 arXiv — NLP / Computation & Language research 5h ago Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models arXiv:2608.12391v1 Announce Type: new Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input… 36 arXiv — NLP / Computation & Language research 5h ago LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning arXiv:2608.12626v1 Announce Type: new Abstract: Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these settings, finite attention resources prevent the model from maintaining strategic coherence… 29 arXiv — NLP / Computation & Language research 5h ago CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives arXiv:2608.12779v1 Announce Type: new Abstract: Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors.… 14 arXiv — NLP / Computation & Language research 5h ago From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options arXiv:2608.12836v1 Announce Type: new Abstract: Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and… 8 arXiv — NLP / Computation & Language research 5h ago HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA arXiv:2608.13004v1 Announce Type: new Abstract: Knowledge-base question answering (KBQA) systems rely on effective retrieval and reasoning mechanisms to generate accurate answers from external knowledge sources. However, developing reliable KBQA systems for low-resource… 19 arXiv — NLP / Computation & Language research 5h ago GEM: A Generative Embedding Model Bridging Reasoning and Retrieval arXiv:2608.13200v1 Announce Type: new Abstract: Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents,… 32 arXiv — NLP / Computation & Language research 5h ago Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits arXiv:2608.13244v1 Announce Type: new Abstract: Local chemical perception and property reasoning are both essential for understanding how molecular structure determines properties. Current LLM-based chemical reasoning methods either receive SMILES/molecular images together with… 36 arXiv — NLP / Computation & Language research 5h ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 5h ago Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability… 20 arXiv — NLP / Computation & Language research 5h ago RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory arXiv:2608.13334v1 Announce Type: new Abstract: LLM-based agents increasingly rely on external memory to support long-horizon reasoning and interaction. However, the main bottleneck is not simply storing past experience, but recovering the right set of evidence when relevant… 38 arXiv — NLP / Computation & Language research 5h ago Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction arXiv:2608.12426v1 Announce Type: cross Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled… 33 arXiv — NLP / Computation & Language research 5h ago MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination arXiv:2608.13476v1 Announce Type: cross Abstract: We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized… 17 r/LocalLLaMA community 13h ago Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release Qwen just released their first 3.8 model. The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting reasoning_effort to xhigh , medium , or low . However, the official template still has some serious problems: You cannot… 14 Hugging Face Daily Papers research 14h ago Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning Abstract Latent Dynamics Reasoning integrates kinematic dynamics in structured latent space to enable video world models that extrapolate physical laws far beyond training distributions with far fewer parameters and faster inference. Generated by thinkingmachines/Inkling-Small… 35 Hugging Face Daily Papers research 21h ago AVA-Encoder: Towards Agent-Native Video Representation Learning Abstract AVA-Encoder learns structured video representations via agentic auto-encoding using knowledge graphs to enable cinematic video generation and reasoning with reduced token usage. Generated by thinkingmachines/Inkling-Small Creative agents still lack an effective way to… 30 Hugging Face Daily Papers research 1d ago AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models Abstract AtlasVLA improves embodied AI by replacing reactive control with proactive reasoning via persistent world-ego memory, enabling robust long-horizon manipulation from a single wrist camera. Generated by thinkingmachines/Inkling-Small While Vision-Language-Action (VLA)… 36 arXiv — Machine Learning research 1d ago PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR arXiv:2608.11368v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this cost by assigning budgets to prompts, rollouts, or tokens according to… 26 arXiv — NLP / Computation & Language research 1d ago LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection arXiv:2608.11691v1 Announce Type: cross Abstract: Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a… 18 arXiv — Machine Learning research 1d ago Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity arXiv:2608.11716v1 Announce Type: new Abstract: Chain of Thought (CoT) lifts the expressive ceiling of bounded-depth Transformers, with characterizations tying the number of CoT steps to circuit complexity classes. What remains largely missing are concrete instantiations with… 6 arXiv — NLP / Computation & Language research 1d ago Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling arXiv:2608.11829v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledge from a stronger teacher model, thereby… 29 arXiv — Machine Learning research 1d ago LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation arXiv:2608.11967v1 Announce Type: new Abstract: Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress,… 31 arXiv — NLP / Computation & Language research 1d ago Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing arXiv:2608.08514v1 Announce Type: cross Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B… 35 arXiv — NLP / Computation & Language research 1d ago Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs arXiv:2608.11573v1 Announce Type: new Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement… 35 arXiv — NLP / Computation & Language research 1d ago AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention arXiv:2608.11758v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic… 34 arXiv — NLP / Computation & Language research 1d ago GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation arXiv:2608.11787v1 Announce Type: new Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision… 6 arXiv — NLP / Computation & Language research 1d ago Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology arXiv:2608.11420v1 Announce Type: cross Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are… 9 arXiv — NLP / Computation & Language research 1d ago Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder arXiv:2608.11650v1 Announce Type: cross Abstract: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This… 8 arXiv — NLP / Computation & Language research 1d ago LookBack: Where and How to Score LVLM Responses via Visual Reference Usage arXiv:2608.11847v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations;… 14 arXiv — NLP / Computation & Language research 1d ago Claim-Level Reliability Assessment for Efficient Test-Time Reasoning arXiv:2608.11994v1 Announce Type: cross Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution… 38 arXiv — NLP / Computation & Language research 1d ago A Reality Check of Language Models as Formalizers on Constraint Satisfaction Problems arXiv:2505.13252v5 Announce Type: replace Abstract: Recent work shows superior performance when using large language models (LLMs) as formalizers instead of as end-to-end solvers for symbolic reasoning problems. Given the problem description, the LLM generates a formal program… 17 Hugging Face Daily Papers research 1d ago AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Abstract Stronger models can build inference-time harnesses that substantially improve weaker models' task performance without parameter updates by offloading reasoning into structured code and routing. Generated by thinkingmachines/Inkling-Small Recent work on distillation… 36 NVIDIA Developer Blog official-blog 1d ago Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72 Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open... 35 r/LocalLLaMA community 1d ago All your reasoning are belong to us   submitted by   /u/indicava [link]   [comments] 20 r/LocalLLaMA community 1d ago Hidden Reasoning from Claude and GPT are Decoded, and it is interesting Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs . check it out, they have published lots of example reasonings. this is very relevant for open soruce; for the… 36 Hugging Face Daily Papers research 1d ago InSight-doc: Agentic Visual Perception for Long-Document Understanding Abstract InSight-doc adaptively allocates visual resolution during reasoning to improve long-document understanding while reducing latency and hallucinations. Generated by thinkingmachines/Inkling-Small Long-document understanding often requires reasoning over many visually rich… 27 Hugging Face Daily Papers research 2d ago 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents Abstract A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning. Generated by thinkingmachines/Inkling-Small We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of… 8 Latent.Space news-outlet 2d ago [AINews] How to steal a Reasoning Trace Speculative Decoding by any other name would distil as sweet 8 arXiv — Machine Learning research 2d ago REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting arXiv:2608.10149v1 Announce Type: new Abstract: Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed… 38 arXiv — Machine Learning research 2d ago MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods… 34 arXiv — NLP / Computation & Language research 2d ago Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth… 30 arXiv — NLP / Computation & Language research 2d ago When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning arXiv:2608.09942v1 Announce Type: new Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds… 21 arXiv — NLP / Computation & Language research 2d ago From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models arXiv:2608.10444v1 Announce Type: new Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively… 21 arXiv — NLP / Computation & Language research 2d ago Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse arXiv:2608.10810v1 Announce Type: new Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks… 5 Page 1 of 10 · 500 articles Older →