News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow Hugging Face Daily Papers research 14d ago β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Abstract On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD… 27 Hugging Face Daily Papers research 14d ago See2Think: Do Multimodal Models Really Use Intermediate Visual States? Abstract Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage… 9 arXiv — Machine Learning research 14d ago Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models arXiv:2607.27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral… 14 arXiv — Machine Learning research 14d ago Policy Gradient Steering: Interventions from Behavioral Objectives arXiv:2607.27574v1 Announce Type: new Abstract: Activation steering has emerged in large language models as a lightweight alternative for dynamically changing a model's behavior at inference time. However, we show that existing steering methods fail to steer even a simple policy… 4 arXiv — Machine Learning research 14d ago Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, as LRMs personalization for downstream users takes center stage, the demand for… 26 arXiv — Machine Learning research 14d ago Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning arXiv:2607.27610v1 Announce Type: new Abstract: Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current… 20 arXiv — Machine Learning research 14d ago Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold arXiv:2607.27770v1 Announce Type: new Abstract: A single reinforcement-learning run can produce a strong reasoner yet an incomplete teacher: it often amplifies only a subset of the valid solution modes. We argue that reinforcement learning (RL)-trained policies should therefore… 24 arXiv — Machine Learning research 14d ago LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv:2607.27787v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is… 4 arXiv — Machine Learning research 14d ago Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control arXiv:2607.27914v1 Announce Type: new Abstract: Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model predictive control and reinforcement learning are widely studied, but… 31 arXiv — Machine Learning research 14d ago ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute… 18 arXiv — Machine Learning research 14d ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger arXiv:2607.28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate… 18 arXiv — NLP / Computation & Language research 14d ago Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different… 26 arXiv — NLP / Computation & Language research 14d ago Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or… 29 arXiv — NLP / Computation & Language research 14d ago Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation arXiv:2607.27783v1 Announce Type: new Abstract: Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously… 12 arXiv — NLP / Computation & Language research 14d ago Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory arXiv:2607.27919v1 Announce Type: new Abstract: Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only… 13 arXiv — NLP / Computation & Language research 14d ago LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models arXiv:2607.28077v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals.… 27 arXiv — NLP / Computation & Language research 14d ago Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a… 32 arXiv — NLP / Computation & Language research 14d ago RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning arXiv:2607.28156v1 Announce Type: new Abstract: Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When… 18 arXiv — NLP / Computation & Language research 14d ago Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the… 35 arXiv — NLP / Computation & Language research 14d ago Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning arXiv:2607.28478v1 Announce Type: new Abstract: As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a… 14 arXiv — NLP / Computation & Language research 14d ago ORCA-bench: How Ready Are Language Model Agents for Oncall? arXiv:2607.28545v1 Announce Type: new Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports,… 18 arXiv — NLP / Computation & Language research 14d ago ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning arXiv:2607.27631v1 Announce Type: cross Abstract: Reinforcement learning has emerged as an effective paradigm for enhancing the mathematical reasoning capabilities of large language models. Among existing policy optimization methods, Proximal Policy Optimization (PPO) remains… 16 arXiv — NLP / Computation & Language research 14d ago PCAP-LM: An LLM-Native Text Representation for TLS Bulk Traffic Analysis arXiv:2607.28100v1 Announce Type: cross Abstract: Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equivalents are prohibitively verbose, overflowing LLM context windows by two orders… 37 arXiv — NLP / Computation & Language research 14d ago SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute arXiv:2607.28457v1 Announce Type: cross Abstract: Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR),… 16 arXiv — NLP / Computation & Language research 14d ago Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation arXiv:2607.28495v1 Announce Type: cross Abstract: Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix. We audit that assumption at a whole reasoning-stage… 8 arXiv — NLP / Computation & Language research 14d ago OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation,… 17 Hugging Face Daily Papers research 14d ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Abstract Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct… 37 Hugging Face Daily Papers research 14d ago Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Abstract Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In… 37 Hugging Face Daily Papers research 14d ago Beacon: Knowing When and How to Perform Agentic Visual Reasoning Abstract The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic… 14 Hugging Face Daily Papers research 14d ago Metis: Memory Foundation Model Abstract Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external… 7 Hugging Face Daily Papers research 14d ago VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Abstract Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought… 11 Hugging Face Daily Papers research 14d ago SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them Abstract Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason… 14 r/LocalLLaMA community 14d ago Making a synthetic dataset for fine-tuning I've been thinking about building a pipeline to generate reasoning training data for LLMs, but I want to avoid the common failure mode of synthetic data where you just generate the same template with different numbers. The rough idea: Generate an abstract reasoning task (logic,… 9 r/LocalLLaMA community 15d ago Benchmarked: MindControl for Llama.cpp I recently shared the original MindControl PoC (and on github ) - sampler-level guided reasoning budgets for llama.cpp, nudging the model with self-aware statements about its own thinking budget instead of just hard-truncating it. We received some great feedback, and the most… 27 Hugging Face Daily Papers research 15d ago CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation Abstract On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii)… 18 arXiv — Machine Learning research 15d ago From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models arXiv:2607.26504v1 Announce Type: new Abstract: Many discrete reasoning tasks, such as code generation, are inherently non-causal: programmers move between high-level structure and local details, a process we call any-order inference. For autoregressive language models, which… 12 arXiv — Machine Learning research 15d ago ReCo: Reweighting GRPO Against Distributional Concentration arXiv:2607.26862v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in… 24 arXiv — Machine Learning research 15d ago Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models arXiv:2607.26922v1 Announce Type: new Abstract: Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial models. In this study, we investigate Parishad, a structured multi-agent system… 22 arXiv — NLP / Computation & Language research 15d ago Steering Instruction Hierarchies at Inference Time arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs from users or tools. Yet frontier LLMs often… 33 arXiv — NLP / Computation & Language research 15d ago Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification arXiv:2607.26397v1 Announce Type: new Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC… 38 arXiv — NLP / Computation & Language research 15d ago ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models arXiv:2607.26455v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing… 37 arXiv — NLP / Computation & Language research 15d ago CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG arXiv:2607.26470v1 Announce Type: new Abstract: Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG systems typically represent conversational memory as raw dialogue history,… 32 arXiv — NLP / Computation & Language research 15d ago Metis: Memory Foundation Model arXiv:2607.26760v1 Announce Type: new Abstract: Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still… 35 arXiv — NLP / Computation & Language research 15d ago SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning arXiv:2607.26873v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where… 6 arXiv — NLP / Computation & Language research 15d ago Dual-Path LLM Reasoning for Multimodal Few-Shot Knowledge Graph Completion arXiv:2607.26909v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing facts in knowledge graphs (KGs), thereby improving their completeness and supporting downstream intelligent applications. However, emerging entities and relations in real-world… 36 arXiv — NLP / Computation & Language research 15d ago Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? arXiv:2607.26952v1 Announce Type: new Abstract: We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers… 33 arXiv — NLP / Computation & Language research 15d ago TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days… 29 arXiv — NLP / Computation & Language research 15d ago Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this… 35 arXiv — NLP / Computation & Language research 15d ago ARC-Encoder: learning compressed text representations for large language models arXiv:2510.20535v2 Announce Type: replace Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Context compression techniques can reduce these costs, but the most effective… 8 arXiv — NLP / Computation & Language research 15d ago LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs arXiv:2603.08286v2 Announce Type: replace Abstract: Legal argument mining aims to identify and classify the functional components of judicial reasoning, such as facts, issues, rules, analysis, and conclusions. Progress in this area is limited by the lack of large-scale,… 22 Page 6 of 10 · 500 articles ← Newer Older →