News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow arXiv — NLP / Computation & Language research 5d ago Beyond Static Charts: Can Language and Vision Language Models Generate Interactive Data Visualization Interfaces? arXiv:2609.26208v1 Announce Type: new Abstract: Data visualization is central to analytical reasoning, but real-world analysis increasingly requires language-driven interactive interfaces rather than static charts. Although recent large language and vision language models… 25 arXiv — NLP / Computation & Language research 5d ago Spoken Language Models that Think Aloud arXiv:2609.26488v1 Announce Type: new Abstract: While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial "think-then-speak" paradigm,… 32 arXiv — NLP / Computation & Language research 5d ago Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation arXiv:2609.26536v1 Announce Type: new Abstract: In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this,… 24 arXiv — NLP / Computation & Language research 5d ago Semantic Abstraction for Natural Language Inference: a Methodological Framework for Discovering and Compensating Semantic Knowledge and Reasoning Gaps in Large Language Models arXiv:2609.26610v1 Announce Type: new Abstract: Despite their outstanding performance on many NLP tasks, LLMs face serious challenges related to semantic abstraction. In this study, we are interested in understanding how LLMs leverage abstract semantic knowledge in natural… 11 arXiv — NLP / Computation & Language research 5d ago Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models arXiv:2609.26637v1 Announce Type: new Abstract: The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool… 35 arXiv — NLP / Computation & Language research 5d ago Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning arXiv:2609.26704v1 Announce Type: new Abstract: Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such… 11 The Information — AI news-outlet 5d ago Anthropic releases cheaper model in first launch since slowdown calls Anthropic announced the release of its newest model, Claude Opus 5.5 , saying it performs as well or better than Anthropic’s previous most powerful models, Fable 5.1 and Mythos 5.1, in domains including coding, reasoning, business workflows, and safety. At the same time, it will… 18 arXiv — Machine Learning research 6d ago A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents arXiv:2609.22194v1 Announce Type: new Abstract: Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT)… 18 arXiv — Machine Learning research 6d ago Dissecting Hierarchical Reasoning Models: A Mechanistic Study arXiv:2609.22197v1 Announce Type: new Abstract: We study Hierarchical Reasoning Model (HRM), a representative hierarchical Transformer-based latent reasoning model with many variants, on Sudoku, Maze, and ARC-AGI-2. We mechanistically understand how HRM reasons and what… 21 arXiv — Machine Learning research 6d ago Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks arXiv:2609.22237v1 Announce Type: new Abstract: Merging low-rank adapters (LoRAs) promises to eliminate the overhead of swapping task-specific weights at inference time. However, existing merging methods assume every layer needs the same rank budget. Further, some methods assume… 38 arXiv — Machine Learning research 6d ago Causilo Technical Report arXiv:2609.22866v1 Announce Type: new Abstract: We introduce Causilo, a tabular foundation model (TFM) that combines frontier predictive performance with exceptionally fast inference. On TabArena, Causilo achieves 1785.4 Elo, at a median inference time of 0.10 seconds per 1K… 25 arXiv — Machine Learning research 6d ago Towards Full Pipeline FP8 Reinforcement Learning for LLMs arXiv:2609.22870v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8… 19 arXiv — NLP / Computation & Language research 6d ago Balancing Reasoning and Hardware Constraints in RAG Pipelines for Ukrainian Multi-Domain Document Understanding arXiv:2609.22124v1 Announce Type: new Abstract: This paper describes the system submitted to the UNLP 2026 Shared Task on Multi-Domain Document Understanding. The challenge required extracting precise answers, document IDs, and page numbers from a diverse corpus of Ukrainian PDF… 18 arXiv — NLP / Computation & Language research 6d ago Quantifying Hidden Salt for Precision Healthcare: Sodium Assessment via Joint-Factor Retrieval and Chain-of-Thought Inference arXiv:2609.22171v1 Announce Type: new Abstract: Precision healthcare, particularly for conditions like hypertension and cardiovascular disease, necessitates monitoring of dietary sodium intake. However, tracking this is hindered by the prevalence of hidden salt in cooking, such… 17 arXiv — NLP / Computation & Language research 6d ago Assessing Adversarial Robustness of Latent Reasoning Models arXiv:2609.22228v1 Announce Type: new Abstract: Large language models increasingly rely on long chain-of-thought (CoT) trajectories for complex reasoning, but autoregressive generation brings substantial memory and inference costs. Latent reasoning models (LRMs) offer a more… 37 arXiv — NLP / Computation & Language research 6d ago Do Chess Explanations Reflect Model Decisions? Behavioral and Token-Level Tests of LLM Reasoning Faithfulness arXiv:2609.22245v1 Announce Type: new Abstract: Large language models can produce fluent explanations for chess moves, but plausible language does not necessarily reflect the reasoning behind a decision. We study this question in chess, where the board state is fully observable,… 29 arXiv — NLP / Computation & Language research 6d ago The Corroboration Illusion: When More News Makes LLM Forecasts Less True arXiv:2609.22246v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to forecast real-world events by retrieving and reasoning over news. We show that this dependence on an open, crawlable news corpus creates a new attack surface: an adversary who… 12 arXiv — NLP / Computation & Language research 6d ago Correct Diagnosis, Better Feedback: A Symbolic-Verifier for Faithful LLM Tutoring Feedback in Logic Proofs arXiv:2609.22553v1 Announce Type: new Abstract: Effective LLM tutoring depends on correctly identifying the specific error in a student's reasoning before generating feedback. We study this problem in propositional-logic proof tutoring, where student actions can be checked… 22 arXiv — NLP / Computation & Language research 6d ago COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural… 7 arXiv — NLP / Computation & Language research 6d ago LLaDA-PRM: A Bidirectional Step-Level Reasoning Evaluator arXiv:2609.22700v1 Announce Type: new Abstract: Step-level reasoning evaluators are commonly based on autoregressive language models, whose causal attention restricts each step representation to the problem, previous steps, and the current step. Yet, when the complete solution… 12 r/LocalLLaMA community 6d ago 50+ Hours and 100M+ Tokens Later, Open Source Autonomous Agent is GETTING CLOSER at Solving an Open Math problem This experiment is live, you can inspect all the internal reasoning, memories, attempts here: https://artificium-covering-experiment.gr.bio/ The problem that the agent is trying to solve is a covering design problem: https://en.wikipedia.org/wiki/Covering_design Known as… 19 r/LocalLLaMA community 6d ago [Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff" https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=webp&s=74765dbd409f4c221640f9f6000a685f6fdbb242 Spent weekend benchmarking the Splash engine (by Incoai) and extending its architecture to native 8-bit on Apple Silicon (M5 Pro, 64 GB unified memory). Splash is… 10 arXiv — Machine Learning research 7d ago M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection arXiv:2609.21164v1 Announce Type: new Abstract: Integrating diverse data modalities --- such as clinical notes, laboratory results, and medical imaging --- is essential for advancing clinical decision-making. While Large Language Models (LLMs) have shown remarkable performance… 25 arXiv — NLP / Computation & Language research 7d ago HERMES: Contrast-Aware Knowledge Graph Reasoning from Clinical Notes for Patient Outcome Prediction arXiv:2609.20825v1 Announce Type: new Abstract: Clinical predictive models often rely on structured Electronic Health Record data, such as time-series and procedure codes. While recent approaches have begun leveraging unstructured clinical notes, they typically encode them as… 20 arXiv — NLP / Computation & Language research 7d ago VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering arXiv:2609.20843v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) enables models to answer natural-language questions through structured graph reasoning and has achieved substantial progress across many benchmarks and applications. Recently, multimodal… 28 arXiv — NLP / Computation & Language research 7d ago Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models arXiv:2609.20846v1 Announce Type: new Abstract: While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from… 4 arXiv — Machine Learning research 7d ago Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework arXiv:2609.20973v1 Announce Type: cross Abstract: Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect… 35 arXiv — NLP / Computation & Language research 7d ago Enhancing Audio Reasoning via Semantic Summary Prediction arXiv:2609.20849v1 Announce Type: new Abstract: Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning… 9 arXiv — NLP / Computation & Language research 7d ago When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces arXiv:2609.21247v1 Announce Type: new Abstract: Large Reasoning Models increasingly use intermediate traces for machine translation, but it remains unclear when such reasoning helps or hurts. We analyze reasoning traces across models, languages, domains, and datasets, focusing… 17 arXiv — NLP / Computation & Language research 7d ago MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance arXiv:2609.21554v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks.… 14 arXiv — NLP / Computation & Language research 7d ago When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap arXiv:2609.21662v1 Announce Type: new Abstract: Activation steering has become a widely used approach for controlling language models during explicit chain-of-thought (CoT) reasoning, motivating its extension to latent CoT. However, we find that steering continuous thoughts… 36 r/LocalLLaMA community 7d ago You can use any LLM just like JEV You can simply run any GGUF with llama.cpp with n_predict=1 and n_probs=10, disable reasoning, and prompt it such as "If the following email is spam, respond with 1, if not spam, respond with 0. Do not respond with anything other than 1 or 0. Email: ...." And that is it! It… 35 Vercel — AI dev-tools 7d ago Grok 4.7 now available and 40% off on AI Gateway, fx, and eve Grok 4.7 from SpaceXAI is now available on AI Gateway and 40% off through September 27. The discount applies automatically when you call spacexai/grok-4.7 . Grok 4.7 has a 500K token context window and supports low, medium, high, and xhigh reasoning levels, giving you control… 31 Vercel — AI dev-tools 7d ago MiMo V2.6 models now available on AI Gateway MiMo V2.6 Pro , MiMo V2.6 Flash , and MiMo V2.6 Pro UltraSpeed from Xiaomi are now available on AI Gateway . MiMo V2.6 combines coding, reasoning, and tool use with native text, image, audio, and video understanding. Its 1M token context supports long repositories, tool traces,… 14 arXiv — Machine Learning research 10d ago Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling arXiv:2609.19499v1 Announce Type: new Abstract: Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candidates, N.… 38 arXiv — NLP / Computation & Language research 10d ago Learn Your Own Thoughts: Abstract Token Curriculum arXiv:2609.19717v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on… 29 arXiv — NLP / Computation & Language research 10d ago Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning arXiv:2609.19878v1 Announce Type: cross Abstract: Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the… 6 arXiv — NLP / Computation & Language research 10d ago Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry arXiv:2609.19154v1 Announce Type: new Abstract: While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training… 4 arXiv — NLP / Computation & Language research 10d ago Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes arXiv:2609.19156v1 Announce Type: new Abstract: Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning… 30 arXiv — NLP / Computation & Language research 10d ago To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives arXiv:2609.19167v1 Announce Type: new Abstract: As AI systems evolve into personalized digital companions, a central capability is reasoning over a user's long-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences.… 5 arXiv — NLP / Computation & Language research 10d ago CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives arXiv:2609.19585v1 Announce Type: new Abstract: In mental health care, reasoning over patient journeys is a key task for clinicians. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate… 37 arXiv — NLP / Computation & Language research 10d ago Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction arXiv:2609.19606v1 Announce Type: new Abstract: This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model's chain-of-thought entropy trajectory predicts whether the final answer is correct, while the… 15 arXiv — NLP / Computation & Language research 10d ago PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces arXiv:2609.19883v1 Announce Type: new Abstract: Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained,… 8 arXiv — NLP / Computation & Language research 10d ago Schema-Anchored Latent Reasoning for Semantic Parsing-Based Knowledge Base Question Answering arXiv:2609.20398v1 Announce Type: new Abstract: Semantic parsing (SP)-based knowledge base question answering aims to answer natural language questions by generating executable logical forms (LFs) over knowledge bases (KBs). When applying Large Language Models (LLMs) to this… 31 arXiv — NLP / Computation & Language research 10d ago On-Demand Attention: Language Models Know When to Recall arXiv:2609.20734v1 Announce Type: new Abstract: Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained… 30 arXiv — NLP / Computation & Language research 10d ago CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning arXiv:2609.19189v1 Announce Type: cross Abstract: Design verification remains one of the most resource-intensive stages of hardware development, often consuming up to 70% of the total design effort. While recent work has explored using Large Language Models (LLMs) to automate… 33 arXiv — NLP / Computation & Language research 10d ago Think Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking arXiv:2609.20131v1 Announce Type: cross Abstract: Reasoning-based reranking with Large Language Models (LLMs) has shown promising improvements in text ranking. However, current methods predominantly rely on a single reasoning trajectory, resulting in rankings that are… 29 arXiv — NLP / Computation & Language research 10d ago Language-model groups overstate consensus when replaying human deliberation on a reasoning task arXiv:2609.20543v1 Announce Type: cross Abstract: Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model… 34 Hugging Face Daily Papers research 11d ago Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Abstract Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators,… 16 arXiv — Machine Learning research 11d ago When the Gradient Sees Rank: Provable Necessity, Causal Recruitment, and Composition in Trained Matrix Memories arXiv:2609.17594v1 Announce Type: new Abstract: Can gradient-based training learn the rank needed to store and compose associations in a matrix memory? In our earlier study, we used a matrix-augmented reasoner on a task that admits a rank-1 solution, leaving this question open.… 9 Page 2 of 10 · 500 articles ← Newer Older →