arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 4d ago
TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation
arXiv:2608.06396v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relevant experts during downstream adaptation. Yet current approaches have two…
18 -
arXiv — NLP / Computation & Language research 4d ago
Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
arXiv:2608.06409v1 Announce Type: new Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a…
4 -
arXiv — NLP / Computation & Language research 4d ago
NTDH: Complex Reasoning for Comprehensive Affective Analysis
arXiv:2608.06425v1 Announce Type: new Abstract: Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring conflicting cues to…
9 -
arXiv — NLP / Computation & Language research 4d ago
Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models
arXiv:2608.06429v1 Announce Type: new Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce…
28 -
arXiv — NLP / Computation & Language research 4d ago
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
arXiv:2608.06485v1 Announce Type: new Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key…
21 -
arXiv — NLP / Computation & Language research 4d ago
ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives
arXiv:2608.06495v1 Announce Type: new Abstract: Construction accident narratives contain rich causal information, but the evidence is often implicit, long-span, and distributed. We introduce ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA…
5 -
arXiv — NLP / Computation & Language research 4d ago
Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand
arXiv:2608.06506v1 Announce Type: new Abstract: Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while…
16 -
arXiv — NLP / Computation & Language research 4d ago
GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization
arXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into a privacy risk. Adversarial anonymization defends against this by rewriting a…
36 -
arXiv — NLP / Computation & Language research 4d ago
Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models
arXiv:2608.06529v1 Announce Type: new Abstract: Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs). Existing formulations build this blend with linear interpolation (LERP) in the raw embedding space, which implicitly treats that space as…
31 -
arXiv — NLP / Computation & Language research 4d ago
Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding
arXiv:2608.06532v1 Announce Type: new Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometimes one the model produced without reading the…
21 -
arXiv — NLP / Computation & Language research 4d ago
Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance
arXiv:2608.06539v1 Announce Type: new Abstract: False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than reinforcing false assumptions. The common approach reduces the task to prompting LLMs…
19 -
arXiv — NLP / Computation & Language research 4d ago
TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade
arXiv:2608.06549v1 Announce Type: new Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where…
11 -
arXiv — NLP / Computation & Language research 4d ago
Beyond "AI Language": The case for the idiolectal nature of LLM output
arXiv:2608.06589v1 Announce Type: new Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human…
10 -
arXiv — NLP / Computation & Language research 4d ago
Pre-Inference Routing for Cost-Efficient Document Field Extraction
arXiv:2608.06607v1 Announce Type: new Abstract: Most document-extraction systems use a single model for all documents. This is simple but can be costly for easy cases and less effective for difficult ones. We examine whether we can predict a document's difficulty before…
25 -
arXiv — NLP / Computation & Language research 4d ago
Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
arXiv:2608.06614v1 Announce Type: new Abstract: Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and…
36 -
arXiv — NLP / Computation & Language research 4d ago
Discovering Conceptual Metaphors Across Topics and Media Types
arXiv:2608.06652v1 Announce Type: new Abstract: Conceptual metaphors guide our thinking and actions by allowing us to reason about more abstract experiences (e.g., paying taxes) in terms of more concrete or embodied experiences (e.g., carrying a physical load) (Lakoff and…
10 -
arXiv — NLP / Computation & Language research 4d ago
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents
arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done,…
9 -
arXiv — NLP / Computation & Language research 4d ago
TA-RAG: Tone Awareness as a Design Imperative for Retrieval-Augmented Generation
arXiv:2608.06672v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a robust architecture for grounding large language models (LLMs) in trusted knowledge. However, standard RAG systems exhibit a structural limitation: retrieved documents carry their…
34 -
arXiv — NLP / Computation & Language research 4d ago
Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response…
13 -
arXiv — NLP / Computation & Language research 4d ago
Progressive Content Refinement with Decaying Reward Joint LinUCB
arXiv:2608.06750v1 Announce Type: new Abstract: Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the…
5 -
arXiv — NLP / Computation & Language research 4d ago
Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control
arXiv:2608.06758v1 Announce Type: new Abstract: We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16. The central goal of this work is structured document parsing via capability…
17 -
arXiv — NLP / Computation & Language research 4d ago
Multi-Perspective Triad Interaction Graph Neural Network for Cognitive Distortion Detection
arXiv:2608.06785v1 Announce Type: new Abstract: Cognitive distortion detection is a key task in computational mental health, yet existing approaches often overlook the psychological structure of distorted thoughts. We propose MTI-GNN (Multi-Perspective Triad Interaction Graph…
17 -
arXiv — NLP / Computation & Language research 4d ago
Simple-OPD: Demystifying Warm-up for On-policy Distillation
arXiv:2608.06802v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage before OPD. In this paper, we demystify warm-up for…
37 -
arXiv — NLP / Computation & Language research 4d ago
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
arXiv:2608.06819v1 Announce Type: new Abstract: Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods either use LLM-generated intervention tokens or rank candidates with the LLM's…
8 -
arXiv — NLP / Computation & Language research 4d ago
Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry
arXiv:2608.06849v1 Announce Type: new Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime…
27 -
arXiv — NLP / Computation & Language research 4d ago
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
arXiv:2608.06867v1 Announce Type: new Abstract: No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair…
4 -
arXiv — NLP / Computation & Language research 4d ago
Georeferencing Non-Gazetteered Place Names using Biological Specimen Records
arXiv:2608.06884v1 Announce Type: new Abstract: Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing biodiversity information about regional landscapes as they were recorded at different times.…
19 -
arXiv — NLP / Computation & Language research 4d ago
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests
arXiv:2608.06908v1 Announce Type: new Abstract: We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI…
27 -
arXiv — NLP / Computation & Language research 4d ago
Ask-E: An Environment for Calibrated Question Generation
arXiv:2608.06933v1 Announce Type: new Abstract: Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to probe model limits and generalize beyond existing…
22 -
arXiv — NLP / Computation & Language research 4d ago
Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
arXiv:2608.06953v1 Announce Type: new Abstract: Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask what governs whether it does. Matched notes carry the…
20 -
arXiv — NLP / Computation & Language research 4d ago
Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs
arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar…
36 -
arXiv — NLP / Computation & Language research 4d ago
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue
arXiv:2608.06975v1 Announce Type: new Abstract: Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally…
14 -
arXiv — NLP / Computation & Language research 4d ago
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
arXiv:2608.06977v1 Announce Type: new Abstract: It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases…
5 -
arXiv — NLP / Computation & Language research 4d ago
GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base
arXiv:2608.06992v1 Announce Type: new Abstract: We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated…
35 -
arXiv — NLP / Computation & Language research 4d ago
Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?
arXiv:2608.07006v1 Announce Type: new Abstract: Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be passed to the generator. We show that this…
35 -
arXiv — NLP / Computation & Language research 4d ago
An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
arXiv:2608.07023v1 Announce Type: new Abstract: Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we…
30 -
arXiv — NLP / Computation & Language research 4d ago
HNR-DAC: Hard-Negative Reranking and Distribution-Aligned Classification for Scientific Claim Verification
arXiv:2608.07204v1 Announce Type: new Abstract: Scientific claim verification over a cited paper requires predicting the claim--paper relation and identifying the paragraphs that justify that prediction. This setting poses two linked challenges: within-paper distractors often…
10 -
arXiv — NLP / Computation & Language research 4d ago
Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes
arXiv:2608.07208v1 Announce Type: new Abstract: Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it.…
11 -
arXiv — NLP / Computation & Language research 4d ago
From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL
arXiv:2608.07213v1 Announce Type: new Abstract: Test-time scaling can correct difficult text-to-SQL queries, but the extra computation is normally discarded after each answer. Systems increasingly retain verified repair episodes, yet evaluations still report one end-to-end…
18 -
arXiv — NLP / Computation & Language research 4d ago
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
arXiv:2608.07222v1 Announce Type: new Abstract: Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying…
23 -
arXiv — NLP / Computation & Language research 4d ago
Stoicheia: Character-Level Masked Diffusion for Ancient Greek Textual Restoration, Parsing, and Metrical Scansion
arXiv:2608.07249v1 Announce Type: new Abstract: We introduce Stoicheia, a 405M-parameter character-level masked-diffusion encoder for Ancient Greek whose input factors into five aligned, independently maskable planes: letters, word and sentence boundaries, diacritics,…
20 -
arXiv — NLP / Computation & Language research 4d ago
Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models
arXiv:2608.07261v1 Announce Type: new Abstract: Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctly store each individual hop, it often fails to combine them. To understand the…
31 -
arXiv — NLP / Computation & Language research 4d ago
Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders
arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly…
27 -
arXiv — NLP / Computation & Language research 4d ago
Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks
arXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cross-linguistic consistency with language-specific properties. This paper…
4 -
arXiv — NLP / Computation & Language research 4d ago
Natural Language Processing Psychometrics
arXiv:2608.07316v1 Announce Type: new Abstract: Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from…
26 -
arXiv — NLP / Computation & Language research 4d ago
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
arXiv:2608.07341v1 Announce Type: new Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and…
7 -
arXiv — NLP / Computation & Language research 4d ago
Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
arXiv:2608.07353v1 Announce Type: new Abstract: Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine concept understanding. Prior work has evaluated…
18 -
arXiv — NLP / Computation & Language research 4d ago
LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering
arXiv:2608.07370v1 Announce Type: new Abstract: Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent…
15 -
arXiv — NLP / Computation & Language research 4d ago
An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis
arXiv:2608.07439v1 Announce Type: new Abstract: Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat) is one of its theoretically grounded formulations. Prior work on financial…
29 -
arXiv — NLP / Computation & Language research 4d ago
CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
arXiv:2608.07458v1 Announce Type: new Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise…
35