arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 3d ago
Explicit Boundary Markers for Subword Vocabularies
arXiv:2608.08847v1 Announce Type: new Abstract: Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have separate embeddings in models, so occurrences of one word are divided across rows…
37 -
arXiv — NLP / Computation & Language research 3d ago
Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State
arXiv:2608.08868v1 Announce Type: new Abstract: Many modern AI systems analyze conversational traces to infer aspects of human interaction and state, implicitly assuming that such information is recoverable from conversation. We study observability: whether a target is…
28 -
arXiv — NLP / Computation & Language research 3d ago
Position Bias in Ordinal Classification: A Systematic Evaluation
arXiv:2608.08869v1 Announce Type: new Abstract: Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from…
36 -
arXiv — NLP / Computation & Language research 3d ago
Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving
arXiv:2608.08910v1 Announce Type: new Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses the decomposition into a single uniform nine-level quantizer, a known…
29 -
arXiv — NLP / Computation & Language research 3d ago
Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue
arXiv:2608.08915v1 Announce Type: new Abstract: Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much…
35 -
arXiv — NLP / Computation & Language research 3d ago
Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access
arXiv:2608.08942v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide comparable assistance regardless of a user's literacy, communication style, or…
22 -
arXiv — NLP / Computation & Language research 3d ago
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
arXiv:2608.08975v1 Announce Type: new Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how…
24 -
arXiv — NLP / Computation & Language research 3d ago
How Far Do Foundation Models Transfer to Infant Signals? A Cross-Dataset Transfer Audit with a Unified Need Ontology
arXiv:2608.08989v1 Announce Type: new Abstract: Public infant cry corpora are small, label-incompatible, and almost always evaluated one corpus at a time. We ask what this practice hides and what fixes it. Across four cry corpora screened by a multi-level leakage audit…
15 -
arXiv — NLP / Computation & Language research 3d ago
ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making
arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prioritize limited clinical resources. At presentation, however, information may be…
5 -
arXiv — NLP / Computation & Language research 3d ago
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
arXiv:2608.09043v1 Announce Type: new Abstract: Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We formalize this setting as streaming dialogue summarization, where a system must…
37 -
arXiv — NLP / Computation & Language research 3d ago
Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents
arXiv:2608.09044v1 Announce Type: new Abstract: Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectories or abstract shared knowledge from related…
27 -
arXiv — NLP / Computation & Language research 3d ago
Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production
arXiv:2608.09045v1 Announce Type: new Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR),…
4 -
arXiv — NLP / Computation & Language research 3d ago
Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities
arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is…
35 -
arXiv — NLP / Computation & Language research 3d ago
Security and Privacy Taxonomy Generation from Mobile App Reviews
arXiv:2608.09049v1 Announce Type: new Abstract: Mobile app reviews are a rich, continuously renewing source of how users experience privacy and security, yet existing taxonomies of these concerns are hand-crafted and cannot keep pace with the evolving nature of the data.…
28 -
arXiv — NLP / Computation & Language research 3d ago
When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information
arXiv:2608.09080v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for…
4 -
arXiv — NLP / Computation & Language research 3d ago
The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora
arXiv:2608.09093v1 Announce Type: new Abstract: How a document's arrangement is written down, its notation, is a training variable that no dataset card records. The field has established that text-extraction choices change model behaviour, and has never once measured the…
35 -
arXiv — NLP / Computation & Language research 3d ago
Evo-Bench: Can Language Models Improve Agent Harness?
arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously…
27 -
arXiv — NLP / Computation & Language research 3d ago
LexKairos: Benchmarking Legal Temporal Capabilities in LLMs
arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the validity of statutes, the progression of legal cases, and the…
37 -
arXiv — NLP / Computation & Language research 3d ago
Subjective Multi-Bias Detection with Large Language Models
arXiv:2608.09126v1 Announce Type: new Abstract: In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identification of subjective bias, a type of bias that introduces improper attitudes or…
20 -
arXiv — NLP / Computation & Language research 3d ago
Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
arXiv:2608.09128v1 Announce Type: new Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social…
29 -
arXiv — NLP / Computation & Language research 3d ago
An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer
arXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many…
21 -
arXiv — NLP / Computation & Language research 3d ago
UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following
arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a reference document (i.e., back-translation) is a widely used method to…
30 -
arXiv — NLP / Computation & Language research 3d ago
Failure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation System
arXiv:2608.09187v1 Announce Type: new Abstract: A long-form translation request can succeed at the API layer and still produce an unusable result. The output may be empty, truncated, filtered, dominated by source or prompt material, or interrupted after producing text worth…
32 -
arXiv — NLP / Computation & Language research 3d ago
EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models
arXiv:2608.09189v1 Announce Type: new Abstract: Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a…
7 -
arXiv — NLP / Computation & Language research 3d ago
UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
arXiv:2608.09209v1 Announce Type: new Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on…
33 -
arXiv — NLP / Computation & Language research 3d ago
Reading Cognition as Decisions Unfold in Words: A Factorized Inverse Decision Model
arXiv:2608.09222v1 Announce Type: new Abstract: Inverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action trajectories. In verbalized cognitive tasks, task execution also produces response…
10 -
arXiv — NLP / Computation & Language research 3d ago
Verifiably grounded machine interpretation of lunar geology
arXiv:2608.09276v1 Announce Type: new Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an automated "machine intelligence geologist" by embedding this distinct…
33 -
arXiv — NLP / Computation & Language research 3d ago
Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025
arXiv:2608.09280v1 Announce Type: new Abstract: Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Checklists to…
27 -
arXiv — NLP / Computation & Language research 3d ago
Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing
arXiv:2608.09289v1 Announce Type: new Abstract: Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipeline that…
18 -
arXiv — NLP / Computation & Language research 3d ago
Universal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic Languages
arXiv:2608.09356v1 Announce Type: new Abstract: Closely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare two approaches to script unification: the general-purpose uroman romanizer and…
11 -
arXiv — NLP / Computation & Language research 3d ago
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
arXiv:2608.09393v1 Announce Type: new Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the…
17 -
arXiv — NLP / Computation & Language research 3d ago
Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation
arXiv:2608.09420v1 Announce Type: new Abstract: User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple…
12 -
arXiv — NLP / Computation & Language research 3d ago
Reducing Pretraining-Generation Mismatch in Diffusion Language Models
arXiv:2608.09424v1 Announce Type: new Abstract: Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left context. Diffusion language models offer parallel denoising, but native dLLM…
27 -
arXiv — NLP / Computation & Language research 3d ago
ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models
arXiv:2608.09432v1 Announce Type: new Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by…
11 -
arXiv — NLP / Computation & Language research 3d ago
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
arXiv:2608.09507v1 Announce Type: new Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full…
29 -
arXiv — NLP / Computation & Language research 3d ago
Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts
arXiv:2608.09510v1 Announce Type: new Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring…
38 -
arXiv — NLP / Computation & Language research 3d ago
TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
arXiv:2608.09538v1 Announce Type: new Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top…
16 -
arXiv — NLP / Computation & Language research 3d ago
Mawqif-v2: An Arabic Benchmark Dataset for Cross-Target Stance Detection
arXiv:2608.09539v1 Announce Type: new Abstract: Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This paper presents the Mawqif-v2 Extension, consisting of 996 manually annotated…
9 -
arXiv — NLP / Computation & Language research 3d ago
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
arXiv:2608.09548v1 Announce Type: new Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed…
32 -
arXiv — NLP / Computation & Language research 3d ago
Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models
arXiv:2608.09551v1 Announce Type: new Abstract: In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks directly exploit…
26 -
arXiv — NLP / Computation & Language research 3d ago
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization
arXiv:2608.09568v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing equally to the preference signal. However, the contribution of individual…
30 -
arXiv — NLP / Computation & Language research 3d ago
MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL
arXiv:2608.09588v1 Announce Type: new Abstract: Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, heterogeneous database collection. We therefore study schema linking in a…
31 -
arXiv — NLP / Computation & Language research 3d ago
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
arXiv:2608.09624v1 Announce Type: new Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the…
6 -
arXiv — NLP / Computation & Language research 3d ago
How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans
arXiv:2608.09717v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction…
16 -
arXiv — NLP / Computation & Language research 3d ago
REFRAMED: Towards Realistic Audio Description Generation for Movies
arXiv:2608.09765v1 Announce Type: new Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into…
35 -
arXiv — NLP / Computation & Language research 3d ago
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
arXiv:2608.09766v1 Announce Type: new Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale…
30 -
arXiv — NLP / Computation & Language research 3d ago
Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification
arXiv:2608.09767v1 Announce Type: new Abstract: Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based…
35 -
arXiv — NLP / Computation & Language research 3d ago
PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial…
8 -
arXiv — NLP / Computation & Language research 3d ago
KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs
arXiv:2608.09779v1 Announce Type: new Abstract: Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to…
9 -
arXiv — NLP / Computation & Language research 3d ago
Comparing British and American Audio Description of Movies
arXiv:2608.09792v1 Announce Type: new Abstract: Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than most…
24