arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 2d ago
LLM Agents Factory: Retrieval of Domain-Specific LLM Agents
arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the…
24 -
arXiv — NLP / Computation & Language research 2d ago
Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025
arXiv:2608.09936v1 Announce Type: new Abstract: Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We examine 28,592 headlines about La France insoumise (LFI) and Rassemblement…
12 -
arXiv — NLP / Computation & Language research 2d ago
Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory
arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural…
29 -
arXiv — NLP / Computation & Language research 2d ago
The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs
arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a…
30 -
arXiv — NLP / Computation & Language research 2d ago
When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning
arXiv:2608.09942v1 Announce Type: new Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds…
21 -
arXiv — NLP / Computation & Language research 2d ago
Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling
arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or…
11 -
arXiv — NLP / Computation & Language research 2d ago
PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing
arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs,…
27 -
arXiv — NLP / Computation & Language research 2d ago
The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding
arXiv:2608.10137v1 Announce Type: new Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. However, rigid masking distorts the model's underlying probability…
16 -
arXiv — NLP / Computation & Language research 2d ago
Multimodal Item Parameter Estimation using Simulated Response Probabilitie
arXiv:2608.10154v1 Announce Type: new Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to…
19 -
arXiv — NLP / Computation & Language research 2d ago
Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems
arXiv:2608.10216v1 Announce Type: new Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question:…
32 -
arXiv — NLP / Computation & Language research 2d ago
Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So
arXiv:2608.10251v1 Announce Type: new Abstract: A transformer's answer lives on one axis: the direction its unembedding reads. Its intermediate states largely do not, and that off-axis position is usually treated as an obstacle to interpretation. We show it is functional. A…
38 -
arXiv — NLP / Computation & Language research 2d ago
TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent
arXiv:2608.10258v1 Announce Type: new Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after…
29 -
arXiv — NLP / Computation & Language research 2d ago
Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies
arXiv:2608.10273v1 Announce Type: new Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic…
4 -
arXiv — NLP / Computation & Language research 2d ago
Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a…
32 -
arXiv — NLP / Computation & Language research 2d ago
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic…
36 -
arXiv — NLP / Computation & Language research 2d ago
Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility
arXiv:2608.10315v1 Announce Type: new Abstract: Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an…
19 -
arXiv — NLP / Computation & Language research 2d ago
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise…
14 -
arXiv — NLP / Computation & Language research 2d ago
How Robust Are LLMs to Vietnamese Dialects?
arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect…
22 -
arXiv — NLP / Computation & Language research 2d ago
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
arXiv:2608.10444v1 Announce Type: new Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively…
21 -
arXiv — NLP / Computation & Language research 2d ago
MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection
arXiv:2608.10459v1 Announce Type: new Abstract: As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models.…
25 -
arXiv — NLP / Computation & Language research 2d ago
Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection
arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) therefore aims to determine whether a given text is a…
28 -
arXiv — NLP / Computation & Language research 2d ago
Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases
arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks.…
21 -
arXiv — NLP / Computation & Language research 2d ago
ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS
arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct…
5 -
arXiv — NLP / Computation & Language research 2d ago
Simplex Relaxation for Discrete Diffusion
arXiv:2608.10615v1 Announce Type: new Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem. We study uniform discrete diffusion and ask…
31 -
arXiv — NLP / Computation & Language research 2d ago
Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue
arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently multi-turn and path-dependent: users disclose concerns gradually, emotions evolve…
27 -
arXiv — NLP / Computation & Language research 2d ago
Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text
arXiv:2608.10627v1 Announce Type: new Abstract: Decompose-then-verify pipelines, including FActScore-style fact-checkers and long-form factuality evaluators, first split a passage into atomic claims before checking each one. Decomposition itself is treated as a neutral…
38 -
arXiv — NLP / Computation & Language research 2d ago
Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR
arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first…
28 -
arXiv — NLP / Computation & Language research 2d ago
Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics
arXiv:2608.10678v1 Announce Type: new Abstract: Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too…
11 -
arXiv — NLP / Computation & Language research 2d ago
Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus
arXiv:2608.10688v1 Announce Type: new Abstract: Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, existing keyphrase extraction (KPE) studies mainly focus on improving textual…
34 -
arXiv — NLP / Computation & Language research 2d ago
Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?
arXiv:2608.10690v1 Announce Type: new Abstract: Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific token groups from released tokenizer vocabularies;…
27 -
arXiv — NLP / Computation & Language research 2d ago
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
arXiv:2608.10692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the…
21 -
arXiv — NLP / Computation & Language research 2d ago
EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection
arXiv:2608.10698v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios involving human-written text (HWT), LLM-generated text (LGT), and…
5 -
arXiv — NLP / Computation & Language research 2d ago
Most biomedical publications show signs of LLM-assisted writing
arXiv:2608.10715v1 Announce Type: new Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about…
25 -
arXiv — NLP / Computation & Language research 2d ago
Mitigating Context Interference for Reliable and Efficient Search Agents
arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the contexts of multi-turn search agents are lengthy and…
38 -
arXiv — NLP / Computation & Language research 2d ago
Assessing Reliability of BERT-Based Models on Question Answering Tasks
arXiv:2608.10806v1 Announce Type: new Abstract: Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical applications. Recent advancements in natural language…
9 -
arXiv — NLP / Computation & Language research 2d ago
Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
arXiv:2608.10810v1 Announce Type: new Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks…
5 -
arXiv — NLP / Computation & Language research 2d ago
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
arXiv:2608.10812v1 Announce Type: new Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a…
30 -
arXiv — NLP / Computation & Language research 2d ago
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
arXiv:2608.10875v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs…
20 -
arXiv — NLP / Computation & Language research 2d ago
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction
arXiv:2608.10878v1 Announce Type: new Abstract: Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular…
27 -
arXiv — NLP / Computation & Language research 2d ago
Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift
arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a $\beta$-fraction of shifted target traffic with at most an $\alpha$-fraction of answers wrong. Under…
4 -
arXiv — NLP / Computation & Language research 2d ago
FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation
arXiv:2608.10916v1 Announce Type: new Abstract: Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these systems. Existing approaches require expensive…
37 -
arXiv — NLP / Computation & Language research 2d ago
A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models
arXiv:2608.10939v1 Announce Type: new Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and…
23 -
arXiv — NLP / Computation & Language research 2d ago
REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs
arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a budget of at most 32B parameters and no model fine-tuning. Our system combines…
14 -
arXiv — NLP / Computation & Language research 2d ago
ReLTEx: Reliable LLM-based Taxonomy Expansion
arXiv:2608.10970v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment. However, directly relying on…
8 -
arXiv — NLP / Computation & Language research 2d ago
MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales
arXiv:2608.10974v1 Announce Type: new Abstract: Scientific papers contain fine-grained records of problem solving: authors mention technical obstacles and methods that were used to address them, often along with reasoning on why those methods were chosen. We introduce MUSE…
32 -
arXiv — NLP / Computation & Language research 2d ago
What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model
arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring…
5 -
arXiv — NLP / Computation & Language research 2d ago
ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering
arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical questions lack comparably cheap outcome verifiers:…
26 -
arXiv — NLP / Computation & Language research 2d ago
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
arXiv:2608.11002v1 Announce Type: new Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects…
37 -
arXiv — NLP / Computation & Language research 2d ago
Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance
arXiv:2608.11008v1 Announce Type: new Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild,…
25 -
arXiv — NLP / Computation & Language research 2d ago
Data Attribution of Emergent Misalignment with Persona Features
arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions…
29