arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 3d ago
Unified Hallucination Fuzzing for Multimodal Large Language Models
arXiv:2608.07525v1 Announce Type: new Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from…
16 -
arXiv — NLP / Computation & Language research 3d ago
DocAtlas: Long-Document Understanding as Mutable-State Interaction
arXiv:2608.07527v1 Announce Type: new Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually select evidence from a static index before generation,…
29 -
arXiv — NLP / Computation & Language research 3d ago
WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management
arXiv:2608.07529v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than…
6 -
arXiv — NLP / Computation & Language research 3d ago
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
arXiv:2608.07531v1 Announce Type: new Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from…
19 -
arXiv — NLP / Computation & Language research 3d ago
Scaling Inherently Interpretable Language Models
arXiv:2608.07594v1 Announce Type: new Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this…
27 -
arXiv — NLP / Computation & Language research 3d ago
Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation
arXiv:2608.07629v1 Announce Type: new Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. When fine-tuning these models for an unseen language,…
18 -
arXiv — NLP / Computation & Language research 3d ago
SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators
arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly…
36 -
arXiv — NLP / Computation & Language research 3d ago
Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages
arXiv:2608.07727v1 Announce Type: new Abstract: Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so it's not clear how much per-language ability these models actually keep. I have…
36 -
arXiv — NLP / Computation & Language research 3d ago
The No-Meaning Falsity: The Structural Impossibility of the Arbitrary Sign in Classical Arabic
arXiv:2608.07737v1 Announce Type: new Abstract: This paper investigates whether the postmodern claim of unrestricted semantic indeterminacy, and its foundational Saussurean axiom of the arbitrary sign, are compatible with the structural architecture of Classical Arabic. We…
29 -
arXiv — NLP / Computation & Language research 3d ago
Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation
arXiv:2608.07763v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which…
19 -
arXiv — NLP / Computation & Language research 3d ago
On the use of foundation models in cognitive science
arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive…
34 -
arXiv — NLP / Computation & Language research 3d ago
"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders
arXiv:2608.07852v1 Announce Type: new Abstract: How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker representations using a dataset of user-expressed…
7 -
arXiv — NLP / Computation & Language research 3d ago
SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.…
21 -
arXiv — NLP / Computation & Language research 3d ago
Detection of Self-Introductions in Legislative Testimony
arXiv:2608.07891v1 Announce Type: new Abstract: Self-introductions are common in legislative committee testimonies. Successfully detecting them and extracting the speaker's name is enormously helpful in the task of speaker identification in the context of government meetings. In…
4 -
arXiv — NLP / Computation & Language research 3d ago
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
arXiv:2608.07968v1 Announce Type: new Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency…
33 -
arXiv — NLP / Computation & Language research 3d ago
Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States
arXiv:2608.08024v1 Announce Type: new Abstract: Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Probes (PEP), a white-box method for answer-level hallucination detection from the hidden…
33 -
arXiv — NLP / Computation & Language research 3d ago
APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain
arXiv:2608.08059v1 Announce Type: new Abstract: Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents. Despite substantial work…
25 -
arXiv — NLP / Computation & Language research 3d ago
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
arXiv:2608.08067v1 Announce Type: new Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the…
14 -
arXiv — NLP / Computation & Language research 3d ago
Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models
arXiv:2608.08082v1 Announce Type: new Abstract: Classifier-free guidance (CFG) is usually kept on throughout masked diffusion language model decoding, although its benefit varies across prompts and over time. We study when CFG is actually needed by comparing, from any partial…
4 -
arXiv — NLP / Computation & Language research 3d ago
Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models
arXiv:2608.08086v1 Announce Type: new Abstract: Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability distinguishes them from irreversible autoregressive generation, but makes…
37 -
arXiv — NLP / Computation & Language research 3d ago
Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs
arXiv:2608.08090v1 Announce Type: new Abstract: Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clearer understanding of the contribution of translated multilingual supervision. We…
17 -
arXiv — NLP / Computation & Language research 3d ago
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
arXiv:2608.08107v1 Announce Type: new Abstract: Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this phenomenon from the perspective…
8 -
arXiv — NLP / Computation & Language research 3d ago
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
arXiv:2608.08160v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of…
14 -
arXiv — NLP / Computation & Language research 3d ago
STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs
arXiv:2608.08164v1 Announce Type: new Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller…
29 -
arXiv — NLP / Computation & Language research 3d ago
Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit Thinking mode from direct answer generation (NoThinking mode) remain…
29 -
arXiv — NLP / Computation & Language research 3d ago
A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization
arXiv:2608.08180v1 Announce Type: new Abstract: Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities and events. Such relation-level hallucinations undermine the reliability of…
28 -
arXiv — NLP / Computation & Language research 3d ago
Focus particles and scalar inferences across humans and language models
arXiv:2608.08227v1 Announce Type: new Abstract: Focus particles such as "even" and "only" are central to formal semantic theories that posit structured representations over sets of alternatives. "Even" highlights unexpected or extreme alternatives, while "only" enforces…
8 -
arXiv — NLP / Computation & Language research 3d ago
AraSSM: A bidirectional state-space encoder for Arabic masked language modeling
arXiv:2608.08256v1 Announce Type: new Abstract: Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but their self-attention mechanism scales quadratically with sequence length,…
17 -
arXiv — NLP / Computation & Language research 3d ago
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
arXiv:2608.08283v1 Announce Type: new Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic…
30 -
arXiv — NLP / Computation & Language research 3d ago
Safety Cost of Steering Vectors Is Separable and Reducible
arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechanisms and increase compliance with harmful requests,…
34 -
arXiv — NLP / Computation & Language research 3d ago
Hidden Language Consistency Phenomena in Reasoning LLMs
arXiv:2608.08447v1 Announce Type: new Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingual…
18 -
arXiv — NLP / Computation & Language research 3d ago
Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization
arXiv:2608.08451v1 Announce Type: new Abstract: Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity…
22 -
arXiv — NLP / Computation & Language research 3d ago
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
arXiv:2608.08459v1 Announce Type: new Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare, education, transportation, and enterprise…
22 -
arXiv — NLP / Computation & Language research 3d ago
VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use
arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to a 1.04B Spanish/LATAM security decoder via an MLP. To our knowledge, it is the…
9 -
arXiv — NLP / Computation & Language research 3d ago
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
arXiv:2608.08510v1 Announce Type: new Abstract: Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This "cocktail party" scenario still…
33 -
arXiv — NLP / Computation & Language research 3d ago
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories
arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for…
17 -
arXiv — NLP / Computation & Language research 3d ago
Mitigating Gender Bias in English to Romanian Machine Translation
arXiv:2608.08606v1 Announce Type: new Abstract: Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a gendered target language such as Romanian. This bias results in translations…
14 -
arXiv — NLP / Computation & Language research 3d ago
North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings
arXiv:2608.08607v1 Announce Type: new Abstract: Mental health disorders are a leading cause of disability worldwide, yet Natural Language Processing (NLP) research for mental healthcare has remained concentrated in high-income, English-language settings. North Africa, and…
7 -
arXiv — NLP / Computation & Language research 3d ago
Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach
arXiv:2608.08636v1 Announce Type: new Abstract: Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts. Recently, large language models (LLMs) have demonstrated the capacity to achieve competitive…
19 -
arXiv — NLP / Computation & Language research 3d ago
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism
arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution cannot be explained by a chronological list of model releases alone. This…
33 -
arXiv — NLP / Computation & Language research 3d ago
LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization
arXiv:2608.08721v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing…
25 -
arXiv — NLP / Computation & Language research 3d ago
Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs
arXiv:2608.08744v1 Announce Type: new Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds the one-time cost of fine-tuning. Yet most efficiency interventions target either…
8 -
arXiv — NLP / Computation & Language research 3d ago
Multilingual Emotion Neurons in Large Audio-Language Models
arXiv:2608.08772v1 Announce Type: new Abstract: Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion…
23 -
arXiv — NLP / Computation & Language research 3d ago
OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically…
18 -
arXiv — NLP / Computation & Language research 3d ago
Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models
arXiv:2608.08791v1 Announce Type: new Abstract: Diffusion language models use broad context to create text, suggesting they might handle input noise better than standard models. Testing reveals this is only partially true. Internally, diffusion models detect text errors highly…
38 -
arXiv — NLP / Computation & Language research 3d ago
Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents
arXiv:2608.08793v1 Announce Type: new Abstract: Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly represented by session-, model-, or tool-centric traces: a Skill can be discovered…
24 -
arXiv — NLP / Computation & Language research 3d ago
Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages
arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed to reach a certain performance. We validate this…
18 -
arXiv — NLP / Computation & Language research 3d ago
IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical Requirements
arXiv:2608.08801v1 Announce Type: new Abstract: Translating technical requirements across languages can introduce semantic drift, altering numerical constraints, polarities, modalities, or other specification-critical meaning. IDRAAK is presented as an interpretable framework…
11 -
arXiv — NLP / Computation & Language research 3d ago
Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers
arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the right trade-off changes with the workload. In the context of information retrieval…
12 -
arXiv — NLP / Computation & Language research 3d ago
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
arXiv:2608.08829v1 Announce Type: new Abstract: Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task. We argue that the best layers are an…
32