arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 5h ago
RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation
arXiv:2608.13010v1 Announce Type: new Abstract: Retrieval-augmented generation treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims. Existing detectors depend on trusted references, specific attack artifacts, or global…
19 -
arXiv — NLP / Computation & Language research 5h ago
CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
arXiv:2608.13101v1 Announce Type: new Abstract: Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and…
36 -
arXiv — NLP / Computation & Language research 5h ago
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
arXiv:2608.13136v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas.…
32 -
arXiv — NLP / Computation & Language research 5h ago
Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering
arXiv:2608.13160v1 Announce Type: new Abstract: Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved…
15 -
arXiv — NLP / Computation & Language research 5h ago
Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory
arXiv:2608.13168v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied for emotional companionship, evaluating their behavior and capabilities in intimate relationships has become a pressing issue. However, existing assessments primarily…
24 -
arXiv — NLP / Computation & Language research 5h ago
GEM: A Generative Embedding Model Bridging Reasoning and Retrieval
arXiv:2608.13200v1 Announce Type: new Abstract: Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents,…
32 -
arXiv — NLP / Computation & Language research 5h ago
Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits
arXiv:2608.13244v1 Announce Type: new Abstract: Local chemical perception and property reasoning are both essential for understanding how molecular structure determines properties. Current LLM-based chemical reasoning methods either receive SMILES/molecular images together with…
36 -
arXiv — NLP / Computation & Language research 5h ago
Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models
arXiv:2608.13258v1 Announce Type: new Abstract: Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated,…
12 -
arXiv — NLP / Computation & Language research 5h ago
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty…
20 -
arXiv — NLP / Computation & Language research 5h ago
Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model
arXiv:2608.13277v1 Announce Type: new Abstract: We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular…
36 -
arXiv — NLP / Computation & Language research 5h ago
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form…
38 -
arXiv — NLP / Computation & Language research 5h ago
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation
arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability…
20 -
arXiv — NLP / Computation & Language research 5h ago
It's How You Ask: Gender-Associated Linguistic Bias in LLMs
arXiv:2608.13328v1 Announce Type: new Abstract: Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective…
10 -
arXiv — NLP / Computation & Language research 5h ago
RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory
arXiv:2608.13334v1 Announce Type: new Abstract: LLM-based agents increasingly rely on external memory to support long-horizon reasoning and interaction. However, the main bottleneck is not simply storing past experience, but recovering the right set of evidence when relevant…
38 -
arXiv — NLP / Computation & Language research 5h ago
CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
arXiv:2608.13387v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation…
33 -
arXiv — NLP / Computation & Language research 5h ago
Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection
arXiv:2608.13425v1 Announce Type: new Abstract: Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related…
28 -
arXiv — NLP / Computation & Language research 5h ago
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
arXiv:2608.13430v1 Announce Type: new Abstract: Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be…
5 -
arXiv — NLP / Computation & Language research 5h ago
Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity
arXiv:2608.13484v1 Announce Type: new Abstract: When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative…
25 -
arXiv — NLP / Computation & Language research 5h ago
Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
arXiv:2608.13515v1 Announce Type: new Abstract: Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task…
11 -
arXiv — NLP / Computation & Language research 5h ago
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
arXiv:2608.13517v1 Announce Type: new Abstract: Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter…
32 -
arXiv — NLP / Computation & Language research 5h ago
SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
arXiv:2608.13538v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial…
19 -
arXiv — NLP / Computation & Language research 5h ago
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
arXiv:2608.13545v1 Announce Type: new Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this…
6 -
arXiv — NLP / Computation & Language research 5h ago
When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models
arXiv:2608.12324v1 Announce Type: cross Abstract: People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about…
37 -
arXiv — NLP / Computation & Language research 5h ago
Position: Reasoning is a Learnable Rule-Based Process
arXiv:2608.12325v1 Announce Type: cross Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite…
22 -
arXiv — NLP / Computation & Language research 5h ago
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
arXiv:2608.12345v1 Announce Type: cross Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct…
9 -
arXiv — NLP / Computation & Language research 5h ago
From Observation to Intervention: Memory in Brains and Large Language Models
arXiv:2608.12377v1 Announce Type: cross Abstract: Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related information is represented, how partial cues recover broader…
19 -
arXiv — NLP / Computation & Language research 5h ago
Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
arXiv:2608.12426v1 Announce Type: cross Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled…
33 -
arXiv — NLP / Computation & Language research 5h ago
Geometric and Behavioral Stratification in Transformer Residual Streams
arXiv:2608.12447v1 Announce Type: cross Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the…
6 -
arXiv — NLP / Computation & Language research 5h ago
SoK: From Generation to Consumption of Privacy Documents in Software Systems
arXiv:2608.12511v1 Announce Type: cross Abstract: Privacy documents (e.g., privacy policies) are a central mechanism through which digital services disclose data practices and seek user consent. Over the past decades, research on privacy documents has expanded significantly,…
9 -
arXiv — NLP / Computation & Language research 5h ago
Is this Citation on Point?
arXiv:2608.12571v1 Announce Type: cross Abstract: In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting…
30 -
arXiv — NLP / Computation & Language research 5h ago
EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
arXiv:2608.12627v1 Announce Type: cross Abstract: Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are…
11 -
arXiv — NLP / Computation & Language research 5h ago
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: cross Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy…
38 -
arXiv — NLP / Computation & Language research 5h ago
Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks
arXiv:2608.12713v1 Announce Type: cross Abstract: Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining…
25 -
arXiv — NLP / Computation & Language research 5h ago
Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
arXiv:2608.12717v1 Announce Type: cross Abstract: Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the…
36 -
arXiv — NLP / Computation & Language research 5h ago
Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors
arXiv:2608.12746v1 Announce Type: cross Abstract: Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an individual object mention to what the image shows. Most…
4 -
arXiv — NLP / Computation & Language research 5h ago
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
arXiv:2608.12847v1 Announce Type: cross Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this…
8 -
arXiv — NLP / Computation & Language research 5h ago
I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization
arXiv:2608.12957v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this gap with…
33 -
arXiv — NLP / Computation & Language research 5h ago
Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks"
arXiv:2608.12974v1 Announce Type: cross Abstract: McCoy & Griffiths (2025, henceforth M&G) suggest that a Bayesian prior can be distilled into Artificial Neural Networks (ANNs) through Model-Agnostic Meta-Learning (MAML, Finn et al., 2017). They support this empirically by…
20 -
arXiv — NLP / Computation & Language research 5h ago
Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research
arXiv:2608.12984v1 Announce Type: cross Abstract: Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a…
7 -
arXiv — NLP / Computation & Language research 5h ago
Latent On-Policy Self-Distillation
arXiv:2608.13040v1 Announce Type: cross Abstract: Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to…
21 -
arXiv — NLP / Computation & Language research 5h ago
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
arXiv:2608.13057v1 Announce Type: cross Abstract: In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on…
32 -
arXiv — NLP / Computation & Language research 5h ago
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)
arXiv:2608.13063v1 Announce Type: cross Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement -…
25 -
arXiv — NLP / Computation & Language research 5h ago
TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
arXiv:2608.13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce…
30 -
arXiv — NLP / Computation & Language research 5h ago
When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1
arXiv:2608.13237v1 Announce Type: cross Abstract: Multi-round retrieval-augmented generation (RAG) must decide when to stop searching as evidence accumulates. Because the deployed policy is determined by the first STOP on each trajectory, this is a sequential selection problem…
19 -
arXiv — NLP / Computation & Language research 5h ago
Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference
arXiv:2608.13426v1 Announce Type: cross Abstract: Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free,…
25 -
arXiv — NLP / Computation & Language research 5h ago
MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
arXiv:2608.13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image…
12 -
arXiv — NLP / Computation & Language research 5h ago
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination
arXiv:2608.13476v1 Announce Type: cross Abstract: We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized…
17 -
arXiv — NLP / Computation & Language research 5h ago
Synthetic Persona Pretraining: Alignment from Token Zero
arXiv:2608.13482v1 Announce Type: cross Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only…
9 -
arXiv — NLP / Computation & Language research 5h ago
Intern-S2-Preview: Scientific Agentic Foundation Model
arXiv:2608.13505v1 Announce Type: cross Abstract: Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We…
34 -
arXiv — NLP / Computation & Language research 5h ago
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
arXiv:2608.13558v1 Announce Type: cross Abstract: Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not…
6