arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 7h ago
A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
arXiv:2609.30287v1 Announce Type: new Abstract: AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support…
27 -
arXiv — NLP / Computation & Language research 7h ago
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
arXiv:2609.30288v1 Announce Type: new Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or…
7 -
arXiv — NLP / Computation & Language research 7h ago
Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
arXiv:2609.30289v1 Announce Type: new Abstract: In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations,…
25 -
arXiv — NLP / Computation & Language research 7h ago
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only…
28 -
arXiv — NLP / Computation & Language research 7h ago
A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models
arXiv:2609.30292v1 Announce Type: new Abstract: Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust…
28 -
arXiv — NLP / Computation & Language research 7h ago
Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
arXiv:2609.30293v1 Announce Type: new Abstract: The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool…
20 -
arXiv — NLP / Computation & Language research 7h ago
SlideLab: Audience-Centered Scientific Slide Generation and Evaluation
arXiv:2609.30294v1 Announce Type: new Abstract: Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a…
30 -
arXiv — NLP / Computation & Language research 7h ago
SignTrace: Describe a Sign, Find the Word
arXiv:2609.30295v1 Announce Type: new Abstract: Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access…
26 -
arXiv — NLP / Computation & Language research 7h ago
Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
arXiv:2609.30297v1 Announce Type: new Abstract: Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge…
31 -
arXiv — NLP / Computation & Language research 7h ago
A Benchmark Framework for Screening Automation in Systematic Reviews
arXiv:2609.30298v1 Announce Type: new Abstract: Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by…
20 -
arXiv — NLP / Computation & Language research 7h ago
A Unified Account of Concepts and Chunks
arXiv:2609.30414v1 Announce Type: new Abstract: Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are…
38 -
arXiv — NLP / Computation & Language research 7h ago
All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation
arXiv:2609.30416v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally…
30 -
arXiv — NLP / Computation & Language research 7h ago
Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
arXiv:2609.30439v1 Announce Type: new Abstract: We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech…
18 -
arXiv — NLP / Computation & Language research 7h ago
Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
arXiv:2609.30467v1 Announce Type: new Abstract: Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical…
16 -
arXiv — NLP / Computation & Language research 7h ago
CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
arXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically…
4 -
arXiv — NLP / Computation & Language research 7h ago
Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs
arXiv:2609.30492v1 Announce Type: new Abstract: Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as…
18 -
arXiv — NLP / Computation & Language research 7h ago
Inquesto Score: A reliability Protocol For Voice Agents
arXiv:2609.30514v1 Announce Type: new Abstract: Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto…
22 -
arXiv — NLP / Computation & Language research 7h ago
Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
arXiv:2609.30535v1 Announce Type: new Abstract: Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only…
12 -
arXiv — NLP / Computation & Language research 7h ago
REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles
arXiv:2609.30547v1 Announce Type: new Abstract: Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive…
9 -
arXiv — NLP / Computation & Language research 7h ago
Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
arXiv:2609.30558v1 Announce Type: new Abstract: Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a…
23 -
arXiv — NLP / Computation & Language research 7h ago
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
arXiv:2609.30604v1 Announce Type: new Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations,…
31 -
arXiv — NLP / Computation & Language research 7h ago
Recursive Self-Improvement via On-Policy Distillation for Reasoning
arXiv:2609.30652v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the…
15 -
arXiv — NLP / Computation & Language research 7h ago
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
arXiv:2609.30670v1 Announce Type: new Abstract: Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses…
10 -
arXiv — NLP / Computation & Language research 7h ago
Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4
arXiv:2609.30716v1 Announce Type: new Abstract: When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on…
33 -
arXiv — NLP / Computation & Language research 7h ago
Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction
arXiv:2609.30738v1 Announce Type: new Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention…
16 -
arXiv — NLP / Computation & Language research 7h ago
SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages
arXiv:2609.30739v1 Announce Type: new Abstract: Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing…
10 -
arXiv — NLP / Computation & Language research 7h ago
Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
arXiv:2609.30773v1 Announce Type: new Abstract: Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech…
26 -
arXiv — NLP / Computation & Language research 7h ago
Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment
arXiv:2609.30802v1 Announce Type: new Abstract: Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during…
35 -
arXiv — NLP / Computation & Language research 7h ago
I-Parakeet: Integer-Only Conformer ASR on Mobile NPU
arXiv:2609.30846v1 Announce Type: new Abstract: In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard…
13 -
arXiv — NLP / Computation & Language research 7h ago
Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
arXiv:2609.30849v1 Announce Type: new Abstract: Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose…
28 -
arXiv — NLP / Computation & Language research 7h ago
Persistent Negatives for Adversarial Black-Box On-Policy Distillation
arXiv:2609.30864v1 Announce Type: new Abstract: Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a…
32 -
arXiv — NLP / Computation & Language research 7h ago
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
arXiv:2609.30867v1 Announce Type: new Abstract: Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model…
8 -
arXiv — NLP / Computation & Language research 7h ago
Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos
arXiv:2609.30882v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim…
9 -
arXiv — NLP / Computation & Language research 7h ago
From annotation to reasoning: Culture in language models
arXiv:2609.30897v1 Announce Type: new Abstract: How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave…
30 -
arXiv — NLP / Computation & Language research 7h ago
ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
arXiv:2609.30906v1 Announce Type: new Abstract: Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical…
12 -
arXiv — NLP / Computation & Language research 7h ago
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
arXiv:2609.30924v1 Announce Type: new Abstract: Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P)…
26 -
arXiv — NLP / Computation & Language research 7h ago
Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models
arXiv:2609.30935v1 Announce Type: new Abstract: Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks…
19 -
arXiv — NLP / Computation & Language research 7h ago
FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation
arXiv:2609.30968v1 Announce Type: new Abstract: Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated…
36 -
arXiv — NLP / Computation & Language research 7h ago
Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change
arXiv:2609.30974v1 Announce Type: new Abstract: Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component…
11 -
arXiv — NLP / Computation & Language research 7h ago
THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer
arXiv:2609.30984v1 Announce Type: new Abstract: Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces,…
8 -
arXiv — NLP / Computation & Language research 7h ago
Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
arXiv:2609.30986v1 Announce Type: new Abstract: As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even…
9 -
arXiv — NLP / Computation & Language research 7h ago
ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
arXiv:2609.31002v1 Announce Type: new Abstract: Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These…
18 -
arXiv — NLP / Computation & Language research 7h ago
G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
arXiv:2609.31009v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer…
16 -
arXiv — NLP / Computation & Language research 7h ago
Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference
arXiv:2609.31046v1 Announce Type: new Abstract: Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is…
11 -
arXiv — NLP / Computation & Language research 7h ago
CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
arXiv:2609.31062v1 Announce Type: new Abstract: Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological…
32 -
arXiv — NLP / Computation & Language research 7h ago
LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
arXiv:2609.31122v1 Announce Type: new Abstract: Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the…
16 -
arXiv — NLP / Computation & Language research 7h ago
Do we need to answer that question? Salience and Answerability of Potential Questions in Naturalistic Dialogue
arXiv:2609.31130v1 Announce Type: new Abstract: We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we…
30 -
arXiv — NLP / Computation & Language research 7h ago
Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting
arXiv:2609.31169v1 Announce Type: new Abstract: Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the…
16 -
arXiv — NLP / Computation & Language research 7h ago
Where a Model Sends Its Own Repeated Token
arXiv:2609.31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We…
26 -
arXiv — NLP / Computation & Language research 7h ago
RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models
arXiv:2609.31245v1 Announce Type: new Abstract: Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs…
35