News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID arXiv:2609.30287v1 Announce Type: new Abstract: AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support… 27 arXiv — NLP / Computation & Language research 7h ago Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling arXiv:2609.30288v1 Announce Type: new Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or… 7 arXiv — NLP / Computation & Language research 7h ago Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents arXiv:2609.30289v1 Announce Type: new Abstract: In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations,… 25 arXiv — NLP / Computation & Language research 7h ago Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only… 28 arXiv — NLP / Computation & Language research 7h ago A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models arXiv:2609.30292v1 Announce Type: new Abstract: Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust… 28 arXiv — NLP / Computation & Language research 7h ago Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents arXiv:2609.30293v1 Announce Type: new Abstract: The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool… 20 arXiv — NLP / Computation & Language research 7h ago SlideLab: Audience-Centered Scientific Slide Generation and Evaluation arXiv:2609.30294v1 Announce Type: new Abstract: Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a… 30 arXiv — NLP / Computation & Language research 7h ago SignTrace: Describe a Sign, Find the Word arXiv:2609.30295v1 Announce Type: new Abstract: Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access… 26 arXiv — NLP / Computation & Language research 7h ago Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops arXiv:2609.30297v1 Announce Type: new Abstract: Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge… 31 arXiv — NLP / Computation & Language research 7h ago A Benchmark Framework for Screening Automation in Systematic Reviews arXiv:2609.30298v1 Announce Type: new Abstract: Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by… 20 arXiv — NLP / Computation & Language research 7h ago A Unified Account of Concepts and Chunks arXiv:2609.30414v1 Announce Type: new Abstract: Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are… 38 arXiv — NLP / Computation & Language research 7h ago All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation arXiv:2609.30416v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally… 30 arXiv — NLP / Computation & Language research 7h ago Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition arXiv:2609.30439v1 Announce Type: new Abstract: We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech… 18 arXiv — NLP / Computation & Language research 7h ago Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification arXiv:2609.30467v1 Announce Type: new Abstract: Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical… 16 arXiv — NLP / Computation & Language research 7h ago CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production arXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically… 4 arXiv — NLP / Computation & Language research 7h ago Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs arXiv:2609.30492v1 Announce Type: new Abstract: Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as… 18 arXiv — NLP / Computation & Language research 7h ago Inquesto Score: A reliability Protocol For Voice Agents arXiv:2609.30514v1 Announce Type: new Abstract: Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto… 22 arXiv — NLP / Computation & Language research 7h ago Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence arXiv:2609.30535v1 Announce Type: new Abstract: Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only… 12 arXiv — NLP / Computation & Language research 7h ago REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles arXiv:2609.30547v1 Announce Type: new Abstract: Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive… 9 arXiv — NLP / Computation & Language research 7h ago Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms arXiv:2609.30558v1 Announce Type: new Abstract: Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a… 23 arXiv — NLP / Computation & Language research 7h ago The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge arXiv:2609.30604v1 Announce Type: new Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations,… 31 arXiv — NLP / Computation & Language research 7h ago Recursive Self-Improvement via On-Policy Distillation for Reasoning arXiv:2609.30652v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the… 15 arXiv — NLP / Computation & Language research 7h ago TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding arXiv:2609.30670v1 Announce Type: new Abstract: Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses… 10 arXiv — NLP / Computation & Language research 7h ago Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4 arXiv:2609.30716v1 Announce Type: new Abstract: When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on… 33 arXiv — NLP / Computation & Language research 7h ago Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction arXiv:2609.30738v1 Announce Type: new Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention… 16 arXiv — NLP / Computation & Language research 7h ago SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages arXiv:2609.30739v1 Announce Type: new Abstract: Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing… 10 arXiv — NLP / Computation & Language research 7h ago Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance arXiv:2609.30773v1 Announce Type: new Abstract: Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech… 26 arXiv — NLP / Computation & Language research 7h ago Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment arXiv:2609.30802v1 Announce Type: new Abstract: Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during… 35 arXiv — NLP / Computation & Language research 7h ago I-Parakeet: Integer-Only Conformer ASR on Mobile NPU arXiv:2609.30846v1 Announce Type: new Abstract: In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard… 13 arXiv — NLP / Computation & Language research 7h ago Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength arXiv:2609.30849v1 Announce Type: new Abstract: Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose… 28 arXiv — NLP / Computation & Language research 7h ago Persistent Negatives for Adversarial Black-Box On-Policy Distillation arXiv:2609.30864v1 Announce Type: new Abstract: Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a… 32 arXiv — NLP / Computation & Language research 7h ago Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations arXiv:2609.30867v1 Announce Type: new Abstract: Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model… 8 arXiv — NLP / Computation & Language research 7h ago Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos arXiv:2609.30882v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim… 9 arXiv — NLP / Computation & Language research 7h ago From annotation to reasoning: Culture in language models arXiv:2609.30897v1 Announce Type: new Abstract: How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave… 30 arXiv — NLP / Computation & Language research 7h ago ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning arXiv:2609.30906v1 Announce Type: new Abstract: Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical… 12 arXiv — NLP / Computation & Language research 7h ago Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring arXiv:2609.30924v1 Announce Type: new Abstract: Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P)… 26 arXiv — NLP / Computation & Language research 7h ago Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models arXiv:2609.30935v1 Announce Type: new Abstract: Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks… 19 arXiv — NLP / Computation & Language research 7h ago FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation arXiv:2609.30968v1 Announce Type: new Abstract: Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated… 36 arXiv — NLP / Computation & Language research 7h ago Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change arXiv:2609.30974v1 Announce Type: new Abstract: Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component… 11 arXiv — NLP / Computation & Language research 7h ago THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer arXiv:2609.30984v1 Announce Type: new Abstract: Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces,… 8 arXiv — NLP / Computation & Language research 7h ago Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries arXiv:2609.30986v1 Announce Type: new Abstract: As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even… 9 arXiv — NLP / Computation & Language research 7h ago ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker arXiv:2609.31002v1 Announce Type: new Abstract: Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These… 18 arXiv — NLP / Computation & Language research 7h ago G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation arXiv:2609.31009v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer… 16 arXiv — NLP / Computation & Language research 7h ago Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference arXiv:2609.31046v1 Announce Type: new Abstract: Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is… 11 arXiv — NLP / Computation & Language research 7h ago CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings arXiv:2609.31062v1 Announce Type: new Abstract: Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological… 32 arXiv — NLP / Computation & Language research 7h ago LocUS: Head Selection and Subspace Projection for Targeted Activation Steering arXiv:2609.31122v1 Announce Type: new Abstract: Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the… 16 arXiv — NLP / Computation & Language research 7h ago Do we need to answer that question? Salience and Answerability of Potential Questions in Naturalistic Dialogue arXiv:2609.31130v1 Announce Type: new Abstract: We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we… 30 arXiv — NLP / Computation & Language research 7h ago Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting arXiv:2609.31169v1 Announce Type: new Abstract: Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the… 16 arXiv — NLP / Computation & Language research 7h ago Where a Model Sends Its Own Repeated Token arXiv:2609.31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We… 26 arXiv — NLP / Computation & Language research 7h ago RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models arXiv:2609.31245v1 Announce Type: new Abstract: Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs… 35 Page 1 of 10 · 500 articles Older →