News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding arXiv:2609.31255v1 Announce Type: new Abstract: General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is… 35 arXiv — NLP / Computation & Language research 7h ago MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries arXiv:2609.31261v1 Announce Type: new Abstract: The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We… 13 arXiv — NLP / Computation & Language research 7h ago Identifying Scientists on X arXiv:2609.31264v1 Announce Type: new Abstract: With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an… 32 arXiv — NLP / Computation & Language research 7h ago Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers arXiv:2609.31342v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document… 36 arXiv — NLP / Computation & Language research 7h ago Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding arXiv:2609.31382v1 Announce Type: new Abstract: Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content.… 26 arXiv — NLP / Computation & Language research 7h ago Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers arXiv:2609.31403v1 Announce Type: new Abstract: Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the… 24 arXiv — NLP / Computation & Language research 7h ago ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs arXiv:2609.31448v1 Announce Type: new Abstract: Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from… 29 arXiv — NLP / Computation & Language research 7h ago Evaluating Cultural Awareness of LLMs for Haitian Creole arXiv:2609.31506v1 Announce Type: new Abstract: Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented… 22 arXiv — NLP / Computation & Language research 7h ago Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge arXiv:2609.31511v1 Announce Type: new Abstract: We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a… 9 arXiv — NLP / Computation & Language research 7h ago MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos arXiv:2609.31553v1 Announce Type: new Abstract: Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of… 24 arXiv — NLP / Computation & Language research 7h ago Strategically Diverse Sampling for Self-Training arXiv:2609.31571v1 Announce Type: new Abstract: Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically… 19 arXiv — NLP / Computation & Language research 7h ago PALM: Point-in-Time Adaptation for Financial Language Models arXiv:2609.30316v1 Announce Type: cross Abstract: Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time… 30 arXiv — NLP / Computation & Language research 7h ago When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess arXiv:2609.30328v1 Announce Type: cross Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for.… 11 arXiv — NLP / Computation & Language research 7h ago What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study arXiv:2609.30402v1 Announce Type: cross Abstract: Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices… 4 arXiv — NLP / Computation & Language research 7h ago RAZOR: Pruning Replaceable Experts in LLMs arXiv:2609.30465v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output… 22 arXiv — NLP / Computation & Language research 7h ago Asymmetric Classifier-Free Guidance for Target-Speaker ASR arXiv:2609.30476v1 Announce Type: cross Abstract: Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech… 7 arXiv — NLP / Computation & Language research 7h ago AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth arXiv:2609.30483v1 Announce Type: cross Abstract: Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against… 4 arXiv — NLP / Computation & Language research 7h ago Don't CLAP: Are Music-Text Models Bag-of-Words? arXiv:2609.30540v1 Announce Type: cross Abstract: Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective… 7 arXiv — NLP / Computation & Language research 7h ago Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content arXiv:2609.30563v1 Announce Type: cross Abstract: Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on… 21 arXiv — NLP / Computation & Language research 7h ago Epstein Files Engine: Agentic Search for Investigative Journalism arXiv:2609.30611v1 Announce Type: cross Abstract: On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times… 24 arXiv — NLP / Computation & Language research 7h ago Prompt Injection Detection for Email Agents Through Attack Chain Modeling arXiv:2609.30657v1 Announce Type: cross Abstract: Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection… 6 arXiv — NLP / Computation & Language research 7h ago LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information arXiv:2609.30706v1 Announce Type: cross Abstract: "System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a… 30 arXiv — NLP / Computation & Language research 7h ago Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models arXiv:2609.30784v1 Announce Type: cross Abstract: This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes… 23 arXiv — NLP / Computation & Language research 7h ago Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness arXiv:2609.30820v1 Announce Type: cross Abstract: Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails… 9 arXiv — NLP / Computation & Language research 7h ago Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces arXiv:2609.30914v1 Announce Type: cross Abstract: Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and… 25 arXiv — NLP / Computation & Language research 7h ago Does Uniform Discrete Diffusion Need Time? arXiv:2609.30977v1 Announce Type: cross Abstract: Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends… 36 arXiv — NLP / Computation & Language research 7h ago Same Text, Different Numbers: The Divergence of LLM-Based Measures arXiv:2609.31013v1 Announce Type: cross Abstract: Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen… 21 arXiv — NLP / Computation & Language research 7h ago KuaFu: Compressing Long User Behavior into Understanding at Billion Scale arXiv:2609.31045v1 Announce Type: cross Abstract: Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant… 37 arXiv — NLP / Computation & Language research 7h ago JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models arXiv:2609.31142v1 Announce Type: cross Abstract: Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a… 33 arXiv — NLP / Computation & Language research 7h ago Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring arXiv:2609.31293v1 Announce Type: cross Abstract: Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or… 29 arXiv — NLP / Computation & Language research 7h ago The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models arXiv:2609.31341v1 Announce Type: cross Abstract: Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed)… 11 arXiv — NLP / Computation & Language research 7h ago Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models arXiv:2609.31397v1 Announce Type: cross Abstract: Automated and highly usable Quality-of-Service (QoS) enforcement requires translating high-level service intents into deployable traffic-management policies. Although intent-based networking (IBN) has simplified policy… 35 arXiv — NLP / Computation & Language research 7h ago Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis arXiv:2609.31422v1 Announce Type: cross Abstract: Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with… 24 arXiv — NLP / Computation & Language research 7h ago PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents arXiv:2609.31468v1 Announce Type: cross Abstract: LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean… 19 arXiv — NLP / Computation & Language research 7h ago Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing arXiv:2609.31513v1 Announce Type: cross Abstract: We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures… 16 arXiv — NLP / Computation & Language research 7h ago Two Conformal Constructions for Adaptive Within-Document AI-Text Screening arXiv:2609.31547v1 Announce Type: cross Abstract: We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection… 17 arXiv — NLP / Computation & Language research 7h ago Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer arXiv:2609.31587v1 Announce Type: cross Abstract: We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether… 14 arXiv — NLP / Computation & Language research 7h ago User Model Extraction via Belief Self-Distillation arXiv:2609.31603v1 Announce Type: cross Abstract: Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a… 9 arXiv — NLP / Computation & Language research 7h ago Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency arXiv:2609.31619v1 Announce Type: cross Abstract: Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly… 19 arXiv — NLP / Computation & Language research 7h ago NaijaNLP: A Survey of Nigerian Low-Resource Languages arXiv:2502.19784v3 Announce Type: replace Abstract: With over 500 languages in Nigeria, three languages - Hausa, Yor\`ub\'a and Igbo spoken by more than 175 million people, account for about 65% of the languages. However, these languages are classed as low-resource due to… 10 arXiv — NLP / Computation & Language research 7h ago MedHal: a Synthetic Dataset for Medical Hallucination Detection arXiv:2504.08596v3 Announce Type: replace Abstract: Hallucination, the generation of non factual content by AI systems, poses serious risks in medical contexts, where errors can directly affect patient outcomes. We present MedHal, a large-scale dataset specifically designed to… 7 arXiv — NLP / Computation & Language research 7h ago Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning arXiv:2505.09738v2 Announce Type: replace Abstract: Pretrained language models (LLMs) are often constrained by their fixed tokenization schemes, leading to inefficiencies and performance limitations, particularly for multilingual or specialized applications. This tokenizer… 38 arXiv — NLP / Computation & Language research 7h ago Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries arXiv:2601.01842v2 Announce Type: replace Abstract: Dictionary definitions are an essential resource for learning word senses, but manually creating them is costly. We thus study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for… 12 arXiv — NLP / Computation & Language research 7h ago Affective Flow Language Model for Emotional Support Conversation arXiv:2602.08826v3 Announce Type: replace Abstract: Large language models (LLMs) have advanced emotional support conversation, but existing alignment methods rely mainly on sparse preferences at the response level or outcomes at the dialogue level, providing limited supervision… 36 arXiv — NLP / Computation & Language research 7h ago Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders arXiv:2603.18863v2 Announce Type: replace Abstract: Cross-lingual alignment is often assumed to improve cross-lingual transfer by bringing representations of different languages closer together. However, improvements in representational alignment do not consistently translate… 13 arXiv — NLP / Computation & Language research 7h ago Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR arXiv:2604.06487v3 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection… 34 arXiv — NLP / Computation & Language research 7h ago VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation arXiv:2604.21375v3 Announce Type: replace Abstract: Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without… 38 arXiv — NLP / Computation & Language research 7h ago Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations arXiv:2604.23295v3 Announce Type: replace Abstract: Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open,… 21 arXiv — NLP / Computation & Language research 7h ago UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification arXiv:2605.06221v2 Announce Type: replace Abstract: As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing,… 30 arXiv — NLP / Computation & Language research 7h ago StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction arXiv:2605.06642v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both… 4 Page 2 of 10 · 500 articles ← Newer Older →