arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 4d ago
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
arXiv:2608.07460v1 Announce Type: new Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as…
27 -
arXiv — NLP / Computation & Language research 4d ago
ADIAS: Automated Design of Interactive Agentic Systems
arXiv:2608.06410v1 Announce Type: cross Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents,…
36 -
arXiv — NLP / Computation & Language research 4d ago
Latent Fact-Checking: Detecting Misinformation through Activation Engineering
arXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level linguistic features or external knowledge retrieval, we examine truthfulness as a…
36 -
arXiv — NLP / Computation & Language research 4d ago
Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing
arXiv:2608.06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing…
29 -
arXiv — NLP / Computation & Language research 4d ago
StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection
arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a…
5 -
arXiv — NLP / Computation & Language research 4d ago
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
arXiv:2608.06501v1 Announce Type: cross Abstract: Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented…
34 -
arXiv — NLP / Computation & Language research 4d ago
Quantization Damage Is Multiplicative, Not Additive
arXiv:2608.06564v1 Announce Type: cross Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's decisions will change at a given bit-width. The damage is silent: a compressed…
38 -
arXiv — NLP / Computation & Language research 4d ago
Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier
arXiv:2608.06571v1 Announce Type: cross Abstract: Deployed vision-language systems often gate their answers on confidence, making confidence robustness relevant to oversight. We study confidence readouts under white-box, image-only attacks constrained to preserve the generated…
11 -
arXiv — NLP / Computation & Language research 4d ago
Online Monitoring and Corrective Steering of Programming Agents
arXiv:2608.06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it. As a result,…
15 -
arXiv — NLP / Computation & Language research 4d ago
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents
arXiv:2608.06735v1 Announce Type: cross Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and code execution. In these settings, the environment…
8 -
arXiv — NLP / Computation & Language research 4d ago
Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inference
arXiv:2608.06752v1 Announce Type: cross Abstract: This paper proposes DKG-MTI, a dual knowledge graph framework for unified multi-task user intent inference from online travel reviews. Existing approaches often rely on hierarchical pipelines that suffer from error propagation or…
38 -
arXiv — NLP / Computation & Language research 4d ago
Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence
arXiv:2608.06778v1 Announce Type: cross Abstract: Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly and does not scale. The ATT&CK taxonomy comprises several hundred attack…
17 -
arXiv — NLP / Computation & Language research 4d ago
Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models
arXiv:2608.06779v1 Announce Type: cross Abstract: Large Language Models (LLMs) have accelerated drug discovery, particularly in the automated design of antimicrobial peptides (AMPs). However, current validation pipelines for peptide generation models overlook historical…
22 -
arXiv — NLP / Computation & Language research 4d ago
LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes
arXiv:2608.06795v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, untrusted adapters introduce a supply-chain threat: a backdoored adapter can cause a model…
23 -
arXiv — NLP / Computation & Language research 4d ago
DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding
arXiv:2608.06869v1 Announce Type: cross Abstract: We describe DAEP, team BIGC's submission to NLPCC 2026 Shared Task 1 Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The task requires retrieving the target video from 50 candidates and localizing…
9 -
arXiv — NLP / Computation & Language research 4d ago
How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots
arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social…
14 -
arXiv — NLP / Computation & Language research 4d ago
DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
arXiv:2608.07067v1 Announce Type: cross Abstract: Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-round memory. Mainstream single-round methods…
34 -
arXiv — NLP / Computation & Language research 4d ago
Modular TTT: Rethinking Test-Time Training as Composable Modules
arXiv:2608.07110v1 Announce Type: cross Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code…
34 -
arXiv — NLP / Computation & Language research 4d ago
Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models
arXiv:2608.07243v1 Announce Type: cross Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM…
20 -
arXiv — NLP / Computation & Language research 4d ago
Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications
arXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for…
38 -
arXiv — NLP / Computation & Language research 4d ago
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning
arXiv:2608.07371v1 Announce Type: cross Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear. We introduce…
7 -
arXiv — NLP / Computation & Language research 4d ago
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
arXiv:2608.07411v1 Announce Type: cross Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a…
36 -
arXiv — NLP / Computation & Language research 4d ago
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
arXiv:2608.07418v1 Announce Type: cross Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of…
13 -
arXiv — NLP / Computation & Language research 4d ago
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and…
22 -
arXiv — NLP / Computation & Language research 4d ago
PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
arXiv:2608.07438v1 Announce Type: cross Abstract: Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents…
15 -
arXiv — NLP / Computation & Language research 4d ago
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
arXiv:2608.07449v1 Announce Type: cross Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent…
27 -
arXiv — NLP / Computation & Language research 4d ago
Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction
arXiv:2411.00028v3 Announce Type: replace Abstract: Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and commercial activity level, which plays an important role in understanding urban regions and…
8 -
arXiv — NLP / Computation & Language research 4d ago
When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
arXiv:2505.16170v4 Announce Type: replace Abstract: We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in their previously generated false assertions. Using model-specific testbeds, we…
38 -
arXiv — NLP / Computation & Language research 4d ago
Better Together: Quantifying the Benefits of AI-Assisted Recruitment
arXiv:2507.08029v2 Announce Type: replace Abstract: Hiring algorithms have mostly scored the materials recruiters already see. Large language models (LLMs) can instead generate new information about candidates by conducting, at scale, structured interviews once reserved for a…
25 -
arXiv — NLP / Computation & Language research 4d ago
Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs
arXiv:2509.16462v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities. Although prior work has examined intrinsic representational bias…
8 -
arXiv — NLP / Computation & Language research 4d ago
Kimi K2.5: Visual Agentic Intelligence
arXiv:2602.02276v2 Announce Type: replace Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This…
28 -
arXiv — NLP / Computation & Language research 4d ago
AfriNLLB: Efficient Translation Models for African Languages
arXiv:2602.09373v2 Announce Type: replace Abstract: In this work, we present AfriNLLB, a series of lightweight models for efficient translation from and into African languages. AfriNLLB supports 15 language pairs (30 translation directions), including Swahili, Hausa, Yoruba,…
37 -
arXiv — NLP / Computation & Language research 4d ago
Improving Attributed Long-form Question Answering with Intent Awareness
arXiv:2603.27435v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they are not exposed to the…
26 -
arXiv — NLP / Computation & Language research 4d ago
How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality
arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack…
17 -
arXiv — NLP / Computation & Language research 4d ago
Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent
arXiv:2604.07269v2 Announce Type: replace Abstract: Clinical expertise improves not only by acquiring medical knowledge, but by accumulating experience that yields reusable diagnostic patterns. Recent LLMs-based diagnostic agents have shown promising progress in clinical…
34 -
arXiv — NLP / Computation & Language research 4d ago
Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing
arXiv:2604.15771v4 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a foundational paradigm for grounding large language models in external knowledge. While adaptive retrieval mechanisms have improved retrieval efficiency, existing approaches…
31 -
arXiv — NLP / Computation & Language research 4d ago
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
arXiv:2604.26355v4 Announce Type: replace Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains underexplored. We observe that reasoning tokens split into two functional…
9 -
arXiv — NLP / Computation & Language research 4d ago
Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
arXiv:2605.02608v2 Announce Type: replace Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers---the…
32 -
arXiv — NLP / Computation & Language research 7d ago
Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
arXiv:2608.05151v1 Announce Type: new Abstract: Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as "why is N2O rising?" or "what happens if I cut…
24 -
arXiv — NLP / Computation & Language research 7d ago
Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
arXiv:2608.05152v1 Announce Type: new Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization. In this study, we…
24 -
arXiv — NLP / Computation & Language research 7d ago
Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability
arXiv:2608.05153v1 Announce Type: new Abstract: GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-robustness analysis that holds the retrieval architecture fixed and varies three orthogonal…
27 -
arXiv — NLP / Computation & Language research 7d ago
RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates
arXiv:2608.05154v1 Announce Type: new Abstract: Rotary positional encoding (RoPE) is a core component of modern language models and has been extended to multimodal LLMs through multidimensional variants such as multimodal RoPE (M-RoPE), which split positional channels into…
22 -
arXiv — NLP / Computation & Language research 7d ago
Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation
arXiv:2608.05155v1 Announce Type: new Abstract: Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideological, and framing dimensions of political discourse -- dimensions that are central to…
9 -
arXiv — NLP / Computation & Language research 7d ago
Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs
arXiv:2608.05156v1 Announce Type: new Abstract: Post-training of large language models optimizes only parameters, while inference-time procedural scaffolds are typically designed independently of parameter training. This disconnect makes it difficult to automatically acquire and…
22 -
arXiv — NLP / Computation & Language research 7d ago
Large Language Models Threaten Double-blind Review
arXiv:2608.05157v1 Announce Type: new Abstract: Double blind peer review serves as the scientific community primary defense against status and affiliation bias. Its effectiveness rests on the assumption that anonymized manuscripts convey scientific merit without revealing their…
4 -
arXiv — NLP / Computation & Language research 7d ago
Safe Evolution with Circuit Anchors
arXiv:2608.05158v1 Announce Type: new Abstract: In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature's solution is \textit{developmental constraints},…
21 -
arXiv — NLP / Computation & Language research 7d ago
SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters
arXiv:2608.05161v1 Announce Type: new Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining remains an unsolved practical challenge. We present SemiAdapt-Instruct, a modular…
16 -
arXiv — NLP / Computation & Language research 7d ago
PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs
arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing…
24 -
arXiv — NLP / Computation & Language research 7d ago
Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages
arXiv:2608.05163v1 Announce Type: new Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test this on an English-source synthetic-PII corpus with five query languages and a…
31 -
arXiv — NLP / Computation & Language research 7d ago
Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study
arXiv:2608.05164v1 Announce Type: new Abstract: Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but whether this geometric similarity has functional consequences for cross-model…
11