arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 4h ago
LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning
arXiv:2608.12321v1 Announce Type: new Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional…
10 -
arXiv — NLP / Computation & Language research 4h ago
What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting
arXiv:2608.12322v1 Announce Type: new Abstract: Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence…
6 -
arXiv — NLP / Computation & Language research 4h ago
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
arXiv:2608.12323v1 Announce Type: new Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI…
27 -
arXiv — NLP / Computation & Language research 4h ago
On Measuring Semantic Preservation in Legal Ontology Learning
arXiv:2608.12326v1 Announce Type: new Abstract: Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on…
26 -
arXiv — NLP / Computation & Language research 4h ago
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
arXiv:2608.12327v1 Announce Type: new Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium,…
34 -
arXiv — NLP / Computation & Language research 4h ago
LoRA-Diffusion: Parameter-Efficient Fine-Tuning via Low-Rank Trajectory Decomposition
arXiv:2608.12328v1 Announce Type: new Abstract: Parameter-efficient fine-tuning methods such as LoRA have transformed the adaptation of large autoregressive language models, enabling task-specific customization with substantially fewer trainable parameters. However, these…
11 -
arXiv — NLP / Computation & Language research 4h ago
AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement
arXiv:2608.12329v1 Announce Type: new Abstract: Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic…
32 -
arXiv — NLP / Computation & Language research 4h ago
Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring
arXiv:2608.12330v1 Announce Type: new Abstract: The detection of online sexism remains an open problem. Sexism detection is inherently subjective, yet most existing systems reduce multi-annotator labels to a single majority decision and treat all instances uniformly. This…
6 -
arXiv — NLP / Computation & Language research 4h ago
Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching
arXiv:2608.12331v1 Announce Type: new Abstract: Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat…
25 -
arXiv — NLP / Computation & Language research 4h ago
Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?
arXiv:2608.12332v1 Announce Type: new Abstract: In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters. In this work, we uncover…
24 -
arXiv — NLP / Computation & Language research 4h ago
Vision-Language Models are Fragile Multilingual Associators
arXiv:2608.12333v1 Announce Type: new Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark…
37 -
arXiv — NLP / Computation & Language research 4h ago
Steering the Language Axis: From Linear Decodability to Causal Control
arXiv:2608.12334v1 Announce Type: new Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood. In this work, we ask whether language identity is merely linearly decodable from…
11 -
arXiv — NLP / Computation & Language research 4h ago
HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings
arXiv:2608.12335v1 Announce Type: new Abstract: Financial question answering over annual reports requires more than retrieving semantically similar passages. It often involves identifying relevant companies and fiscal years, locating standardized filing sections, collecting…
10 -
arXiv — NLP / Computation & Language research 4h ago
StorySpark: Module-wise Evolutionary Search for Story Premise Generation
arXiv:2608.12336v1 Announce Type: new Abstract: A story premise is the creative spark from which a full narrative can grow. Yet LLM-based story generation has mostly emphasized later-stage planning, controllability, coherence, and prose expansion, while premise-level ideation…
32 -
arXiv — NLP / Computation & Language research 4h ago
From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning
arXiv:2608.12337v1 Announce Type: new Abstract: Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off in long-form hallucination RL. Instead of using…
6 -
arXiv — NLP / Computation & Language research 4h ago
SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL
arXiv:2608.12338v1 Announce Type: new Abstract: Text-to-SQL aims to convert natural language questions into executable SQL queries. While memory-based agent system improves complex SQL generation, existing memory design neglect historical experience and suffer from weak…
6 -
arXiv — NLP / Computation & Language research 4h ago
Mimicry without understanding: the origins of decision bias in large language models
arXiv:2608.12339v1 Announce Type: new Abstract: Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training…
35 -
arXiv — NLP / Computation & Language research 4h ago
Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification
arXiv:2608.12340v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), generative data augmentation has attracted considerable attention for imbalanced text classification in natural language processing. However, no empirical benchmark to…
7 -
arXiv — NLP / Computation & Language research 4h ago
The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models
arXiv:2608.12341v1 Announce Type: new Abstract: Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or…
11 -
arXiv — NLP / Computation & Language research 4h ago
Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
arXiv:2608.12342v1 Announce Type: new Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial…
7 -
arXiv — NLP / Computation & Language research 4h ago
Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark
arXiv:2608.12343v1 Announce Type: new Abstract: AI chatbots are widely used by students as knowledge sources, yet LLM benchmarks rarely assess interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exams (Matura) in history - three…
12 -
arXiv — NLP / Computation & Language research 4h ago
Predicting consumer-technology ownership without a diffusion history
arXiv:2608.12344v1 Announce Type: new Abstract: We test whether the perceived attributes of a consumer technology predict how widely it is owned. In a 2022 Prolific survey of US adults (n = 678), respondents rated 65 consumer technologies on six attributes. We then elicited the…
38 -
arXiv — NLP / Computation & Language research 4h ago
New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
arXiv:2608.12361v1 Announce Type: new Abstract: Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "country girl" as a stigmatizing label targeting feminism. Such toxic neologisms appear benign but have evolved into toxic usage in…
29 -
arXiv — NLP / Computation & Language research 4h ago
Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities
arXiv:2608.12374v1 Announce Type: new Abstract: Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation…
28 -
arXiv — NLP / Computation & Language research 4h ago
Query Timing Produces Opposite Positional Biases Between LLMs and Humans
arXiv:2608.12387v1 Announce Type: new Abstract: Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and…
11 -
arXiv — NLP / Computation & Language research 4h ago
Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
arXiv:2608.12391v1 Announce Type: new Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input…
36 -
arXiv — NLP / Computation & Language research 4h ago
DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution
arXiv:2608.12486v1 Announce Type: new Abstract: Large language models (LLMs) cannot retain post-deployment experience without parameter updates. We introduce DIVE, a diversity-driven framework that enables frozen LLMs to improve by evolving persistent natural-language skills…
36 -
arXiv — NLP / Computation & Language research 4h ago
Intensional Anaphora
arXiv:2608.12598v1 Announce Type: new Abstract: Intensional operators are often treated as quantifiers over possible worlds, parallel to the treatment of determiners as quantifiers over individuals. Yet individuals introduced in intensional contexts cannot serve as antecedents…
19 -
arXiv — NLP / Computation & Language research 4h ago
When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
arXiv:2608.12623v1 Announce Type: new Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the defender has only clean calibration data without trigger information but…
38 -
arXiv — NLP / Computation & Language research 4h ago
LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning
arXiv:2608.12626v1 Announce Type: new Abstract: Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these settings, finite attention resources prevent the model from maintaining strategic coherence…
29 -
arXiv — NLP / Computation & Language research 4h ago
Novels generated by language models show compressed formal variation
arXiv:2608.12630v1 Announce Type: new Abstract: While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as…
33 -
arXiv — NLP / Computation & Language research 4h ago
Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection
arXiv:2608.12652v1 Announce Type: new Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or…
7 -
arXiv — NLP / Computation & Language research 4h ago
ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
arXiv:2608.12720v1 Announce Type: new Abstract: While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits…
17 -
arXiv — NLP / Computation & Language research 4h ago
PatientAct: Theory-Grounded Mental Health Client Simulation
arXiv:2608.12750v1 Announce Type: new Abstract: LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept…
30 -
arXiv — NLP / Computation & Language research 4h ago
ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization
arXiv:2608.12756v1 Announce Type: new Abstract: Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. We introduce ReconSpan, which divides text into chunks that a backward decoder can…
13 -
arXiv — NLP / Computation & Language research 4h ago
ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts
arXiv:2608.12776v1 Announce Type: new Abstract: This paper introduces ViTOED, a novel dataset for target-oriented emotion detection in Vietnamese social media texts. The ViTOED comprises 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target,…
23 -
arXiv — NLP / Computation & Language research 4h ago
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
arXiv:2608.12779v1 Announce Type: new Abstract: Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors.…
14 -
arXiv — NLP / Computation & Language research 4h ago
FastThaiG2P: Lightning-fast Thai Grapheme-to-phoneme Conversion for Voice Agent Pipelines
arXiv:2608.12814v1 Announce Type: new Abstract: FastThaiG2P provides sub-millisecond Thai grapheme-to-phoneme conversion for text-to-speech pipelines (International Phonetic Alphabet and Kokoro-TTS conventions) using a PyThaiNLP-tokenized, extensible dictionary and normalization…
31 -
arXiv — NLP / Computation & Language research 4h ago
From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
arXiv:2608.12836v1 Announce Type: new Abstract: Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and…
8 -
arXiv — NLP / Computation & Language research 4h ago
AQuA: Recursively Self-Improving Quantitative Trading Research Agents
arXiv:2608.12841v1 Announce Type: new Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We…
33 -
arXiv — NLP / Computation & Language research 4h ago
Falsehood and Impossibility Are Different Directions in an AI's Representation of Language
arXiv:2608.12852v1 Announce Type: new Abstract: Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study…
8 -
arXiv — NLP / Computation & Language research 4h ago
The Embedder's Dilemma: LLMs Are Better, but at What Cost?
arXiv:2608.12875v1 Announce Type: new Abstract: Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks…
24 -
arXiv — NLP / Computation & Language research 4h ago
When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory
arXiv:2608.12888v1 Announce Type: new Abstract: Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any question is asked. We ask how much of that benefit…
32 -
arXiv — NLP / Computation & Language research 4h ago
BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian
arXiv:2608.12894v1 Announce Type: new Abstract: Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian…
35 -
arXiv — NLP / Computation & Language research 4h ago
Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code
arXiv:2608.12905v1 Announce Type: new Abstract: The behavior of contemporary generative Large Language Models (LLMs) is directly shaped by prompts, unstructured texts that describe the desired output and model behavior. In this paper we argue that prompts are linguistic objects…
35 -
arXiv — NLP / Computation & Language research 4h ago
Decoupled Contrastive Decoding via Expert-Aligned Drafting
arXiv:2608.12913v1 Announce Type: new Abstract: Contrastive Decoding (CD) improves generation quality, but its amateur-model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal-alignment question: should the contrastive signal shape the…
32 -
arXiv — NLP / Computation & Language research 4h ago
Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization
arXiv:2608.12953v1 Announce Type: new Abstract: Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression…
14 -
arXiv — NLP / Computation & Language research 4h ago
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
arXiv:2608.12990v1 Announce Type: new Abstract: Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory systems typically rely on eager consolidation, invoking LLMs after each interaction to extract, summarize, or update…
20 -
arXiv — NLP / Computation & Language research 4h ago
HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA
arXiv:2608.13004v1 Announce Type: new Abstract: Knowledge-base question answering (KBQA) systems rely on effective retrieval and reasoning mechanisms to generate accurate answers from external knowledge sources. However, developing reliable KBQA systems for low-resource…
19 -
arXiv — NLP / Computation & Language research 4h ago
EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval
arXiv:2608.13006v1 Announce Type: new Abstract: Multi-hop retrieval must recover passages that provide sufficient evidence together. An initial passage often resolves an entity or relation implicit in the question, making the missing evidence easier to describe only after…
10