arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 2d ago
myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR
arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality…
14 -
arXiv — NLP / Computation & Language research 2d ago
TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification
arXiv:2608.11044v1 Announce Type: new Abstract: Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on large language models (LLMs) struggle to be efficiently…
29 -
arXiv — NLP / Computation & Language research 2d ago
Multiclass Sentiment Analysis for Identifying Political Viewpoints
arXiv:2608.11049v1 Announce Type: new Abstract: The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core…
13 -
arXiv — NLP / Computation & Language research 2d ago
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product:…
38 -
arXiv — NLP / Computation & Language research 2d ago
Attention-Path Fragility as an Uncertainty Signal in Large Language Models
arXiv:2608.11138v1 Announce Type: new Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We…
27 -
arXiv — NLP / Computation & Language research 2d ago
The Illusion of Cross-Lingual Safety in Low-Resource Languages
arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in…
14 -
arXiv — NLP / Computation & Language research 2d ago
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
arXiv:2608.11171v1 Announce Type: new Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc…
21 -
arXiv — NLP / Computation & Language research 2d ago
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
arXiv:2608.11200v1 Announce Type: new Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion…
8 -
arXiv — NLP / Computation & Language research 2d ago
Divergent Response Modes in Frontier Language Models Under Steering Pressure
arXiv:2608.06578v1 Announce Type: cross Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study…
35 -
arXiv — NLP / Computation & Language research 2d ago
TRIBE: Predicting Team Performance via Communication Behavior Ensembles
arXiv:2608.06926v1 Announce Type: cross Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics…
30 -
arXiv — NLP / Computation & Language research 2d ago
OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents
arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that…
31 -
arXiv — NLP / Computation & Language research 2d ago
Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness
arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate; none ask whether the model knew it was hallucinating. We jointly audit…
36 -
arXiv — NLP / Computation & Language research 2d ago
Procedural Fairness Failures in RLHF from Preference Averaging
arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural…
13 -
arXiv — NLP / Computation & Language research 2d ago
Edge Phoneme Recognition for Children's Speech through Age-Aware Training
arXiv:2608.10206v1 Announce Type: cross Abstract: Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a…
4 -
arXiv — NLP / Computation & Language research 2d ago
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
arXiv:2608.10218v1 Announce Type: cross Abstract: AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate…
9 -
arXiv — NLP / Computation & Language research 2d ago
Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output
arXiv:2608.10279v1 Announce Type: cross Abstract: Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, whereas repeated semantic classification of partial text can be costly and unstable. We study a…
21 -
arXiv — NLP / Computation & Language research 2d ago
Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
arXiv:2608.10288v1 Announce Type: cross Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned,…
28 -
arXiv — NLP / Computation & Language research 2d ago
Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking
arXiv:2608.10329v1 Announce Type: cross Abstract: Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive capacity to shape rule text. Existing strategies operate at the rule or…
11 -
arXiv — NLP / Computation & Language research 2d ago
Narrative Keyframing for Generative Creative Writing
arXiv:2608.10337v1 Announce Type: cross Abstract: We introduce narrative keyframing, an interaction technique for AI-assisted creative writing that lets writers specify different types of narrative constraints at selected moments in a story, then use AI to generate intervening…
32 -
arXiv — NLP / Computation & Language research 2d ago
VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation
arXiv:2608.10359v1 Announce Type: cross Abstract: As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual…
23 -
arXiv — NLP / Computation & Language research 2d ago
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
arXiv:2608.10366v1 Announce Type: cross Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and…
19 -
arXiv — NLP / Computation & Language research 2d ago
Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation
arXiv:2608.10392v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared-expert designs preserve reusable knowledge, fine-grained methods vary computation within experts, and dynamic routers…
13 -
arXiv — NLP / Computation & Language research 2d ago
Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry
arXiv:2608.10416v1 Announce Type: cross Abstract: We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realization (Riemann GeoResolver). The Euclidean part establishes three core theorems: (1) circuit…
16 -
arXiv — NLP / Computation & Language research 2d ago
Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents
arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth…
30 -
arXiv — NLP / Computation & Language research 2d ago
Evaluating Rational Contracting in Natural Language
arXiv:2608.10475v1 Announce Type: cross Abstract: The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements…
37 -
arXiv — NLP / Computation & Language research 2d ago
Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models
arXiv:2608.10484v1 Announce Type: cross Abstract: Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2…
21 -
arXiv — NLP / Computation & Language research 2d ago
RadFusion: Towards Threshold-Controllable Radiology Report Generation
arXiv:2608.10505v1 Announce Type: cross Abstract: Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of…
33 -
arXiv — NLP / Computation & Language research 2d ago
InSight-doc: Agentic Visual Perception for Long-Document Understanding
arXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual…
17 -
arXiv — NLP / Computation & Language research 2d ago
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
arXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either train a smaller multi-vector encoder from scratch…
29 -
arXiv — NLP / Computation & Language research 2d ago
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
arXiv:2608.10672v1 Announce Type: cross Abstract: Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users experience these systems, leaving the systems' role in relationship formation poorly…
25 -
arXiv — NLP / Computation & Language research 2d ago
ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
arXiv:2608.10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across…
7 -
arXiv — NLP / Computation & Language research 2d ago
The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces
arXiv:2608.10689v1 Announce Type: cross Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision…
28 -
arXiv — NLP / Computation & Language research 2d ago
Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization
arXiv:2608.10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total…
25 -
arXiv — NLP / Computation & Language research 2d ago
Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report…
18 -
arXiv — NLP / Computation & Language research 2d ago
DuplexWorld: Can voice agents help you get through the day?
arXiv:2608.10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing…
38 -
arXiv — NLP / Computation & Language research 2d ago
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
arXiv:2608.10720v1 Announce Type: cross Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a…
22 -
arXiv — NLP / Computation & Language research 2d ago
When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision
arXiv:2608.10731v1 Announce Type: cross Abstract: Whether an additional general dimension is necessary beyond correlated first-order factors is a property of the population covariance matrix, not of any estimator or design. This research establishes when that property can be…
24 -
arXiv — NLP / Computation & Language research 2d ago
Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences
arXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and…
38 -
arXiv — NLP / Computation & Language research 2d ago
StreamFlow: Dynamic Memory Flows for Streaming Video Understanding
arXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded memory. Yet existing paradigms remain limited:…
18 -
arXiv — NLP / Computation & Language research 2d ago
Mapping and Measuring the Behavioral Evolution of Large Language Models
arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using…
27 -
arXiv — NLP / Computation & Language research 2d ago
Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching
arXiv:2608.11030v1 Announce Type: cross Abstract: Patent retrieval and matching based on large language models (LLMs) play a vital role in intellectual property protection. However, due to the complex structure of patent documents, dense technical terminology, and multi-modal…
10 -
arXiv — NLP / Computation & Language research 2d ago
ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization
arXiv:2608.11045v1 Announce Type: cross Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization…
30 -
arXiv — NLP / Computation & Language research 2d ago
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment…
16 -
arXiv — NLP / Computation & Language research 2d ago
Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
arXiv:2608.11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt…
8 -
arXiv — NLP / Computation & Language research 2d ago
Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders
arXiv:2608.11197v1 Announce Type: cross Abstract: Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We…
22 -
arXiv — NLP / Computation & Language research 2d ago
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to…
37 -
arXiv — NLP / Computation & Language research 2d ago
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
arXiv:2506.03922v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical…
15 -
arXiv — NLP / Computation & Language research 2d ago
InternAgentHarness: A Scalable Synthetic Environment for Enhancing LLM Agentic Abilities
arXiv:2508.08636v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable and diverse environments that support repeated…
11 -
arXiv — NLP / Computation & Language research 2d ago
Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning
arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pretraining priors. For example, models will confidently assert a five-legged dog has four legs; consequently, on the VLMBias…
13 -
arXiv — NLP / Computation & Language research 2d ago
Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety
arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky behaviours for online safety, yet labelling such information is often…
34