arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 3d ago
Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025
arXiv:2609.28487v1 Announce Type: new Abstract: News headlines frame public issues both by what they select and by how they word it, yet computational framing work typically collapses these operations into a single score. We introduce a two-dimensional framework that separates…
8 -
arXiv — NLP / Computation & Language research 3d ago
Reward Hacking Challenges Oversight of Autonomous Research Agents
arXiv:2609.28614v1 Announce Type: new Abstract: Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the…
5 -
arXiv — NLP / Computation & Language research 3d ago
Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks
arXiv:2609.28673v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on…
23 -
arXiv — NLP / Computation & Language research 3d ago
An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection
arXiv:2609.28703v1 Announce Type: new Abstract: Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary…
37 -
arXiv — NLP / Computation & Language research 3d ago
PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs
arXiv:2609.28727v1 Announce Type: new Abstract: Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level…
35 -
arXiv — NLP / Computation & Language research 3d ago
Temporal Taxation Compounds Under Post-Training Compression of Whisper Models
arXiv:2609.28739v1 Announce Type: new Abstract: Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which…
31 -
arXiv — NLP / Computation & Language research 3d ago
Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus
arXiv:2609.28747v1 Announce Type: new Abstract: A language model's confidence in an answer is often read as a proxy for how well it knows the corresponding fact. This manual documents an open toolkit built to test that reading directly: a small causal language model is…
27 -
arXiv — NLP / Computation & Language research 3d ago
Script Choice in LLMs: Evidence for Late-Layer Commitment
arXiv:2609.28784v1 Announce Type: new Abstract: In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit-lens analysis. Our probing experiments reveal a…
32 -
arXiv — NLP / Computation & Language research 3d ago
COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages
arXiv:2609.28826v1 Announce Type: new Abstract: Machine translation (MT) for Indian languages remains constrained by the limited availability of high-quality, Indic-centric parallel corpora and evaluation benchmarks. Existing multilingual resources are largely constructed from…
24 -
arXiv — NLP / Computation & Language research 3d ago
Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding
arXiv:2609.28854v1 Announce Type: new Abstract: Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context…
21 -
arXiv — NLP / Computation & Language research 3d ago
Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms
arXiv:2609.29001v1 Announce Type: new Abstract: Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two…
8 -
arXiv — NLP / Computation & Language research 3d ago
Empath: Tracing Multi-Level Emotion Dynamics in Crisis Counseling Dialogues
arXiv:2609.29056v1 Announce Type: new Abstract: Emotion dynamics are critical for understanding crisis-support conversations, yet most computational work treats emotion as static utterance-level labels. We introduce EMPATH, a framework for understanding affective dynamics in…
24 -
arXiv — NLP / Computation & Language research 3d ago
Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study
arXiv:2609.29090v1 Announce Type: new Abstract: We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated…
6 -
arXiv — NLP / Computation & Language research 3d ago
ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks
arXiv:2609.29102v1 Announce Type: new Abstract: Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning…
38 -
arXiv — NLP / Computation & Language research 3d ago
Tag-Aware Structured Text Translation: Towards a Systematic Understanding
arXiv:2609.29131v1 Announce Type: new Abstract: Internet texts are replete with format tags that carry structural, semantic, and functional meaning. Current large language model (LLM)-based translation systems struggle to balance translation fluency with tag fidelity when…
23 -
arXiv — NLP / Computation & Language research 3d ago
BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech
arXiv:2609.29146v1 Announce Type: new Abstract: Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving…
13 -
arXiv — NLP / Computation & Language research 3d ago
Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space
arXiv:2609.29183v1 Announce Type: new Abstract: Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to…
37 -
arXiv — NLP / Computation & Language research 3d ago
EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards
arXiv:2609.29230v1 Announce Type: new Abstract: End-to-end event extraction remains challenging for large language models as it requires simultaneous identification of event triggers, classification of event types, and extraction of schema-grounded argument spans. We present…
10 -
arXiv — NLP / Computation & Language research 3d ago
Post-Training Leaves Behavioral Shadows on Unrelated Decisions
arXiv:2609.29233v1 Announce Type: new Abstract: We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these…
17 -
arXiv — NLP / Computation & Language research 3d ago
No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow
arXiv:2609.29245v1 Announce Type: new Abstract: Given a large corpus, the questions one might ask can vary -- from "When was the first human heart transplant?" to "What are all the contradictory claims in this literature?" -- but what makes some questions more challenging than…
32 -
arXiv — NLP / Computation & Language research 3d ago
pylazaro: a Python package for anglicism extraction in Spanish
arXiv:2609.29276v1 Announce Type: new Abstract: Lexical borrowings are words from one language that are introduced into another language. Identifying lexical borrowings in text is a relevant task for data-centric fields in Linguistics such as lexicography or corpus linguistics,…
30 -
arXiv — NLP / Computation & Language research 3d ago
Reasoning Instructions Can Break Answer Decoding in Vision--Language Models
arXiv:2609.29278v1 Announce Type: new Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On…
7 -
arXiv — NLP / Computation & Language research 3d ago
Grammatical "grandmother neurons" are rare in LLMs
arXiv:2609.29328v1 Announce Type: new Abstract: Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face…
16 -
arXiv — NLP / Computation & Language research 3d ago
Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams
arXiv:2609.29333v1 Announce Type: new Abstract: One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded…
24 -
arXiv — NLP / Computation & Language research 3d ago
ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
arXiv:2609.29349v1 Announce Type: new Abstract: ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for…
29 -
arXiv — NLP / Computation & Language research 3d ago
Parts-of-Speech as Emergent Categories in SAE Latent Space
arXiv:2609.29362v1 Announce Type: new Abstract: Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test…
4 -
arXiv — NLP / Computation & Language research 3d ago
From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring
arXiv:2609.29370v1 Announce Type: new Abstract: Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are…
20 -
arXiv — NLP / Computation & Language research 3d ago
BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech
arXiv:2609.29371v1 Announce Type: new Abstract: This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining…
8 -
arXiv — NLP / Computation & Language research 3d ago
Likelihood Ranking doesn't Scale Like Prompting in LLMs
arXiv:2609.29390v1 Announce Type: new Abstract: LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still…
21 -
arXiv — NLP / Computation & Language research 3d ago
Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters
arXiv:2609.29397v1 Announce Type: new Abstract: Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block…
38 -
arXiv — NLP / Computation & Language research 3d ago
Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?
arXiv:2609.29410v1 Announce Type: new Abstract: Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in…
28 -
arXiv — NLP / Computation & Language research 3d ago
Controlling Backchannels in Streamable Full-duplex Models
arXiv:2609.29418v1 Announce Type: new Abstract: Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a…
22 -
arXiv — NLP / Computation & Language research 3d ago
Rufus-Air: An Open LLM Post-Training Recipe
arXiv:2609.29421v1 Announce Type: new Abstract: Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search…
31 -
arXiv — NLP / Computation & Language research 3d ago
agentic-ger: terminology recovery in long-form speech using global context
arXiv:2609.29428v1 Announce Type: new Abstract: Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the…
25 -
arXiv — NLP / Computation & Language research 3d ago
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
arXiv:2609.29444v1 Announce Type: new Abstract: Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning,…
18 -
arXiv — NLP / Computation & Language research 3d ago
Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure
arXiv:2609.29445v1 Announce Type: new Abstract: We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be…
9 -
arXiv — NLP / Computation & Language research 3d ago
YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
arXiv:2609.29448v1 Announce Type: new Abstract: We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date,…
30 -
arXiv — NLP / Computation & Language research 3d ago
Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark
arXiv:2609.29479v1 Announce Type: new Abstract: Prospective clinical actions, the follow-ups, orders, referrals, and instructions that deter-mine what happens to a patient next, are annotated today in thin fragments across incom-patible corpora: each records a text span and one…
11 -
arXiv — NLP / Computation & Language research 3d ago
Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report
arXiv:2609.29494v1 Announce Type: new Abstract: Large language models make statements concerning their own "minds". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are;…
36 -
arXiv — NLP / Computation & Language research 3d ago
Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions
arXiv:2609.29496v1 Announce Type: new Abstract: Natural language explanation generation serves as a key mechanism for exposing and evaluating vision-language reasoning. Prior work on explanation-driven vision-language models predominantly follows a post-hoc (answer-first)…
11 -
arXiv — NLP / Computation & Language research 3d ago
PROOF: Profiling Reliability of Object-Level Facts in Large Language Models
arXiv:2609.29504v1 Announce Type: new Abstract: Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for…
14 -
arXiv — NLP / Computation & Language research 3d ago
What a Cross-Model Fixed-Point Census Can and Cannot Arbitrate About Repetition
arXiv:2609.29507v1 Announce Type: new Abstract: Two accounts of neural text degeneration coexist. One locates the cause in the training data -- repetition in the corpus produces repetition in the output, established by training on repetition-sorted data -- the other in the…
5 -
arXiv — NLP / Computation & Language research 3d ago
EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation
arXiv:2609.29511v1 Announce Type: new Abstract: Machine Translation (MT) for low-resource languages remains far behind that of high-resource languages, and the gap is widest in specialised domains, where parallel data is scarce or entirely absent. We present EnSiTa, a trilingual…
27 -
arXiv — NLP / Computation & Language research 3d ago
StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection
arXiv:2609.29549v1 Announce Type: new Abstract: Post-training pipelines must select one language-model policy from many checkpoints, prompts, and decoding rules. Mean evaluator scores can conceal rare failures, whereas simultaneous candidate-wise confidence bounds can be…
16 -
arXiv — NLP / Computation & Language research 3d ago
Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap
arXiv:2609.29559v1 Announce Type: new Abstract: Arabic-Russian machine translation (MT) remains under-explored due to the rich morphology of Arabic and low lexical overlap between the two languages. We benchmark seven fine-tuned neural machine translation (NMT) models against…
11 -
arXiv — NLP / Computation & Language research 3d ago
ModularSQL: A Runtime Guardrail for the Multiplicity Blind Spot in Text-to-SQL
arXiv:2609.29573v1 Announce Type: new Abstract: Text-to-SQL systems are increasingly deployed on production databases, where queries that pass benchmark evaluation can still produce results that distort downstream workflows. Standard set-based execution accuracy (Set-EX)…
36 -
arXiv — NLP / Computation & Language research 3d ago
An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model
arXiv:2609.29618v1 Announce Type: new Abstract: We report an exploratory, single-seed ablation of TALH (Adaptive Latent Hybrid), a decoder-only language model with parallel Multi-head Latent Attention (MLA) and a custom recurrent state-space (SSM) branch. Five variants, spanning…
34 -
arXiv — NLP / Computation & Language research 3d ago
TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification
arXiv:2609.29633v1 Announce Type: new Abstract: We present TTLab's submission to the AlexandriaX-2026 Subtask~3 on Arabic MT error span detection and classification. Our system frames the task as token-level classification over surface forms, preserving character offsets to…
14 -
arXiv — NLP / Computation & Language research 3d ago
Operator Packages, Proposer Strength, and Construction-Family Plateaus in Office-Scale Verified Search
arXiv:2609.29636v1 Announce Type: new Abstract: Verified search, in which a language model proposes programs, a hard evaluator scores them, and selection keeps the best, has recently moved mathematical records; controlled ablations of the proposer-side components remain rare. We…
36 -
arXiv — NLP / Computation & Language research 3d ago
How To Do Things With Prompts
arXiv:2609.29657v1 Announce Type: new Abstract: When users address large language models, they produce directive speech acts whose pragmatic features differ from those of both everyday conversation and traditional human-computer interaction, and these features change as users…
24