News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — Machine Learning research 3d ago Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation arXiv:2609.29931v1 Announce Type: new Abstract: Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities… 17 arXiv — Machine Learning research 3d ago When Temporal Perturbations Act Like Sensor Biases: Label-Free Auditing of Wearable Activity Recognizers arXiv:2609.29937v1 Announce Type: new Abstract: Wearable human-activity recognition (HAR) models operate across sensors, subjects, and backbones, yet a smooth waveform may appear temporal while exploiting a persistent sensor offset primarily. We introduce SpectrumAudit, a… 34 arXiv — Machine Learning research 3d ago MF-SCBO : Multi-fidelity Scalable Constrained Bayesian Optimization arXiv:2609.29941v1 Announce Type: new Abstract: Many real-world optimization problems rely on expensive simulations or experiments, making the efficient use of available data essential. Multi-fidelity optimization of high-dimensional black-box functions subject to black-box… 34 arXiv — NLP / Computation & Language research 3d ago Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025 arXiv:2609.28487v1 Announce Type: new Abstract: News headlines frame public issues both by what they select and by how they word it, yet computational framing work typically collapses these operations into a single score. We introduce a two-dimensional framework that separates… 8 arXiv — NLP / Computation & Language research 3d ago Reward Hacking Challenges Oversight of Autonomous Research Agents arXiv:2609.28614v1 Announce Type: new Abstract: Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the… 5 arXiv — NLP / Computation & Language research 3d ago Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks arXiv:2609.28673v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on… 23 arXiv — NLP / Computation & Language research 3d ago An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection arXiv:2609.28703v1 Announce Type: new Abstract: Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary… 37 arXiv — NLP / Computation & Language research 3d ago PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs arXiv:2609.28727v1 Announce Type: new Abstract: Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level… 35 arXiv — NLP / Computation & Language research 3d ago Temporal Taxation Compounds Under Post-Training Compression of Whisper Models arXiv:2609.28739v1 Announce Type: new Abstract: Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which… 31 arXiv — NLP / Computation & Language research 3d ago Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus arXiv:2609.28747v1 Announce Type: new Abstract: A language model's confidence in an answer is often read as a proxy for how well it knows the corresponding fact. This manual documents an open toolkit built to test that reading directly: a small causal language model is… 27 arXiv — NLP / Computation & Language research 3d ago Script Choice in LLMs: Evidence for Late-Layer Commitment arXiv:2609.28784v1 Announce Type: new Abstract: In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit-lens analysis. Our probing experiments reveal a… 32 arXiv — NLP / Computation & Language research 3d ago COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages arXiv:2609.28826v1 Announce Type: new Abstract: Machine translation (MT) for Indian languages remains constrained by the limited availability of high-quality, Indic-centric parallel corpora and evaluation benchmarks. Existing multilingual resources are largely constructed from… 24 arXiv — NLP / Computation & Language research 3d ago Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding arXiv:2609.28854v1 Announce Type: new Abstract: Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context… 21 arXiv — NLP / Computation & Language research 3d ago Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms arXiv:2609.29001v1 Announce Type: new Abstract: Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two… 8 arXiv — NLP / Computation & Language research 3d ago Empath: Tracing Multi-Level Emotion Dynamics in Crisis Counseling Dialogues arXiv:2609.29056v1 Announce Type: new Abstract: Emotion dynamics are critical for understanding crisis-support conversations, yet most computational work treats emotion as static utterance-level labels. We introduce EMPATH, a framework for understanding affective dynamics in… 24 arXiv — NLP / Computation & Language research 3d ago Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study arXiv:2609.29090v1 Announce Type: new Abstract: We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated… 6 arXiv — NLP / Computation & Language research 3d ago ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks arXiv:2609.29102v1 Announce Type: new Abstract: Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning… 38 arXiv — NLP / Computation & Language research 3d ago Tag-Aware Structured Text Translation: Towards a Systematic Understanding arXiv:2609.29131v1 Announce Type: new Abstract: Internet texts are replete with format tags that carry structural, semantic, and functional meaning. Current large language model (LLM)-based translation systems struggle to balance translation fluency with tag fidelity when… 23 arXiv — NLP / Computation & Language research 3d ago BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech arXiv:2609.29146v1 Announce Type: new Abstract: Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving… 13 arXiv — NLP / Computation & Language research 3d ago Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space arXiv:2609.29183v1 Announce Type: new Abstract: Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to… 37 arXiv — NLP / Computation & Language research 3d ago EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards arXiv:2609.29230v1 Announce Type: new Abstract: End-to-end event extraction remains challenging for large language models as it requires simultaneous identification of event triggers, classification of event types, and extraction of schema-grounded argument spans. We present… 10 arXiv — NLP / Computation & Language research 3d ago Post-Training Leaves Behavioral Shadows on Unrelated Decisions arXiv:2609.29233v1 Announce Type: new Abstract: We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these… 17 arXiv — NLP / Computation & Language research 3d ago No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow arXiv:2609.29245v1 Announce Type: new Abstract: Given a large corpus, the questions one might ask can vary -- from "When was the first human heart transplant?" to "What are all the contradictory claims in this literature?" -- but what makes some questions more challenging than… 32 arXiv — NLP / Computation & Language research 3d ago pylazaro: a Python package for anglicism extraction in Spanish arXiv:2609.29276v1 Announce Type: new Abstract: Lexical borrowings are words from one language that are introduced into another language. Identifying lexical borrowings in text is a relevant task for data-centric fields in Linguistics such as lexicography or corpus linguistics,… 30 arXiv — NLP / Computation & Language research 3d ago Reasoning Instructions Can Break Answer Decoding in Vision--Language Models arXiv:2609.29278v1 Announce Type: new Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On… 7 arXiv — NLP / Computation & Language research 3d ago Grammatical "grandmother neurons" are rare in LLMs arXiv:2609.29328v1 Announce Type: new Abstract: Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face… 16 arXiv — NLP / Computation & Language research 3d ago Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams arXiv:2609.29333v1 Announce Type: new Abstract: One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded… 24 arXiv — NLP / Computation & Language research 3d ago ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts arXiv:2609.29349v1 Announce Type: new Abstract: ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for… 29 arXiv — NLP / Computation & Language research 3d ago Parts-of-Speech as Emergent Categories in SAE Latent Space arXiv:2609.29362v1 Announce Type: new Abstract: Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test… 4 arXiv — NLP / Computation & Language research 3d ago From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring arXiv:2609.29370v1 Announce Type: new Abstract: Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are… 20 arXiv — NLP / Computation & Language research 3d ago BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech arXiv:2609.29371v1 Announce Type: new Abstract: This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining… 8 arXiv — NLP / Computation & Language research 3d ago Likelihood Ranking doesn't Scale Like Prompting in LLMs arXiv:2609.29390v1 Announce Type: new Abstract: LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still… 21 arXiv — NLP / Computation & Language research 3d ago Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters arXiv:2609.29397v1 Announce Type: new Abstract: Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block… 38 arXiv — NLP / Computation & Language research 3d ago Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code? arXiv:2609.29410v1 Announce Type: new Abstract: Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in… 28 arXiv — NLP / Computation & Language research 3d ago Controlling Backchannels in Streamable Full-duplex Models arXiv:2609.29418v1 Announce Type: new Abstract: Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a… 22 arXiv — NLP / Computation & Language research 3d ago Rufus-Air: An Open LLM Post-Training Recipe arXiv:2609.29421v1 Announce Type: new Abstract: Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search… 31 arXiv — NLP / Computation & Language research 3d ago agentic-ger: terminology recovery in long-form speech using global context arXiv:2609.29428v1 Announce Type: new Abstract: Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the… 25 arXiv — NLP / Computation & Language research 3d ago IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis arXiv:2609.29444v1 Announce Type: new Abstract: Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning,… 18 arXiv — NLP / Computation & Language research 3d ago Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure arXiv:2609.29445v1 Announce Type: new Abstract: We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be… 9 arXiv — NLP / Computation & Language research 3d ago YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech arXiv:2609.29448v1 Announce Type: new Abstract: We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date,… 30 arXiv — NLP / Computation & Language research 3d ago Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark arXiv:2609.29479v1 Announce Type: new Abstract: Prospective clinical actions, the follow-ups, orders, referrals, and instructions that deter-mine what happens to a patient next, are annotated today in thin fragments across incom-patible corpora: each records a text span and one… 11 arXiv — NLP / Computation & Language research 3d ago Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report arXiv:2609.29494v1 Announce Type: new Abstract: Large language models make statements concerning their own "minds". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are;… 36 arXiv — NLP / Computation & Language research 3d ago Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions arXiv:2609.29496v1 Announce Type: new Abstract: Natural language explanation generation serves as a key mechanism for exposing and evaluating vision-language reasoning. Prior work on explanation-driven vision-language models predominantly follows a post-hoc (answer-first)… 11 arXiv — NLP / Computation & Language research 3d ago PROOF: Profiling Reliability of Object-Level Facts in Large Language Models arXiv:2609.29504v1 Announce Type: new Abstract: Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for… 14 arXiv — NLP / Computation & Language research 3d ago What a Cross-Model Fixed-Point Census Can and Cannot Arbitrate About Repetition arXiv:2609.29507v1 Announce Type: new Abstract: Two accounts of neural text degeneration coexist. One locates the cause in the training data -- repetition in the corpus produces repetition in the output, established by training on repetition-sorted data -- the other in the… 5 arXiv — NLP / Computation & Language research 3d ago EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation arXiv:2609.29511v1 Announce Type: new Abstract: Machine Translation (MT) for low-resource languages remains far behind that of high-resource languages, and the gap is widest in specialised domains, where parallel data is scarce or entirely absent. We present EnSiTa, a trilingual… 27 arXiv — NLP / Computation & Language research 3d ago StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection arXiv:2609.29549v1 Announce Type: new Abstract: Post-training pipelines must select one language-model policy from many checkpoints, prompts, and decoding rules. Mean evaluator scores can conceal rare failures, whereas simultaneous candidate-wise confidence bounds can be… 16 arXiv — NLP / Computation & Language research 3d ago Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap arXiv:2609.29559v1 Announce Type: new Abstract: Arabic-Russian machine translation (MT) remains under-explored due to the rich morphology of Arabic and low lexical overlap between the two languages. We benchmark seven fine-tuned neural machine translation (NMT) models against… 11 arXiv — NLP / Computation & Language research 3d ago ModularSQL: A Runtime Guardrail for the Multiplicity Blind Spot in Text-to-SQL arXiv:2609.29573v1 Announce Type: new Abstract: Text-to-SQL systems are increasingly deployed on production databases, where queries that pass benchmark evaluation can still produce results that distort downstream workflows. Standard set-based execution accuracy (Set-EX)… 36 arXiv — NLP / Computation & Language research 3d ago An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model arXiv:2609.29618v1 Announce Type: new Abstract: We report an exploratory, single-seed ablation of TALH (Adaptive Latent Hybrid), a decoder-only language model with parallel Multi-head Latent Attention (MLA) and a custom recurrent state-space (SSM) branch. Five variants, spanning… 34 Page 5 of 10 · 500 articles ← Newer Older →