News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering arXiv:2608.13160v1 Announce Type: new Abstract: Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved… 15 arXiv — NLP / Computation & Language research 7h ago Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory arXiv:2608.13168v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied for emotional companionship, evaluating their behavior and capabilities in intimate relationships has become a pressing issue. However, existing assessments primarily… 24 arXiv — NLP / Computation & Language research 7h ago GEM: A Generative Embedding Model Bridging Reasoning and Retrieval arXiv:2608.13200v1 Announce Type: new Abstract: Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents,… 32 arXiv — NLP / Computation & Language research 7h ago Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits arXiv:2608.13244v1 Announce Type: new Abstract: Local chemical perception and property reasoning are both essential for understanding how molecular structure determines properties. Current LLM-based chemical reasoning methods either receive SMILES/molecular images together with… 36 arXiv — NLP / Computation & Language research 7h ago Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models arXiv:2608.13258v1 Announce Type: new Abstract: Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated,… 12 arXiv — NLP / Computation & Language research 7h ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 7h ago Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model arXiv:2608.13277v1 Announce Type: new Abstract: We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular… 36 arXiv — NLP / Computation & Language research 7h ago Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form… 38 arXiv — NLP / Computation & Language research 7h ago Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability… 20 arXiv — NLP / Computation & Language research 7h ago It's How You Ask: Gender-Associated Linguistic Bias in LLMs arXiv:2608.13328v1 Announce Type: new Abstract: Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective… 10 arXiv — NLP / Computation & Language research 7h ago RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory arXiv:2608.13334v1 Announce Type: new Abstract: LLM-based agents increasingly rely on external memory to support long-horizon reasoning and interaction. However, the main bottleneck is not simply storing past experience, but recovering the right set of evidence when relevant… 38 arXiv — NLP / Computation & Language research 7h ago CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation arXiv:2608.13387v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation… 33 arXiv — NLP / Computation & Language research 7h ago Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection arXiv:2608.13425v1 Announce Type: new Abstract: Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related… 28 arXiv — NLP / Computation & Language research 7h ago Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity arXiv:2608.13430v1 Announce Type: new Abstract: Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be… 5 arXiv — NLP / Computation & Language research 7h ago Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity arXiv:2608.13484v1 Announce Type: new Abstract: When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative… 25 arXiv — NLP / Computation & Language research 7h ago Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining arXiv:2608.13515v1 Announce Type: new Abstract: Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task… 11 arXiv — NLP / Computation & Language research 7h ago DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data arXiv:2608.13517v1 Announce Type: new Abstract: Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter… 32 arXiv — NLP / Computation & Language research 7h ago SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization arXiv:2608.13538v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial… 19 arXiv — NLP / Computation & Language research 7h ago LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure arXiv:2608.13545v1 Announce Type: new Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this… 6 arXiv — NLP / Computation & Language research 7h ago When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models arXiv:2608.12324v1 Announce Type: cross Abstract: People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about… 37 arXiv — NLP / Computation & Language research 7h ago Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists arXiv:2608.12345v1 Announce Type: cross Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct… 9 arXiv — NLP / Computation & Language research 7h ago From Observation to Intervention: Memory in Brains and Large Language Models arXiv:2608.12377v1 Announce Type: cross Abstract: Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related information is represented, how partial cues recover broader… 19 arXiv — NLP / Computation & Language research 7h ago Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction arXiv:2608.12426v1 Announce Type: cross Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled… 33 arXiv — NLP / Computation & Language research 7h ago SoK: From Generation to Consumption of Privacy Documents in Software Systems arXiv:2608.12511v1 Announce Type: cross Abstract: Privacy documents (e.g., privacy policies) are a central mechanism through which digital services disclose data practices and seek user consent. Over the past decades, research on privacy documents has expanded significantly,… 9 arXiv — NLP / Computation & Language research 7h ago Is this Citation on Point? arXiv:2608.12571v1 Announce Type: cross Abstract: In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting… 30 arXiv — NLP / Computation & Language research 7h ago EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory arXiv:2608.12627v1 Announce Type: cross Abstract: Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are… 11 arXiv — NLP / Computation & Language research 7h ago SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries arXiv:2608.12654v1 Announce Type: cross Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy… 38 arXiv — NLP / Computation & Language research 7h ago Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks arXiv:2608.12713v1 Announce Type: cross Abstract: Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining… 25 arXiv — NLP / Computation & Language research 7h ago Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors arXiv:2608.12746v1 Announce Type: cross Abstract: Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an individual object mention to what the image shows. Most… 4 arXiv — NLP / Computation & Language research 7h ago Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories arXiv:2608.12847v1 Announce Type: cross Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this… 8 arXiv — NLP / Computation & Language research 7h ago Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research arXiv:2608.12984v1 Announce Type: cross Abstract: Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a… 7 arXiv — NLP / Computation & Language research 7h ago TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes arXiv:2608.13057v1 Announce Type: cross Abstract: In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on… 32 arXiv — NLP / Computation & Language research 7h ago Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) arXiv:2608.13063v1 Announce Type: cross Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement -… 25 arXiv — NLP / Computation & Language research 7h ago TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint arXiv:2608.13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce… 30 arXiv — NLP / Computation & Language research 7h ago When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1 arXiv:2608.13237v1 Announce Type: cross Abstract: Multi-round retrieval-augmented generation (RAG) must decide when to stop searching as evidence accumulates. Because the deployed policy is determined by the first STOP on each trajectory, this is a sequential selection problem… 19 arXiv — NLP / Computation & Language research 7h ago MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification arXiv:2608.13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image… 12 arXiv — NLP / Computation & Language research 7h ago MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination arXiv:2608.13476v1 Announce Type: cross Abstract: We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized… 17 arXiv — NLP / Computation & Language research 7h ago OmniScientist: An Omni-Modal Omni-Discipline AI Scientist arXiv:2608.13558v1 Announce Type: cross Abstract: Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not… 6 arXiv — Machine Learning research 1d ago FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting arXiv:2608.11254v1 Announce Type: new Abstract: Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids. All-sky imagers (ASI) provide high-resolution observations of clouds, making them well suited for… 37 arXiv — Machine Learning research 1d ago Why AI Detection Fails for Academic Integrity arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains;… 15 arXiv — Machine Learning research 1d ago Basin: Efficient and Extensible Numerical Optimization in Rust arXiv:2608.11279v1 Announce Type: new Abstract: Basin is a numerical optimization library for the Rust programming language. Numerical optimization is the task of finding the inputs that minimize a function, and it is a fundamental element across the sciences: fitting a model to… 30 arXiv — Machine Learning research 1d ago Federated Learning for Distributed CNC Tool Wear Prediction arXiv:2608.11281v1 Announce Type: new Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability. Machine learning methods have shown potential for this task, but their use in… 7 arXiv — Machine Learning research 1d ago Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction arXiv:2608.11318v1 Announce Type: new Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. We develop a decision-resource view of terminal symmetry: process evidence supplies… 28 arXiv — Machine Learning research 1d ago Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings arXiv:2608.11324v1 Announce Type: new Abstract: This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical, water-cooled chiller plant and its associated air side. Rather than converging… 6 arXiv — Machine Learning research 1d ago Long-Horizon Forecasting of Complete Financial Statements with Forma arXiv:2608.11327v1 Announce Type: new Abstract: Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm… 6 arXiv — NLP / Computation & Language research 1d ago Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport arXiv:2608.11342v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining,… 35 arXiv — Machine Learning research 1d ago Dynamics Models for Offline Hyperparameter Selection in Real-World RL arXiv:2608.11349v1 Announce Type: new Abstract: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly. Prior work has proposed calibration models… 31 arXiv — Machine Learning research 1d ago Market-Information-Aware Gated-LoRA of Foundation Models for Transferable Day-Ahead Electricity Price Forecasting arXiv:2608.11359v1 Announce Type: new Abstract: Electricity price forecasting is crucial for market participants but remains difficult because prices are volatile, market-specific, and closely tied to anticipated system conditions. Existing supervised methods depend largely on… 5 arXiv — NLP / Computation & Language research 1d ago Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter arXiv:2608.11361v1 Announce Type: cross Abstract: Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis. We show that the cost-optimal… 26 arXiv — Machine Learning research 1d ago PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR arXiv:2608.11368v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this cost by assigning budgets to prompts, rollouts, or tokens according to… 26 Page 4 of 10 · 500 articles ← Newer Older →