News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — NLP / Computation & Language research 1d ago Benchmarking LLM Judges for Mobile Agent Evaluation arXiv:2608.11434v1 Announce Type: cross Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark… 17 arXiv — NLP / Computation & Language research 1d ago Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation arXiv:2608.11513v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code. While prompt wording and structure are known to influence model performance,… 19 arXiv — NLP / Computation & Language research 1d ago Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning arXiv:2608.11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low… 18 arXiv — NLP / Computation & Language research 1d ago Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder arXiv:2608.11650v1 Announce Type: cross Abstract: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This… 8 arXiv — NLP / Computation & Language research 1d ago FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents arXiv:2608.11683v1 Announce Type: cross Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that… 25 arXiv — NLP / Computation & Language research 1d ago MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques arXiv:2608.11755v1 Announce Type: cross Abstract: Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However,… 8 arXiv — NLP / Computation & Language research 1d ago The Sleeping Agent: What Gist-Based Context Compression Loses and Why arXiv:2608.11775v1 Announce Type: cross Abstract: Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly… 38 arXiv — NLP / Computation & Language research 1d ago How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment arXiv:2608.11816v1 Announce Type: cross Abstract: State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced… 25 arXiv — NLP / Computation & Language research 1d ago Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs arXiv:2608.11830v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench… 19 arXiv — NLP / Computation & Language research 1d ago LookBack: Where and How to Score LVLM Responses via Visual Reference Usage arXiv:2608.11847v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations;… 14 arXiv — NLP / Computation & Language research 1d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused… 25 arXiv — NLP / Computation & Language research 1d ago DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation arXiv:2608.11889v1 Announce Type: cross Abstract: Prompting-based (\textit{i}.\textit{e}., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not changed for the task, face three problems: (\textit{i})~relying on coarse-grained schema… 38 arXiv — NLP / Computation & Language research 1d ago Claim-Level Reliability Assessment for Efficient Test-Time Reasoning arXiv:2608.11994v1 Announce Type: cross Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution… 38 arXiv — NLP / Computation & Language research 1d ago Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence arXiv:2608.12036v1 Announce Type: cross Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly… 26 arXiv — NLP / Computation & Language research 1d ago RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation arXiv:2608.12099v1 Announce Type: cross Abstract: We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size… 21 arXiv — NLP / Computation & Language research 1d ago Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation arXiv:2608.12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has… 5 arXiv — NLP / Computation & Language research 1d ago Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation arXiv:2608.12150v1 Announce Type: cross Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across… 9 arXiv — NLP / Computation & Language research 1d ago VICBench: A Multi-Language Benchmark for Code Vulnerability Detection arXiv:2608.12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the… 7 arXiv — NLP / Computation & Language research 1d ago Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals arXiv:2608.12283v1 Announce Type: cross Abstract: Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction. We study an uncertainty-aware construction that… 15 arXiv — NLP / Computation & Language research 1d ago AVA-Encoder: Towards Agent-Native Video Representation Learning arXiv:2608.12313v1 Announce Type: cross Abstract: Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both… 26 arXiv — NLP / Computation & Language research 1d ago Explainability in Practice: A Survey of Explainable NLP Across Various Domains arXiv:2502.00837v3 Announce Type: replace Abstract: Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o, Gemini, and BERT increasingly inform decisions. The… 26 arXiv — NLP / Computation & Language research 1d ago A Reality Check of Language Models as Formalizers on Constraint Satisfaction Problems arXiv:2505.13252v5 Announce Type: replace Abstract: Recent work shows superior performance when using large language models (LLMs) as formalizers instead of as end-to-end solvers for symbolic reasoning problems. Given the problem description, the LLM generates a formal program… 17 arXiv — NLP / Computation & Language research 1d ago Commonsense on Demand: Generating and Selectively Integrating Commonsense Knowledge for Natural Language Inference arXiv:2507.15100v3 Announce Type: replace Abstract: Natural Language Inference (NLI) determines whether a premise entails, contradicts, or is neutral with respect to a hypothesis. The task is often framed as emulating human inference, in which commonsense knowledge plays a major… 22 arXiv — NLP / Computation & Language research 1d ago Marco-Voice Technical Report arXiv:2508.02038v5 Announce Type: replace Abstract: This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in… 13 arXiv — NLP / Computation & Language research 1d ago Investigating Learner-Aware Design of LLM-Generated Educational Feedback arXiv:2602.11650v2 Announce Type: replace Abstract: Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and information coverage) to support answer revision and learner acceptance… 12 arXiv — NLP / Computation & Language research 1d ago LLM-Powered Automatic Translation and Urgency in Crisis Scenarios arXiv:2602.13452v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. However, their suitability for high-stakes crisis contexts remains insufficiently… 37 arXiv — NLP / Computation & Language research 1d ago Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation arXiv:2603.13891v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring. Across 19 LLMs and two experiments totaling more than 4 million… 31 arXiv — NLP / Computation & Language research 1d ago LLM Router: Rethinking Routing with Prefill Activations arXiv:2603.20895v3 Announce Type: replace Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. We instead route using internal LLM activations, specifically the… 19 Hugging Face Daily Papers research 2d ago 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents Abstract A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning. Generated by thinkingmachines/Inkling-Small We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of… 8 arXiv — Machine Learning research 2d ago Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory arXiv:2608.09997v1 Announce Type: new Abstract: Transformers have had a profound impact on the world of language processing and computer vision. As efforts to answer the million-dollar question of ``How does a Transformer learn?" have been increasing, existing interpretability… 6 arXiv — Machine Learning research 2d ago Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification arXiv:2608.10007v1 Announce Type: new Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and… 30 arXiv — Machine Learning research 2d ago CurveFP: Rational-Radix Logarithmic Datatypes with Closed Products for Language Models arXiv:2608.10010v1 Announce Type: new Abstract: Low-precision datatypes reduce language-model cost, but most formats optimize scalar fidelity while leaving the arithmetic induced by their products unchanged. We introduce CurveFP, a closed-product codebook family that distributes… 22 arXiv — Machine Learning research 2d ago Sheaf-Based Federated Representation Learning arXiv:2608.10016v1 Announce Type: new Abstract: Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model architectures, latent dimensionalities, and local learning… 6 arXiv — Machine Learning research 2d ago DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents arXiv:2608.10037v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the… 38 arXiv — Machine Learning research 2d ago FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows arXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures. However,… 20 arXiv — Machine Learning research 2d ago UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark… 29 arXiv — Machine Learning research 2d ago Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons arXiv:2608.10045v1 Announce Type: new Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. In this problem, the goal is to… 7 arXiv — Machine Learning research 2d ago Detecting Soft Skills in ML Engineering Roles CVs arXiv:2608.10046v1 Announce Type: new Abstract: Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. Job advertisements, surveys, and… 12 arXiv — Machine Learning research 2d ago Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review arXiv:2608.10047v1 Announce Type: new Abstract: In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learning (ML) has largely driven advancements in diagnostics and prognostics, yet purely… 10 arXiv — Machine Learning research 2d ago Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as… 23 arXiv — Machine Learning research 2d ago ChronoSSM: Training for Temporally Aware Representations in Autoregressive State Space Models arXiv:2608.10120v1 Announce Type: new Abstract: Modern sequence models, from Transformers to State Space Models, have enabled powerful generative modeling across diverse domains, yet they are typically trained to predict what happens while treating when it happens as a secondary… 13 arXiv — NLP / Computation & Language research 2d ago Procedural Fairness Failures in RLHF from Preference Averaging arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural… 13 arXiv — Machine Learning research 2d ago SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks arXiv:2608.10144v1 Announce Type: new Abstract: We consider federated parameter efficient fine-tuning of large neural networks with low-rank adaptation (LoRA,~Hu et al.\ 2022). Combining LoRA with federated PEFT introduces challenges absent from either setting alone: clients may… 22 arXiv — Machine Learning research 2d ago The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. We reproduce that result by… 18 arXiv — Machine Learning research 2d ago REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting arXiv:2608.10149v1 Announce Type: new Abstract: Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed… 38 arXiv — Machine Learning research 2d ago Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability arXiv:2608.10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the… 28 arXiv — Machine Learning research 2d ago From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation arXiv:2608.10182v1 Announce Type: new Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When the business goal is incremental impact, as in marketing campaigns, incentives, and… 9 arXiv — Machine Learning research 2d ago ELMER: Evolutionary Language Model that Explores and Refines arXiv:2608.10196v1 Announce Type: new Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a… 7 arXiv — Machine Learning research 2d ago Boundary-Seeking Policy Gradient for Safe Reinforcement Learning arXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies that whenever the constraint is active at… 32 arXiv — Machine Learning research 2d ago A matched-integrator evaluation of Hamiltonian neural networks on pendulum and Kepler dynamics arXiv:2608.10235v1 Announce Type: new Abstract: Hamiltonian Neural Networks (HNNs) parameterize conservative dynamics through a learned scalar Hamiltonian, providing an architectural prior that is absent from generic vector-field neural networks. We evaluate this prior under a… 34 Page 8 of 10 · 500 articles ← Newer Older →