News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — Machine Learning research 14d ago ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute… 18 arXiv — Machine Learning research 14d ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger arXiv:2607.28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate… 18 arXiv — NLP / Computation & Language research 14d ago Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models arXiv:2607.27421v1 Announce Type: new Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness… 17 arXiv — NLP / Computation & Language research 14d ago Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories arXiv:2607.27595v1 Announce Type: new Abstract: Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how… 16 arXiv — NLP / Computation & Language research 14d ago Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or… 29 arXiv — NLP / Computation & Language research 14d ago Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation arXiv:2607.27816v1 Announce Type: new Abstract: Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable… 17 arXiv — NLP / Computation & Language research 14d ago Causal Discovery with Inverted Self-attention for Multivariate Time Series arXiv:2607.28212v1 Announce Type: new Abstract: Causal discovery in multivariate time series data is challenging due to complex interactions, high dimensionality, and nonlinear dependencies among variables. Existing methods often struggle to capture these complexities, resulting… 30 arXiv — NLP / Computation & Language research 14d ago (Towards) Scalable Reliable Automated Evaluation with Large Language Models arXiv:2607.28282v1 Announce Type: new Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated metrics often fail to capture the complexity and variability inherent in… 10 arXiv — NLP / Computation & Language research 14d ago Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation arXiv:2607.28439v1 Announce Type: new Abstract: Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation… 15 arXiv — NLP / Computation & Language research 14d ago Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments arXiv:2607.28591v1 Announce Type: cross Abstract: Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable… 15 arXiv — NLP / Computation & Language research 14d ago OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation,… 17 Hugging Face Daily Papers research 14d ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Abstract Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct… 37 Hugging Face Daily Papers research 14d ago Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Abstract Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability,… 24 Simon Willison community 14d ago Investigating three real-world incidents in our cybersecurity evaluations Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a sandboxed container and hacked into… 10 TechCrunch — AI news-outlet 14d ago Dili raises $21.7M to bring AI compliance to the infrastructure boom The Series A was led by Khosla Ventures, with participation from Allianz, Rebel Fund, Brick and Mortar Ventures’ Darren Bechtel, and Y Combinator’s Garry Tan. 13 Hugging Face Daily Papers research 15d ago Can AI agents conduct open-ended AI research? Early evidence from two case studies Abstract Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit… 32 Hugging Face Daily Papers research 15d ago CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition Abstract Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical… 14 arXiv — Machine Learning research 15d ago SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning arXiv:2607.26417v1 Announce Type: new Abstract: Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages. Reset curricula address this by starting some training rollouts from easier intermediate states,… 27 arXiv — Machine Learning research 15d ago From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data arXiv:2607.26521v1 Announce Type: new Abstract: Conventional subgroup analyses can yield unstable and difficult-to-interpret conclusions, especially in observational biomedical data where each individual is observed under only one exposure state, true individual treatment… 4 arXiv — Machine Learning research 15d ago Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods arXiv:2607.26625v1 Announce Type: new Abstract: Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. Standard random splitting assumes that both partitions share the same underlying distribution, an… 18 arXiv — Machine Learning research 15d ago Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark arXiv:2607.26993v1 Announce Type: new Abstract: Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pretrained… 22 arXiv — Machine Learning research 15d ago BayesAME: Bayesian Active Model Evaluation arXiv:2607.27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items,… 38 arXiv — Machine Learning research 15d ago Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes arXiv:2607.27188v1 Announce Type: new Abstract: Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation,… 30 arXiv — Machine Learning research 15d ago When benchmark inferences do not compose: Projectibility in AI evaluation arXiv:2607.26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and… 34 arXiv — NLP / Computation & Language research 15d ago Position: Evaluation Scores Are Perishable Knowledge Claims arXiv:2607.26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via… 10 arXiv — NLP / Computation & Language research 15d ago Contrastive ESA: Human Evaluation of Multiple Translations at Once arXiv:2607.26640v1 Announce Type: new Abstract: Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that… 29 arXiv — NLP / Computation & Language research 15d ago TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days… 29 arXiv — NLP / Computation & Language research 15d ago ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models. Despite growing global AI adoption, no systematic evaluation exists for ML… 30 r/MachineLearning community 15d ago Open-source tabular model validation toolkit TanML needs feedback [D] We’re developing TanML, an MIT-licensed automated model-validation toolkit for tabular machine-learning models. TanML runs locally and provides an end-to-end workflow covering data profiling, preprocessing, feature-power ranking, model development, evaluation, drift analysis,… 5 TechCrunch — AI news-outlet 16d ago As AI content floods the internet, Pangram raises $9M to detect it Pangram has raised $9 million to scale its AI detection software. The startup has also released a new AI text detection model, Pangram 4, and an AI image detection model in research preview. 15 Hugging Face Daily Papers research 16d ago PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Abstract We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with… 34 Hugging Face Daily Papers research 16d ago Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking Abstract Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without… 28 arXiv — NLP / Computation & Language research 16d ago CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models arXiv:2607.24999v1 Announce Type: new Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We… 31 arXiv — NLP / Computation & Language research 16d ago Evaluation of forced alignment of code-mixed speech: the case of Hindi-English arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We… 35 arXiv — NLP / Computation & Language research 16d ago Evaluation of Adversarial Robustness in Arabic Language Models arXiv:2607.25814v1 Announce Type: new Abstract: The emergence of the recent outstanding capabilities of Arabic Language Models has opened doors for exposing their vulnerabilities. One of the major security risks associated with such Natural Language Processing models is… 18 arXiv — NLP / Computation & Language research 16d ago AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation arXiv:2607.25881v1 Announce Type: new Abstract: We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics,… 29 arXiv — NLP / Computation & Language research 16d ago Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases arXiv:2607.25933v1 Announce Type: new Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating… 32 arXiv — NLP / Computation & Language research 16d ago RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation arXiv:2607.24772v1 Announce Type: cross Abstract: Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in… 24 arXiv — NLP / Computation & Language research 16d ago CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition arXiv:2607.25294v1 Announce Type: cross Abstract: Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus… 4 arXiv — NLP / Computation & Language research 16d ago Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications arXiv:2607.25642v1 Announce Type: cross Abstract: Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward… 32 arXiv — NLP / Computation & Language research 16d ago Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models arXiv:2607.25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no… 22 arXiv — NLP / Computation & Language research 16d ago A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series arXiv:2607.25947v1 Announce Type: cross Abstract: Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable… 26 arXiv — NLP / Computation & Language research 16d ago Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Methods arXiv:2502.11150v5 Announce Type: replace Abstract: Automatic methods for scoring text readability have been studied for over a century, and are widely used in research and in user-facing applications in many domains. Thus far, the development and evaluation of such methods have… 22 arXiv — NLP / Computation & Language research 16d ago Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore arXiv:2601.15050v5 Announce Type: replace Abstract: Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from \textit{factual myopia}: they relentlessly emphasize factual accuracy yet neglect global logical integrity in long-form answer generation. This… 28 Hugging Face Daily Papers research 17d ago Codifying the Judge: Scalable Evaluation via Program Distillation Abstract LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program… 16 arXiv — Machine Learning research 17d ago Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation arXiv:2607.22766v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks,… 25 arXiv — Machine Learning research 17d ago CC-AOS: Cost- and Horizon-Conditioned Amortized Backward Induction for Finite-Horizon Optimal Stopping arXiv:2607.22774v1 Announce Type: new Abstract: Finite-horizon optimal stopping is a central problem in early time-series classification, where a system must decide at each sequence prefix whether the expected benefit of another observation justifies its acquisition cost.… 14 arXiv — Machine Learning research 17d ago Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs arXiv:2607.22786v1 Announce Type: new Abstract: In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. The financial time series are price series such… 38 arXiv — Machine Learning research 17d ago Online Policy Evaluation for MDPs with Dynamic UBSR Measures arXiv:2607.23030v1 Announce Type: new Abstract: Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restrictive classes of risk measures or rely on access… 10 arXiv — Machine Learning research 17d ago In-Context Learning as Implicit Policy Gradient arXiv:2607.23153v1 Announce Type: new Abstract: Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings,… 35 Page 5 of 10 · 500 articles ← Newer Older →