News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 24d ago Otap:Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories arXiv:2607.17082v1 Announce Type: cross Abstract: Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current evaluation metrics reduce such a trajectory to a binary success flag or compare it against… 30 arXiv — NLP / Computation & Language research 24d ago How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv:2607.17152v1 Announce Type: cross Abstract: Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a… 22 arXiv — NLP / Computation & Language research 24d ago DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments arXiv:2607.17291v1 Announce Type: cross Abstract: Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents… 13 arXiv — NLP / Computation & Language research 24d ago It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation arXiv:2607.17627v1 Announce Type: cross Abstract: Prior work on AI-assisted information evaluation has largely focused on what AI systems communicate, comparing explanation types and formats, with responses predominantly cast in directive rhetoric where the system delivers a… 10 arXiv — NLP / Computation & Language research 24d ago WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting arXiv:2607.18084v1 Announce Type: cross Abstract: Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark… 16 arXiv — NLP / Computation & Language research 24d ago LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks arXiv:2607.18110v1 Announce Type: cross Abstract: Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential… 22 Hugging Face Daily Papers research 24d ago LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Abstract Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the… 30 Hugging Face Daily Papers research 24d ago Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Abstract Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational… 32 Hugging Face Daily Papers research 24d ago Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark Abstract Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics fail regularly: acquisition systems start late, individual sensors die, and signal dropouts cluster precisely when a… 13 Hugging Face Daily Papers research 25d ago Loop the Loopies! Abstract We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped Transformers have long faced a… 16 arXiv — Machine Learning research 25d ago AI Trading: Evaluating Large Language Models for Technical Market Analysis arXiv:2607.15414v1 Announce Type: new Abstract: Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs:… 17 arXiv — Machine Learning research 25d ago Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data arXiv:2607.15606v1 Announce Type: new Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing, yet a generator can reproduce every marginal and every foreign-key relationship while emitting timestamps that run backwards or repeat,… 11 arXiv — Machine Learning research 25d ago Scaling Time Series Classification via XAI-Driven Data Reduction arXiv:2607.15774v1 Announce Type: new Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI,… 10 arXiv — Machine Learning research 25d ago Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes arXiv:2607.15799v1 Announce Type: new Abstract: Industrial processes often generate complex, interdependent time-series data from multiple sensors across multiple stages, forming complex dependencies among variables and process stages. Effective monitoring and timely anomaly… 6 arXiv — Machine Learning research 25d ago Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation arXiv:2607.16019v1 Announce Type: new Abstract: AI systems increasingly retrieve from records that revise themselves: issue threads, encyclopedic histories, policy logs, and long conversations. The challenge is not only finding relevant evidence, but deciding which claims remain… 20 arXiv — Machine Learning research 25d ago MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation arXiv:2607.15299v1 Announce Type: cross Abstract: In this paper, we propose MLLM-DataEngine, a novel closed-loop system that bridges data generation, model training, and evaluation. Within each loop iteration, the MLLM-DataEngine first analyzes the weakness of the model based on… 29 arXiv — NLP / Computation & Language research 25d ago Loop the Loopies! arXiv:2607.16051v1 Announce Type: new Abstract: We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active… 5 arXiv — NLP / Computation & Language research 25d ago AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation arXiv:2607.16010v1 Announce Type: cross Abstract: Governments are increasingly mandating that LLM-generated content carry watermarks. The EU AI Act calls for markings that are "sufficiently reliable and robust." California's SB 942 requires disclosure that is "permanent or… 28 arXiv — NLP / Computation & Language research 25d ago LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition arXiv:2607.13347v2 Announce Type: replace Abstract: LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a… 28 arXiv — NLP / Computation & Language research 25d ago Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy arXiv:2602.17229v2 Announce Type: replace-cross Abstract: The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. This study investigates the internal neural representations of cognitive complexity… 30 Hugging Face Daily Papers research 27d ago Rethinking the Evaluation of Harness Evolution for Agents Abstract We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental… 17 TechCrunch — AI news-outlet 27d ago Databricks hits $188B valuation, extending its run as AI’s favorite second act Databricks has remade its image into an AI company and has published research on the cost savings of open weight AI models for coding. 25 arXiv — NLP / Computation & Language research 28d ago Privacy Leakage in Federated Learning in Radiology Reports: A Comparative Evaluation of Tokenizer-Driven Privacy Risks arXiv:2607.14205v1 Announce Type: cross Abstract: Federated learning (FL) enables multi-institutional training on clinical text without sharing raw data, but gradient inversion can reconstruct sensitive information from shared model updates. The extent of this leakage for… 21 arXiv — Machine Learning research 28d ago Active Real-World Factor-Based Evaluation for Generalist Robot Policies arXiv:2607.14439v1 Announce Type: new Abstract: Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks. However, rigorously evaluating these policies remains a fundamental challenge. Real-world… 18 arXiv — Machine Learning research 28d ago Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning arXiv:2607.14817v1 Announce Type: new Abstract: Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to… 24 arXiv — Machine Learning research 28d ago Kernel weighted importance sampling for off-policy evaluation in contextual bandits arXiv:2607.15067v1 Announce Type: new Abstract: This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically… 13 arXiv — Machine Learning research 28d ago Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models arXiv:2607.14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must… 5 arXiv — NLP / Computation & Language research 28d ago Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs arXiv:2607.14099v1 Announce Type: new Abstract: Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation… 10 arXiv — NLP / Computation & Language research 28d ago Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation arXiv:2607.14109v1 Announce Type: new Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation… 8 arXiv — NLP / Computation & Language research 28d ago DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA arXiv:2607.14400v1 Announce Type: new Abstract: This paper describes DS@GT ARC's submission to the CLEF 2026 LongEval Task 4 on Retrieval-Augmented Generation (RAG). In this submission, we examine a divergence between traditional natural language evaluation metrics and citation… 23 arXiv — NLP / Computation & Language research 28d ago How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts arXiv:2607.14591v1 Announce Type: new Abstract: This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best… 6 arXiv — NLP / Computation & Language research 28d ago Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays arXiv:2607.14605v1 Announce Type: new Abstract: This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and… 13 arXiv — NLP / Computation & Language research 28d ago Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration arXiv:2607.14399v1 Announce Type: cross Abstract: Evaluations of language-model honesty read the model's verdicts as evidence about the model. We test the instrument instead. We built a text-adventure world where the game engine, not any model, knows whether the quest can be… 19 arXiv — NLP / Computation & Language research 28d ago WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays arXiv:2607.14524v1 Announce Type: cross Abstract: This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated… 23 arXiv — NLP / Computation & Language research 28d ago Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy arXiv:2607.15176v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis).… 20 arXiv — NLP / Computation & Language research 28d ago Scaling Evaluation-time Compute with Reasoning Models as Evaluators arXiv:2503.19877v3 Announce Type: replace Abstract: As language model (LM) outputs get more and more natural, it is becoming more difficult than ever to evaluate their quality. Simultaneously, increasing LMs' "thinking" time through scaling test-time compute has proven an… 8 Hugging Face Daily Papers research 28d ago MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Abstract Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated… 11 Hugging Face Daily Papers research 28d ago KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Abstract Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the… 14 r/MachineLearning community 28d ago Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R] Hi everyone, I've been working independently on a recurrent architecture called **DABSN (Dynamic Adaptive Bias State Network)** for the past several months, and I finally reached the point where I feel comfortable sharing the first preprint. The paper is mainly about the… 7 VentureBeat — AI news-outlet 28d ago The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty… 26 r/LocalLLaMA community 28d ago To KL Diverge, or Not to KL Diverge: A Question for Quants Hey r/LocalLLaMA ! Apparently, if you draw enough arrows between proxy rankings like KLD, perplexity, and BPW, and real deployment measurements, quantization evaluation starts to look like abstract modern art. Check it out in the second figure! TL;DR: KLD and perplexity can help… 12 TechCrunch — AI news-outlet 28d ago How a former DeepMind researcher raised at a $300M pre-seed valuation before launching a product Drawing on more than a decade spent helping build some of the world's most influential AI systems, including research that later informed the development of ChatGPT, Andrew Dai explains why he believes visual AI is one of the next major frontiers in artificial intelligence. 14 Hugging Face Daily Papers research 29d ago From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Abstract AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for predefined goals such as… 29 Hugging Face Daily Papers research 29d ago AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Abstract As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant… 20 arXiv — Machine Learning research 29d ago Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges arXiv:2607.13045v1 Announce Type: new Abstract: Federated Learning (FL) has emerged as a key paradigm for privacy-preserving collaborative model training across distributed and heterogeneous data sources. By keeping raw data local, FL addresses data confidentiality concerns, yet… 18 arXiv — Machine Learning research 29d ago HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration arXiv:2607.13155v1 Announce Type: new Abstract: Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance target-relevant activity, synthetic accessibility, physicochemical properties,… 21 arXiv — Machine Learning research 29d ago Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation arXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision, a cost-aware evaluation framework for multimodal AI agents on quantum circuit… 11 TechCrunch — AI news-outlet 29d ago Applied Computing wants to give oil and gas operators an AI model for the entire plant Applied Computing has raised a $20M Series A to build a foundation AI model for the oil, gas and petrochemical industry. 24 TechCrunch — AI news-outlet 1mo ago Rime picks up $24M Series A to help enterprises field customer calls Rime is handling over 100 million calls each month across multiple companies 10 TechCrunch — AI news-outlet 1mo ago Indian AI coding startup Emergent becomes a unicorn with $130M Series C The startup has reached a $120 million annualized revenue run rate and more than 200,000 paying customers. 37 Page 8 of 10 · 500 articles ← Newer Older →