News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — Machine Learning research 17d ago In-Context Learning as Implicit Policy Gradient arXiv:2607.23153v1 Announce Type: new Abstract: Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings,… 35 arXiv — Machine Learning research 17d ago Transfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization arXiv:2607.23404v1 Announce Type: new Abstract: Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core. Gaussian processes (GPs)… 7 arXiv — NLP / Computation & Language research 17d ago ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation arXiv:2607.23058v1 Announce Type: new Abstract: Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical… 33 arXiv — NLP / Computation & Language research 17d ago Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining arXiv:2607.23175v1 Announce Type: new Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language… 35 arXiv — NLP / Computation & Language research 17d ago Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking arXiv:2607.23514v1 Announce Type: new Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using… 28 arXiv — NLP / Computation & Language research 17d ago BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes arXiv:2607.24137v1 Announce Type: new Abstract: This paper describes the BioSentinel team's participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct,… 8 arXiv — NLP / Computation & Language research 17d ago Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets arXiv:2607.24268v1 Announce Type: new Abstract: Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation… 6 r/MachineLearning community 17d ago Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R] I ran a solo evaluation project benchmarking six current frontier models: GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro, Gemini Flash, and Grok 4.3. I tested tham across 8 established bias/fairness datasets (WinoBias, BBQ Race/Ethnicity, SeeGULL, OpinionsQA, cajcodes… 26 TechCrunch — AI news-outlet 17d ago Enigma raises $70M to make controlling a robot as easy as adjusting the volume The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners. 32 arXiv — Machine Learning research 18d ago Cloud-Native Evaluation-as-a-Service: A Microservices Architecture for Scalable AI Monitoring with Conformal Guarantees arXiv:2607.21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration… 11 arXiv — Machine Learning research 18d ago Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under… 32 arXiv — Machine Learning research 18d ago MissHyper: Restoring Clinical Synchronicity in Missingness-Guided Hypergraph Forecasting arXiv:2607.21922v1 Announce Type: new Abstract: Clinical irregular multivariate time series are shaped not only by physiological dynamics but also by the measurement process that determines when and what to observe. In event-centric models, however, co-timestamp structure can be… 29 arXiv — Machine Learning research 18d ago Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits arXiv:2607.22012v1 Announce Type: new Abstract: Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods… 32 arXiv — Machine Learning research 18d ago An Insight on Evaluation Metrics Under the Imbalanced Case of Anomaly Detection arXiv:2607.22286v1 Announce Type: new Abstract: Anomaly detection is inherently characterised by severe class imbalance, making the interpretation of evaluation metrics challenging. Although metrics such as AUROC, AUPR, F1-score, and MCC are widely used, their values convey… 4 arXiv — NLP / Computation & Language research 18d ago A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models arXiv:2607.21632v1 Announce Type: new Abstract: Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone… 25 arXiv — NLP / Computation & Language research 18d ago Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark arXiv:2607.21685v1 Announce Type: new Abstract: A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings… 13 arXiv — NLP / Computation & Language research 18d ago Agentic Evaluation of Copyright Law Compliance arXiv:2607.21799v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate, reproducing that content. LLM agents should comply with the law, including… 6 arXiv — NLP / Computation & Language research 18d ago Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contamination problems -- and they overwhelmingly measure short interaction histories. We… 18 arXiv — NLP / Computation & Language research 18d ago From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models arXiv:2607.22182v1 Announce Type: new Abstract: Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities… 30 arXiv — NLP / Computation & Language research 18d ago Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from… 37 arXiv — NLP / Computation & Language research 18d ago grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP arXiv:2607.22456v1 Announce Type: new Abstract: Existing lexical distance, similarity, and evaluation metrics operate on Unicode code points, which can misrepresent errors in writing systems where a single grapheme is represented by multiple Unicode code points. We introduce… 37 arXiv — NLP / Computation & Language research 18d ago Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level… 15 arXiv — NLP / Computation & Language research 18d ago DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write… 17 arXiv — NLP / Computation & Language research 18d ago LMEB: Long-horizon Memory Embedding Benchmark arXiv:2603.12572v5 Announce Type: replace Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchmarks, which narrowly focus on traditional passage retrieval and fail to assess… 27 arXiv — NLP / Computation & Language research 18d ago WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47… 25 r/LocalLLaMA community 19d ago CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system prompts, tool schemas, and conversation history. CachyLLama is a llama.cpp… 19 Hugging Face Daily Papers research 20d ago Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Abstract We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running… 30 arXiv — Machine Learning research 21d ago SevDiff: Severity-Conditioned Diffusion for Long-Tail Conflict Trajectory Generation arXiv:2607.20549v1 Announce Type: new Abstract: Trajectory datasets used in ADAS evaluation are heavily biased toward routine driving; genuine vehicle-to-vehicle conflict events are rare, and the rarer the event, the higher the cost when an ADAS system fails to handle it.… 9 arXiv — Machine Learning research 21d ago StabilityBench: Benchmarking Instability in LLMs arXiv:2607.20558v1 Announce Type: new Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poorly understood due to strong context dependence. Current evaluation protocols… 6 arXiv — NLP / Computation & Language research 21d ago Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single answer-order shuffles whose results confound the bias signal with content-level… 35 arXiv — Machine Learning research 21d ago From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python arXiv:2607.21069v1 Announce Type: new Abstract: The original ALPHA benchmark introduced a taxonomy-aware penalty for evaluating CWE-level vulnerability prediction in Python and proposed that the penalty could theoretically also serve as a training signal. This paper provides… 31 arXiv — Machine Learning research 21d ago GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes arXiv:2607.21117v1 Announce Type: new Abstract: Preprocessing blood glucose time-series data is a critical yet often overlooked step in developing data-driven methods for diabetes management, particularly for type 1 diabetes. The lack of standardized preprocessing workflows and… 30 arXiv — NLP / Computation & Language research 21d ago Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models arXiv:2607.20436v1 Announce Type: new Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same… 35 arXiv — NLP / Computation & Language research 21d ago CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation arXiv:2607.20862v1 Announce Type: new Abstract: At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative criteria underlying human preferences in such tasks. To this end, we propose… 21 arXiv — NLP / Computation & Language research 21d ago Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation… 5 arXiv — NLP / Computation & Language research 21d ago A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset arXiv:2607.21274v1 Announce Type: new Abstract: We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. We evaluate sparse (BM25), dense (sentence-transformers), hybrid, and LLM-assisted… 20 arXiv — NLP / Computation & Language research 21d ago Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin arXiv:2607.21332v1 Announce Type: new Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners… 20 arXiv — NLP / Computation & Language research 21d ago An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations arXiv:2607.21424v1 Announce Type: new Abstract: Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluating this… 8 arXiv — NLP / Computation & Language research 21d ago Expectation Alignment of Language Models for Real-World User Expectations arXiv:2607.20485v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model… 34 arXiv — NLP / Computation & Language research 21d ago DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making arXiv:2607.20491v1 Announce Type: cross Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We introduce DFAH-Bench, a replay benchmark that measures observable behavioral… 22 arXiv — NLP / Computation & Language research 21d ago CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning arXiv:2607.20553v1 Announce Type: cross Abstract: Memory Manager models are pivotal in agent systems. Existing methods rely predominantly on LLM-judged synthetic question-answer (QA) pairs, making memory valuation dependent on sampled queries and the downstream reader. To… 25 TechCrunch — AI news-outlet 21d ago AegisAI, founded by former Google security execs, lands $36M to stop AI-driven spear phishing The Series A was led by Battery Ventures, bringing AegisAI total funding to $49 million. 32 TechCrunch — AI news-outlet 21d ago AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs required, it says. 14 arXiv — NLP / Computation & Language research 22d ago Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance arXiv:2607.19386v1 Announce Type: cross Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For… 25 arXiv — Machine Learning research 22d ago Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely… 11 arXiv — Machine Learning research 22d ago Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation arXiv:2607.19855v1 Announce Type: new Abstract: Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget $\varepsilon$ and on a selective choice of perturbation norms. We argue this formulation is… 29 arXiv — Machine Learning research 22d ago Post-Training in Time Series Foundation Models: A Unifying Framework arXiv:2607.20002v1 Announce Type: new Abstract: Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable downstream deployment. Bridging this gap requires further intervention… 20 arXiv — NLP / Computation & Language research 22d ago Reference-Free Evaluation of Reasoning in Open-Ended Question Answering arXiv:2607.19678v1 Announce Type: new Abstract: AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for… 37 arXiv — NLP / Computation & Language research 22d ago Lightweight Person-Place Relation Extraction from Historical Newspapers with Dependency Graphs and Proximity Features arXiv:2607.19718v1 Announce Type: new Abstract: The HIPE-2026 shared task introduces person-place relation extraction from multilingual historical newspapers as a new evaluation track, classifying the at and isAt relations between pre-annotated person and location mentions in… 24 arXiv — NLP / Computation & Language research 22d ago Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking arXiv:2607.19747v1 Announce Type: new Abstract: As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents… 22 Page 6 of 10 · 500 articles ← Newer Older →