News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow Hugging Face Daily Papers research 1mo ago Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Abstract Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the… 38 arXiv — Machine Learning research 1mo ago From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness arXiv:2607.12166v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predominantly on correlational recovery metrics -- cosine similarity between… 5 arXiv — Machine Learning research 1mo ago ReDiTT: Retrieval Augmented Conditional Diffusion Transformers for Asynchronous Time Series arXiv:2607.12391v1 Announce Type: new Abstract: We present a diffusion based model for asynchronous time series prediction, where the goal is to predict the next inter event time and event type. To address the inherent uncertainty of future events, we introduce ReDiTT, a… 17 arXiv — Machine Learning research 1mo ago Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection arXiv:2607.12454v1 Announce Type: new Abstract: Multivariate Time Series Anomaly Detection (MTSAD) is essential for reliability and safety in domains such as industrial process monitoring and financial risk management, yet conventional approaches rely on application-specific… 13 arXiv — Machine Learning research 1mo ago Sample Efficient Generative Optimization for Molecular Design arXiv:2607.12488v1 Announce Type: new Abstract: Molecular optimization in drug discovery, materials design, and catalysis requires searching vast chemical spaces under tight evaluation budgets, since high-fidelity oracles and experimental measurements are costly. The practical… 37 arXiv — Machine Learning research 1mo ago Lightweight Multi-Scale Anomaly Detection for Resource-Constrained Edge Devices arXiv:2607.12599v1 Announce Type: new Abstract: Time-series anomaly detection is increasingly important in IoT systems, sensor networks, and edge monitoring applications, where models must operate under strict constraints on memory, latency, and power consumption. While recent… 36 arXiv — Machine Learning research 1mo ago Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark arXiv:2607.11915v1 Announce Type: cross Abstract: Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics fail regularly: acquisition systems start late, individual sensors die, and… 19 arXiv — Machine Learning research 1mo ago Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection arXiv:2607.11969v1 Announce Type: cross Abstract: Point-adjustment (PA), long the default scoring protocol in time-series anomaly detection (TSAD), was shown by Kim et al. (2022) to award near-perfect F1 to random scores. The field migrated to replacement metrics: PA%K,… 23 arXiv — NLP / Computation & Language research 1mo ago FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality arXiv:2607.12252v1 Announce Type: new Abstract: Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human experts to define and execute high-quality rubrics. We address this problem by… 22 arXiv — NLP / Computation & Language research 1mo ago Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction arXiv:2607.12835v1 Announce Type: new Abstract: Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repository comparison is prone to hallucination. However,… 18 arXiv — NLP / Computation & Language research 1mo ago LLM Judges Can Be Too Generous When There Is No Reference Answer arXiv:2607.12885v1 Announce Type: new Abstract: LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable. However, can they reliably assess in such evaluation setups? We explore this… 32 arXiv — NLP / Computation & Language research 1mo ago The Sound of Absence: Audio-Language Embedding Models Struggle with Negation arXiv:2607.12290v1 Announce Type: cross Abstract: Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated… 5 arXiv — NLP / Computation & Language research 1mo ago Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents arXiv:2607.12790v1 Announce Type: cross Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not. We… 4 arXiv — NLP / Computation & Language research 1mo ago Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks arXiv:2511.04689v3 Announce Type: replace Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely on average accuracy over… 28 arXiv — NLP / Computation & Language research 1mo ago Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective arXiv:2603.14217v3 Announce Type: replace Abstract: In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by coherence, consistency, and shared understanding. However, many systems for… 21 arXiv — NLP / Computation & Language research 1mo ago Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory arXiv:2603.25112v2 Announce Type: replace Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive… 28 arXiv — NLP / Computation & Language research 1mo ago From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models arXiv:2604.26052v4 Announce Type: replace Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide how risk changes between prompt and response.… 37 TechCrunch — AI news-outlet 1mo ago The founder of Hinge raised $18M to build a new AI dating service, Overtone Overtone describes itself as "a voice- and audio-forward service, enabled by AI, that provides highly curated introductions." 26 r/MachineLearning community 1mo ago I trained a vision-language model to play Snake, and so can you. [P] I built this Snake demo to show how easy it can be to go from data preparation to training and evaluation with FeynRL. The model is overkill for Snake, but that’s not the point. This example walks through the full VLM training pipeline in a simple, visual, and fun setting,… 19 Hugging Face Daily Papers research 1mo ago MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning Abstract Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific… 5 arXiv — Machine Learning research 1mo ago Position: Every Ground Truth is a Human Construction, not an Objective Truth arXiv:2607.09668v1 Announce Type: new Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models. This position paper argues that ground truths are not neutral objective measurements that are naturally… 37 arXiv — Machine Learning research 1mo ago Exploratory Analysis of Deep Learning Models for Forecasting Meteorological Parameters in the Agricultural Sector arXiv:2607.10208v1 Announce Type: new Abstract: Accurate meteorological forecasting is essential for agricultural planning, irrigation management, and environmental decision support. This study conducts a comparative evaluation of recurrent and hybrid deep learning architectures… 20 arXiv — Machine Learning research 1mo ago Multi-Scale Convolution with Optimal Transport Attention Effect on Multivariate Time Series arXiv:2607.10740v1 Announce Type: new Abstract: The analysis of Multivariate Time Series (MTS) plays an important role in a lot of real-world practical applications, but it still remains some challenging problem about capturing multi-granularity structural patterns and… 23 arXiv — NLP / Computation & Language research 1mo ago CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series arXiv:2607.09880v1 Announce Type: new Abstract: Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal… 5 arXiv — NLP / Computation & Language research 1mo ago Index SLM Technical Report arXiv:2607.09885v1 Announce Type: new Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili. The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion… 27 arXiv — NLP / Computation & Language research 1mo ago RouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation arXiv:2607.09908v1 Announce Type: new Abstract: Recommender systems increasingly face a choice among heterogeneous agents -- collaborative filters, sequential models, content-based retrievers, and LLM-based rerankers -- yet no single agent is uniformly best. We study this choice… 27 arXiv — NLP / Computation & Language research 1mo ago CAFE: A Compound-AI Factorial Evaluation Framework arXiv:2607.10380v1 Announce Type: new Abstract: We introduce CAFE (Compound-AI Factorial Evaluation), an open-source platform that brings design of experiments to the evaluation of compound AI systems (CAIS). Such systems expose many interchangeable choices - e.g. which… 6 arXiv — NLP / Computation & Language research 1mo ago Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG arXiv:2607.10626v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to identify. We introduce Eval-Pair Matrix, a controlled… 16 arXiv — NLP / Computation & Language research 1mo ago Knowledge Distillation for Automated AI Tutor Evaluation arXiv:2607.10647v1 Announce Type: new Abstract: The rapid integration of Large Language Models (LLMs) into K-12 and higher education has outpaced the development of reliable methods for evaluating their pedagogical quality. As the research community starts to explore the space… 37 arXiv — NLP / Computation & Language research 1mo ago Capabilities of Claude Fable 5 on Biomedical Challenge Problems arXiv:2607.10849v1 Announce Type: new Abstract: Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language… 20 arXiv — NLP / Computation & Language research 1mo ago ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers arXiv:2607.11074v1 Announce Type: new Abstract: Large language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations. We introduce ResearchQA, a benchmark of 6,211… 25 arXiv — NLP / Computation & Language research 1mo ago Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results arXiv:2607.11183v1 Announce Type: new Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible responses. We study inference-time feed-forward network (FFN) intervention… 10 arXiv — NLP / Computation & Language research 1mo ago Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points arXiv:2607.11363v1 Announce Type: new Abstract: Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not… 17 arXiv — NLP / Computation & Language research 1mo ago GEIS: A Generation-Evaluation-Improvement Loop of Agent Skills for Long-Form Article Generation arXiv:2607.11503v1 Announce Type: new Abstract: Long-form article generation remains difficult for large language models because it combines long context, long instructions, and long outputs. Existing multi-agent pipelines such as STORM improve information coverage by simulating… 19 arXiv — NLP / Computation & Language research 1mo ago MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning arXiv:2607.11736v1 Announce Type: new Abstract: Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct… 20 arXiv — NLP / Computation & Language research 1mo ago A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol arXiv:2607.11873v1 Announce Type: new Abstract: Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented… 22 arXiv — NLP / Computation & Language research 1mo ago Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks arXiv:2607.09739v1 Announce Type: cross Abstract: We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised… 12 Hugging Face Daily Papers research 1mo ago AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Abstract Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity:… 11 TechCrunch — AI news-outlet 1mo ago Video generation startup PixVerse raises $439M, valuation soars past $2B Singapore-based video generation startup PixVerse closed a Series C extension on the strength of 15 million monthly active users, it said. 14 TechCrunch — AI news-outlet 1mo ago Hermes agent maker Nous Research in talks for new funding at $1.5B valuation The company is raising at least $75 million, led by Robot, with significant participation from USV and other prominent investors. 37 MIT Technology Review — AI news-outlet 1mo ago What Anthropic’s latest AI discovery does—and doesn’t—show This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and… 30 arXiv — Machine Learning research 1mo ago HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning arXiv:2607.08784v1 Announce Type: new Abstract: Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets,… 35 arXiv — Machine Learning research 1mo ago TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning arXiv:2607.08940v1 Announce Type: new Abstract: Time series reasoning is essential for real-world problem-solving. While both Large Language Models (LLMs) and Vision-Language Models (VLMs) can reason about time-series data, their capabilities are complementary: LLMs process time… 4 arXiv — Machine Learning research 1mo ago FairSelect: A Systematic Evaluation of Multi-Level and Intersectional Algorithmic Fairness arXiv:2607.08953v1 Announce Type: new Abstract: Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for… 29 arXiv — Machine Learning research 1mo ago NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision arXiv:2607.08961v1 Announce Type: new Abstract: Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative,… 17 arXiv — Machine Learning research 1mo ago Federated Low-Rank Koopman Learning for Multivariate Time-Series Anomaly Detection in IoT Systems arXiv:2607.08978v1 Announce Type: new Abstract: Distributed IoT systems generate multivariate time-series streams for monitoring physical assets, servers, and embedded sensing platforms. Detecting abnormal temporal behavior is critical for fault diagnosis, predictive… 37 arXiv — Machine Learning research 1mo ago Temporal Knowledge Graph Forecasting under Distribution Shifts: A Synthetic Evaluation arXiv:2607.09232v1 Announce Type: new Abstract: Temporal knowledge graphs (TKGs) represent evolving relational systems, whose underlying data-generating processes often change over time. Yet, TKG forecasting models are commonly evaluated only on empirical benchmark datasets that… 22 arXiv — Machine Learning research 1mo ago SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets arXiv:2607.08681v1 Announce Type: cross Abstract: As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve… 5 arXiv — Machine Learning research 1mo ago Secure-by-Disguise: A Systematic Evaluation of Image Disguising for Confidential Medical Image Modeling arXiv:2607.08867v1 Announce Type: cross Abstract: Cloud-based deep learning enables large-scale medical image analysis but raises significant privacy concerns when sensitive patient images are outsourced for model development. Image disguising has recently emerged as a promising… 7 arXiv — Machine Learning research 1mo ago Learning-enabled Parameter Synthesis for Nonlinear Systems from Signal Temporal Logic arXiv:2607.08899v1 Announce Type: cross Abstract: Signal Temporal Logic (STL) is increasingly used to describe interpretable objectives and constraints for optimal control and learning methods, especially when no target time series data is available. In this work, we propose to… 17 Page 9 of 10 · 500 articles ← Newer Older →