News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow Hugging Face Daily Papers research 7d ago OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Abstract Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning.… 25 arXiv — Machine Learning research 7d ago Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language arXiv:2608.05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap. The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is… 31 arXiv — Machine Learning research 7d ago A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric… 33 arXiv — Machine Learning research 7d ago Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping arXiv:2608.06105v1 Announce Type: new Abstract: Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments.… 28 arXiv — Machine Learning research 7d ago Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction arXiv:2608.06262v1 Announce Type: new Abstract: Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. We… 17 arXiv — NLP / Computation & Language research 7d ago PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing… 24 arXiv — NLP / Computation & Language research 7d ago Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles… 38 arXiv — Machine Learning research 7d ago A Unified Causal Inference Framework for the Desirability of Outcome Ranking Paradigm in Benefit-Risk Evaluation arXiv:2608.05244v1 Announce Type: cross Abstract: We developed a unified covariate-adjusted causal inference framework for estimating the desirability of outcome ranking (DOOR) probability for benefit-risk evaluation in randomized trials and observational studies. The framework… 30 arXiv — Machine Learning research 7d ago Physics-Based Molecular Fingerprints from Spectral Graph Theory Provide Efficient Geometry-Aware Measures of Chemical Similarity arXiv:2608.05336v1 Announce Type: cross Abstract: Molecular representations are essential for the evaluation of molecular similarity and the development of structure-property relationships. Despite the known importance of 3D structure to determine chemical and physical… 10 arXiv — NLP / Computation & Language research 7d ago Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation arXiv:2608.05155v1 Announce Type: new Abstract: Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideological, and framing dimensions of political discourse -- dimensions that are central to… 9 arXiv — NLP / Computation & Language research 7d ago Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation arXiv:2608.05353v1 Announce Type: new Abstract: LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information needed for a later verdict. For reasoning-capable… 19 arXiv — NLP / Computation & Language research 7d ago Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation arXiv:2608.05726v1 Announce Type: new Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to… 38 arXiv — NLP / Computation & Language research 7d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark arXiv:2608.05850v1 Announce Type: new Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have… 23 arXiv — NLP / Computation & Language research 7d ago A Study of LLMs' Preferences for Libraries and Programming Languages arXiv:2503.17181v4 Announce Type: cross Abstract: Despite the rapid progress of large language models (LLMs) in code generation, existing evaluations focus on functional correctness or syntactic validity, overlooking how LLMs make critical design choices such as which library or… 34 arXiv — NLP / Computation & Language research 7d ago TriQua: Reconciling Granularity and Context in Factuality Evaluation arXiv:2608.05228v1 Announce Type: cross Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the… 14 arXiv — NLP / Computation & Language research 7d ago From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a… 33 arXiv — NLP / Computation & Language research 7d ago Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI arXiv:2608.06167v1 Announce Type: cross Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold… 31 arXiv — NLP / Computation & Language research 7d ago AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either… 8 arXiv — NLP / Computation & Language research 7d ago Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges arXiv:2405.15604v4 Announce Type: replace Abstract: Text generation has become more accessible than ever, and the growing interest in these systems, especially those using large language models, has spurred a surge in related publications. We provide a systematic literature… 17 Hugging Face Daily Papers research 7d ago What AI Red-Team Evaluations Can and Cannot Prove Abstract Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief… 11 Simon Willison community 7d ago Simon Willison on Technical Blogging Simon Willison on Technical Blogging I was interviewed by Cynthia Dunlop for her "Write that blog!" series back in January, but I just realized I never linked to the interview from my own blog! It includes my answers to the following questions: Why did you start blogging – and… 29 Simon Willison community 7d ago Simon Willison on Technical Blogging Simon Willison on Technical Blogging I was interviewed by Cynthia Dunlop for her "Write that blog!" series back in January, but I just realized I never linked to the interview from my own blog! It includes my answers to the following questions: Why did you start blogging – and… 28 Don't Worry About the Vase community 7d ago AI #180: No Longer In Charge What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. 23 TechCrunch — AI news-outlet 7d ago Omilia raises $67M to scale its customer support platform The Series B is the company's second fundraise since it last raised capital in 2020. In that time, it has increased its ARR by 10x to $60 million. 20 Hugging Face Daily Papers research 8d ago AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Abstract While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on… 15 arXiv — Machine Learning research 8d ago CAMP: A Cycle-Aware Multi-Scale Patch Mixer for Time Series Forecasting arXiv:2608.04051v1 Announce Type: new Abstract: Real-world time series are often governed by recurring patterns, but their dominant periods may vary across datasets, forecasting settings, and individual input windows. Existing cycle-aware forecasters commonly rely on a single… 35 arXiv — Machine Learning research 8d ago MINT: Tensor Decomposition on Stacked Recurrence Matrices for Time Series Data Mining arXiv:2608.04157v1 Announce Type: new Abstract: Recurrence plots are a time series data mining primitive applied to a variety of domains (e.g. star light curves, sound waveforms, CCT telemetry). This work proposes tensorized self-similarity matrices as a primitive for univariate… 6 arXiv — Machine Learning research 8d ago TS2TabPFN: Time Series Classification and Extrinsic Regression through Feature Extraction and a Tabular Foundation Model arXiv:2608.04174v1 Announce Type: new Abstract: Time series data are ubiquitous in practical applications, where classification (TSC) and extrinsic regression (TSER) have emerged as essential tasks for obtaining value from temporal sequences. While the literature has seen… 6 arXiv — Machine Learning research 8d ago Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation arXiv:2608.04333v1 Announce Type: new Abstract: Large language model (LLM) configuration evaluation is challenging due to limited evaluation budgets, varying costs, and multiple competing objectives. In this paper, we formulate LLM configuration evaluation as a cost-aware… 7 arXiv — Machine Learning research 8d ago EvtGraph: Event-Adaptive Compression for Sparse Temporal Graph Learning in Multimodal Time Series arXiv:2608.04368v1 Announce Type: new Abstract: Multimodal temporal data are inherently irregular and uneven in information density, yet most models rely on uniform discretization, leading to inefficient representations. We propose \textbf{EvtGraph}, a unified framework that… 37 arXiv — Machine Learning research 8d ago Active Learning Guided Design Space Refinement for Scalable Multi-Objective Bayesian Optimization in Materials Discovery arXiv:2608.04651v1 Announce Type: new Abstract: Advanced materials discovery increasingly relies on machine learning and Bayesian optimization to explore large discrete design spaces under limited evaluation budgets. However, conventional Bayesian optimization (BO) can become… 6 arXiv — Machine Learning research 8d ago Benchmarking Deep Learning Models for Dense Event Classification of Offshore Wind Infrastructure in Sentinel-1 Time Series arXiv:2608.04706v1 Announce Type: new Abstract: Monitoring of offshore wind energy infrastructure life cycles, especially during the deployment phase, is an important contribution for stakeholders to make informed decisions in a phase of increasing deployment activities. ESA's… 31 arXiv — Machine Learning research 8d ago MGSB: Manifold Gated Signature Branch Pressure-Domain Baseline Architecture for Two-Phase Pipeline Flows Under Distributional Shift arXiv:2608.04805v1 Announce Type: new Abstract: Leak detection models for multiphase pipelines often degrade when deployed under flow regimes that differ from training. Existing evaluations typically assess performance under in-distribution operating conditions, masking failures… 15 arXiv — NLP / Computation & Language research 8d ago Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap arXiv:2608.04160v1 Announce Type: new Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. We test whether the… 34 arXiv — NLP / Computation & Language research 8d ago Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary arXiv:2608.04240v1 Announce Type: new Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to… 17 arXiv — NLP / Computation & Language research 8d ago Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation arXiv:2608.04260v1 Announce Type: new Abstract: Metaphorical language remains a major challenge for multilingual natural language processing because successful interpretation and translation require reasoning beyond literal lexical meaning. Existing research has largely… 14 arXiv — NLP / Computation & Language research 8d ago STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled… 21 arXiv — NLP / Computation & Language research 8d ago Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its… 5 arXiv — NLP / Computation & Language research 8d ago Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv:2608.04899v1 Announce Type: new Abstract: Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extremely sparse outputs. For instance, Qwen3-32B… 27 arXiv — NLP / Computation & Language research 8d ago Simile Understanding in Text-to-Image Models: An Evaluation Framework arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models… 12 Simon Willison community 8d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 22 Simon Willison community 8d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 19 Simon Willison community 8d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 37 Simon Willison community 8d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 14 TechCrunch — AI news-outlet 8d ago AI makes weather prediction better. Can WindBorne make it lucrative? WindBorne Systems has raised $37 million Series B round to scale its weather balloons and AI forecasts. 9 Hugging Face Daily Papers research 9d ago Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Abstract We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in… 8 arXiv — Machine Learning research 9d ago Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment arXiv:2608.02786v1 Announce Type: new Abstract: AI systems can fail silently. The failure propagates through training loops, evaluation pipelines, and production monitoring stacks until downstream harm makes it visible. This paper introduces evaluation blindness: a measurement… 30 arXiv — Machine Learning research 9d ago Forecasting Revenue with its Customer-Base Drivers: When and Why Coordination Helps arXiv:2608.02911v1 Announce Type: new Abstract: Revenue forecasts guide acquisition budgets, demand planning, and customer-based valuations, yet an aggregate forecast does not show whether change reflects acquisition, repeat purchasing, spending per order, or offsetting… 8 arXiv — Machine Learning research 9d ago Paired Recipient-based Evaluation of Survival Prediction for Deceased Donor Kidney Transplants arXiv:2608.03017v1 Announce Type: new Abstract: There has been significant interest in using machine learning algorithms to predict kidney transplant outcomes, such as the number of years until a graft inevitably fails. These prediction algorithms could possibly be used for… 6 arXiv — Machine Learning research 9d ago FinVerse: Financial Time-Series Benchmark arXiv:2608.03259v1 Announce Type: new Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. Existing time-series forecasting benchmarks provide useful… 14 Page 3 of 10 · 500 articles ← Newer Older →