News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 2d ago No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to… 37 Hugging Face Daily Papers research 2d ago TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity Abstract A unified toolbox enables reproducible comparison and extension of time-series dataset similarity methods for forecasting, classification, and generation tasks. Generated by thinkingmachines/Inkling-Small The rapid advancement of artificial intelligence (AI) has… 10 Hugging Face Daily Papers research 2d ago Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness Abstract Researchers propose source-contrastive evaluation via a localized benchmark to detect data contamination and assess localization robustness in multilingual translation models. Generated by thinkingmachines/Inkling-Small Multilingual translation benchmarks are typically… 4 Hugging Face Daily Papers research 2d ago Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Abstract Optimized GPU kernel benchmarks reveal that evolutionary LLM proposals exploit evaluation configurations, causing widespread failure to generalize to held-out settings. Generated by thinkingmachines/Inkling-Small Benchmarks for systems that are optimized against the… 11 Hugging Face Daily Papers research 2d ago MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Abstract MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities. Generated by… 37 Hugging Face Daily Papers research 3d ago A^2E : An End-to-End Agent Auditing Engine Abstract A2E is an end-to-end evaluation engine for agent harnesses that uses a standardized task protocol and execution traces to assess capabilities across efficiency, tool use, planning, and error recovery. Generated by thinkingmachines/Inkling-Small With the rapid… 9 arXiv — Machine Learning research 3d ago Neural Operators for Immersed-Boundary Soft Swimmers Locomotion arXiv:2608.07722v1 Announce Type: new Abstract: High-fidelity immersed-boundary simulation resolves the coupled motion of a deforming swimmer and its surrounding flow, but the resulting cost limits repeated evaluations for engineering design, parameter studies, and control. We… 16 arXiv — Machine Learning research 3d ago When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes arXiv:2608.07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache management an attractive lever: a policy that raised the hit rate would cut… 28 arXiv — Machine Learning research 3d ago TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity arXiv:2608.08119v1 Announce Type: new Abstract: The rapid advancement of artificial intelligence (AI) has significantly accelerated research in time-series analysis, particularly in forecasting, classification, and generation tasks. Recent models, especially foundation models,… 38 arXiv — Machine Learning research 3d ago FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification arXiv:2608.08207v1 Announce Type: new Abstract: Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales while remaining computationally efficient. However, existing approaches generally struggle… 24 arXiv — Machine Learning research 3d ago The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World arXiv:2608.08239v1 Announce Type: new Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged… 5 arXiv — Machine Learning research 3d ago Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure arXiv:2608.08722v1 Announce Type: new Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU-kernel-optimization suites with held-out generalization gates:… 24 arXiv — Machine Learning research 3d ago Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data arXiv:2608.08859v1 Announce Type: new Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under strict computational and memory constraints. A critical yet underexplored challenge… 19 arXiv — NLP / Computation & Language research 3d ago Unified Hallucination Fuzzing for Multimodal Large Language Models arXiv:2608.07525v1 Announce Type: new Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from… 16 arXiv — NLP / Computation & Language research 3d ago SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly… 36 arXiv — NLP / Computation & Language research 3d ago Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation arXiv:2608.07763v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which… 19 arXiv — NLP / Computation & Language research 3d ago On the use of foundation models in cognitive science arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive… 34 arXiv — NLP / Computation & Language research 3d ago SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.… 21 arXiv — NLP / Computation & Language research 3d ago Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions arXiv:2608.07968v1 Announce Type: new Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency… 33 arXiv — NLP / Computation & Language research 3d ago A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization arXiv:2608.08180v1 Announce Type: new Abstract: Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities and events. Such relation-level hallucinations undermine the reliability of… 28 arXiv — NLP / Computation & Language research 3d ago Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations? arXiv:2608.08283v1 Announce Type: new Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic… 30 arXiv — NLP / Computation & Language research 3d ago Position Bias in Ordinal Classification: A Systematic Evaluation arXiv:2608.08869v1 Announce Type: new Abstract: Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from… 36 arXiv — NLP / Computation & Language research 3d ago How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review arXiv:2608.08975v1 Announce Type: new Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how… 24 arXiv — NLP / Computation & Language research 3d ago ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prioritize limited clinical resources. At presentation, however, information may be… 5 arXiv — NLP / Computation & Language research 3d ago Evo-Bench: Can Language Models Improve Agent Harness? arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously… 27 arXiv — NLP / Computation & Language research 3d ago EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models arXiv:2608.09189v1 Announce Type: new Abstract: Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a… 7 arXiv — NLP / Computation & Language research 3d ago Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing arXiv:2608.09289v1 Announce Type: new Abstract: Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipeline that… 18 arXiv — NLP / Computation & Language research 3d ago Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts arXiv:2608.09510v1 Announce Type: new Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring… 38 arXiv — NLP / Computation & Language research 3d ago How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans arXiv:2608.09717v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction… 16 arXiv — NLP / Computation & Language research 3d ago Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness arXiv:2608.09766v1 Announce Type: new Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale… 30 TechCrunch — AI news-outlet 3d ago Discovered Materials is playing AI whack-a-mole to hunt cooler chips Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips. 36 Hugging Face Daily Papers research 4d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination Abstract Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability,… 14 arXiv — Machine Learning research 4d ago Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning arXiv:2608.06511v1 Announce Type: new Abstract: Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly… 35 arXiv — Machine Learning research 4d ago When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series arXiv:2608.07333v1 Announce Type: new Abstract: Modeling multivariate time series by representing them as graphs, where individual series act as nodes and pairwise temporal corre- lations serve as edges, has gained significant traction. Recent advances in Graph Neural Networks… 6 arXiv — NLP / Computation & Language research 4d ago The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done,… 9 arXiv — NLP / Computation & Language research 4d ago Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response… 13 arXiv — NLP / Computation & Language research 4d ago Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar… 36 arXiv — NLP / Computation & Language research 4d ago From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL arXiv:2608.07213v1 Announce Type: new Abstract: Test-time scaling can correct difficult text-to-SQL queries, but the extra computation is normally discarded after each answer. Systems increasingly retain verified repair episodes, yet evaluations still report one end-to-end… 18 arXiv — NLP / Computation & Language research 4d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination arXiv:2608.07341v1 Announce Type: new Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and… 7 arXiv — NLP / Computation & Language research 4d ago An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis arXiv:2608.07439v1 Announce Type: new Abstract: Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat) is one of its theoretically grounded formulations. Prior work on financial… 29 arXiv — NLP / Computation & Language research 4d ago ADIAS: Automated Design of Interactive Agentic Systems arXiv:2608.06410v1 Announce Type: cross Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents,… 36 arXiv — NLP / Computation & Language research 4d ago How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social… 14 arXiv — NLP / Computation & Language research 4d ago Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models arXiv:2608.07243v1 Announce Type: cross Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM… 20 arXiv — NLP / Computation & Language research 4d ago How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack… 17 Hugging Face Daily Papers research 4d ago StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Abstract Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design… 27 r/LocalLLaMA community 5d ago DeepSeek V4 Flash 0731 hits 82.7% on Terminal-Bench 2.1 in an independent public-harness run (445 trials) Disclosure: I’m the author of Ante. DeepSeek recently reported an 82.7% score on Terminal-Bench 2.1 for DeepSeek V4 Flash 0731. Its evaluation used “DeepSeek Harness minimal mode,” which hasn’t been released yet. We wanted to see whether the reported result could be… 14 r/MachineLearning community 5d ago Evaluation metrics - [D] I wanted to ask about the selection of evaluation metrics. In which scenario we use ROC-AUC score and in which scenario we use f1 score as an evaluation metric in a classification problem to define the model's performance on a specific dataset.   submitted by  … 8 r/LocalLLaMA community 5d ago Repeated generation is worth it and self-evaluation is effective I made gemma4 12B write timestamp-anchored summaries of youtube video transcripts. I tested if the summaries have significant qualitative variance and if the SLM can pick the best one by itself. Below is the prompt texts I used. "{{[INPUT]}}<attachement name='original'>",… 18 OpenAI official-blog 6d ago Responding to the next frontier of critical cyber capabilities OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls. 22 Hugging Face Daily Papers research 6d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Abstract We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish… 5 Page 2 of 10 · 500 articles ← Newer Older →