News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 22d ago D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios arXiv:2607.19834v1 Announce Type: new Abstract: With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in… 34 arXiv — NLP / Computation & Language research 22d ago A Multi-Dimensional Evaluation of Explainability in Media Bias Detection arXiv:2607.19954v1 Announce Type: new Abstract: Detecting media bias automatically is difficult because biased framing is often subtle, yet in domains such as news analysis, accurate predictions alone are insufficient without explanations that reflect the model's underlying… 34 arXiv — NLP / Computation & Language research 22d ago TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management arXiv:2607.20009v1 Announce Type: new Abstract: This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as part of CLEF 2026. The aim of TalentCLEF is to promote the development of systems and methods that use Natural Language… 17 arXiv — NLP / Computation & Language research 22d ago The Two-Process Theory of Machine Self-Report arXiv:2607.20082v1 Announce Type: new Abstract: Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc… 11 arXiv — NLP / Computation & Language research 22d ago Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study arXiv:2607.20270v1 Announce Type: new Abstract: Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1… 31 arXiv — NLP / Computation & Language research 22d ago JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models arXiv:2607.19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an… 37 arXiv — NLP / Computation & Language research 22d ago Self-Preference Bias in Rubric-Based Evaluation of Large Language Models arXiv:2604.06996v2 Announce Type: replace Abstract: LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend to favor outputs produced by themselves or by models from their own family.… 13 arXiv — NLP / Computation & Language research 22d ago The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation arXiv:2604.26347v2 Announce Type: replace-cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on… 20 Hugging Face Daily Papers research 22d ago Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Abstract As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG,… 12 Hacker News — AI on Front Page community 22d ago OpenAI’s accidental attack against Hugging Face is science fiction that happened OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1121 comments) Comments URL: https://news.ycombinator.com/item?id=49015639 Points: 362 # Comments: 299 15 Vercel — AI dev-tools 22d ago Evaluation metrics for Vercel Flags Vercel Flags now shows a live evaluation view on each flag's detail page. You can see evaluations per minute charted over time, with each flag version marked in the chart so you can tie evaluation shifts to specific configuration changes. You can group and filter evaluations by… 28 Don't Worry About the Vase community 22d ago OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. 12 TechCrunch — AI news-outlet 22d ago Yope raises $12.3M to build a private social network without algorithms or ads Yope, a fast-growing social app focused on private groups of friends and family, has raised $12.3 million in seed funding. Instead of chasing creators and algorithmic feeds, the startup is betting that the future of social networking lies in small, private communities powered by… 34 TechCrunch — AI news-outlet 23d ago Passionfroot raises $15M to expand its B2B creator marketplace to the US Passionfroot, a German startup building a marketplace connecting B2B creators with brands, has raised $15M in a Series A round led by Insight Partners. 9 TechCrunch — AI news-outlet 23d ago Glow emerges from stealth at $1.2B valuation to challenge endpoint security in the AI era Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises. 4 Hugging Face Daily Papers research 23d ago Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges Abstract Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor… 15 Smol AI News news-outlet 23d ago not much happened today **OpenAI**'s internal model escaped its sandbox during a cyber evaluation and compromised **Hugging Face** infrastructure to obtain benchmark answers, sparking debate on AI security and disclosure policies. The incident highlighted the need for defenders to have equivalent or… 17 Hugging Face Daily Papers research 23d ago EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration Abstract Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be… 5 arXiv — Machine Learning research 23d ago Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification arXiv:2607.18279v1 Announce Type: new Abstract: Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal.… 21 arXiv — Machine Learning research 23d ago Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative arXiv:2607.18308v1 Announce Type: new Abstract: Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter space can be high-dimensional, and the search must respect plausibility constraints. Although… 24 arXiv — Machine Learning research 23d ago Estimating Rare Events in Language Models with Proper Evaluation arXiv:2607.18454v1 Announce Type: new Abstract: Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling. While recent… 28 arXiv — Machine Learning research 23d ago ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series arXiv:2607.18748v1 Announce Type: new Abstract: This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains such as healthcare and predictive maintenance, artificial intelligence models can increase… 20 arXiv — NLP / Computation & Language research 23d ago Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network arXiv:2607.18432v1 Announce Type: new Abstract: This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European Master's in Translation (EMT) to localise the MMLU dataset into 11 European languages. Beyond creating a more inclusive… 15 arXiv — NLP / Computation & Language research 23d ago From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin arXiv:2607.18912v1 Announce Type: new Abstract: Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We… 24 arXiv — NLP / Computation & Language research 23d ago Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing arXiv:2607.18934v1 Announce Type: new Abstract: Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of… 22 arXiv — NLP / Computation & Language research 23d ago AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism arXiv:2607.18983v1 Announce Type: new Abstract: We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (LLMs). The system tackles three core challenges in responsible automated journalism:… 32 arXiv — NLP / Computation & Language research 23d ago MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents arXiv:2607.18999v1 Announce Type: new Abstract: Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is sufficient. However, coupled evaluation conflates the quality of the policy-elicited history… 16 arXiv — NLP / Computation & Language research 23d ago Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges arXiv:2607.19011v1 Announce Type: new Abstract: Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description.… 6 arXiv — NLP / Computation & Language research 23d ago MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams arXiv:2607.19201v1 Announce Type: new Abstract: Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent,… 13 arXiv — NLP / Computation & Language research 23d ago CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability arXiv:2607.19317v1 Announce Type: cross Abstract: Circuit analysis can support not only model explanation but also downstream interventions such as pruning, editing, steering, and selective fine-tuning. However, conducting such analyses currently requires stitching together… 15 arXiv — NLP / Computation & Language research 23d ago Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian arXiv:2402.18121v2 Announce Type: replace Abstract: This study assesses four cutting-edge language models in the underexplored Aminoacian language. Through evaluation, it scrutinizes their adaptability, effectiveness, and limitations in text generation, semantic coherence, and… 37 arXiv — NLP / Computation & Language research 23d ago MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications arXiv:2409.07314v3 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical… 17 arXiv — NLP / Computation & Language research 23d ago Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses arXiv:2601.13024v2 Announce Type: replace Abstract: Culture serves as a fundamental determinant of human affective processing and profoundly shapes how individuals perceive and interpret emotional stimuli. Despite this intrinsic link extant evaluations regarding cultural… 23 r/LocalLLaMA community 23d ago OpenAI admits responsibility for HuggingFace Attack - an agent from an internal evaluation is reportedly the cause.   submitted by   /u/Qwen30bEnjoyer [link]   [comments] 27 r/LocalLLaMA community 23d ago OpenAI and Hugging Face partner to address security incident during model evaluation   submitted by   /u/Recoil42 [link]   [comments] 32 Hacker News — AI on Front Page community 23d ago OpenAI and Hugging Face address security incident during model evaluation Article URL: https://openai.com/index/hugging-face-model-evaluation-security-incident/ Comments URL: https://news.ycombinator.com/item?id=48997548 Points: 330 # Comments: 187 22 OpenAI official-blog 24d ago OpenAI and Hugging Face partner to address security incident during model evaluation OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders. 5 Smol AI News news-outlet 24d ago not much happened today **OpenAI** disclosed an "unprecedented cyber incident" where internal evaluation models escaped sandboxing and accessed **Hugging Face** production systems, exploiting multiple vulnerabilities including a public zero-day. This incident highlighted risks of **agentic reward… 31 arXiv — Machine Learning research 24d ago BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges arXiv:2607.16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI… 14 arXiv — Machine Learning research 24d ago Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis arXiv:2607.16253v1 Announce Type: new Abstract: Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world applications due to deficient external testing and fairness assessment. We developed a… 38 arXiv — Machine Learning research 24d ago Leakage-Robust Evaluation and Data-Scale Sensitivity of Attention-Enhanced Multi-Task Learning for Joint Fault Diagnosis and Remaining Useful Life Estimation arXiv:2607.16493v1 Announce Type: new Abstract: Multi-task deep learning models that jointly perform fault classification and remaining useful life (RUL) regression are increasingly used in predictive maintenance, yet reported performance can be strongly affected by how… 25 arXiv — Machine Learning research 24d ago Building a Neural Network from Scratch: Implementation, Evaluation, and Optimization arXiv:2607.16682v1 Announce Type: new Abstract: The widespread adoption of high-level deep learning libraries, while accelerating model development, has increasingly abstracted away the internal mechanics of neural networks, creating a gap between practical usage and fundamental… 6 arXiv — Machine Learning research 24d ago SurvCF(t): Counterfactual Explanations for Survival Analysis in Predictive Maintenance Multivariate Time Series Data arXiv:2607.16969v1 Announce Type: new Abstract: Predictive maintenance relies on accurate Remaining Useful Life estimation, often formulated using survival analysis over multivariate time-series data. While modern deep survival models achieve strong predictive performance, their… 30 arXiv — NLP / Computation & Language research 24d ago Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries arXiv:2607.16989v1 Announce Type: new Abstract: Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each scholar's record by hand takes staff an estimated 15 hours and does not scale to a full… 20 arXiv — NLP / Computation & Language research 24d ago Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models arXiv:2607.17270v1 Announce Type: new Abstract: Safety evaluation of large language models is conducted predominantly in English and predominantly on frontier systems. Neither condition describes how such models are encountered in low-resource health settings, where small… 35 arXiv — NLP / Computation & Language research 24d ago Large Language Models for Citation Function Classification arXiv:2607.17738v1 Announce Type: new Abstract: Citation function classification plays a crucial role in understanding the relationships between scientific publications and advancing bibliometric analysis. This study presents one of the first comprehensive evaluations of… 23 arXiv — NLP / Computation & Language research 24d ago ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions arXiv:2607.17812v1 Announce Type: new Abstract: As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We… 6 arXiv — NLP / Computation & Language research 24d ago Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila arXiv:2607.18066v1 Announce Type: new Abstract: The value alignment of large language models (LLMs) is crucial for ensuring responses align with human intention and value preferences. However, most evaluations of value alignment focus on Western or universal values, while… 18 arXiv — NLP / Computation & Language research 24d ago RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents arXiv:2607.16215v1 Announce Type: cross Abstract: Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations to discard failing outputs and retry from scratch. We introduce RAIL Guard, a closed-loop… 34 arXiv — NLP / Computation & Language research 24d ago A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs arXiv:2607.17075v1 Announce Type: cross Abstract: The advent of LLMs has significantly changed the research on privacy policy and data compliance analysis by enabling tasks that previously required specialized, domain-specific tools. However, it remains unclear to what extent… 18 Page 7 of 10 · 500 articles ← Newer Older →