News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — Machine Learning research 17d ago An Empirical Measurement of Jailbreaking Evaluators arXiv:2609.10594v1 Announce Type: cross Abstract: Expert evaluation of jailbreak responses is costly and difficult to scale, so the community increasingly relies on automated evaluators to determine whether an attack succeeds. However, jailbreak studies typically validate their… 38 arXiv — Machine Learning research 17d ago Sequence-Informed Geometric Evaluation of RNA 3D Structures arXiv:2609.10644v1 Announce Type: cross Abstract: Computational RNA structure pipelines generate many candidate conformations for the same sequence. Reliable evaluation therefore requires more than recognising plausible geometry, it requires determining whether that geometry is… 37 arXiv — NLP / Computation & Language research 17d ago Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting arXiv:2609.11131v1 Announce Type: new Abstract: Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is designed for rubric-aligned segment-level SI… 5 arXiv — NLP / Computation & Language research 17d ago Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech arXiv:2609.11545v1 Announce Type: new Abstract: Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker… 37 arXiv — NLP / Computation & Language research 17d ago RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety arXiv:2609.11758v1 Announce Type: new Abstract: Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can… 25 arXiv — NLP / Computation & Language research 17d ago The widening evaluation gap in medical large language model research 2023 to 2026 arXiv:2609.11770v1 Announce Type: new Abstract: Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026… 13 arXiv — NLP / Computation & Language research 17d ago Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech arXiv:2609.11786v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains… 33 arXiv — NLP / Computation & Language research 17d ago Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2 arXiv:2609.10935v1 Announce Type: cross Abstract: Membership inference attacks (MIAs) try to determine whether a specific record was used to train a model, a privacy risk that matters in natural language processing (NLP), where training data can contain sensitive user text. This… 19 arXiv — NLP / Computation & Language research 17d ago Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment arXiv:2609.11144v1 Announce Type: cross Abstract: Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where… 6 arXiv — NLP / Computation & Language research 17d ago "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated arXiv:2508.05830v3 Announce Type: replace Abstract: Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror"… 8 Simon Willison community 17d ago Datasette 1.0a39 and 0.65.4 security releases Datasette 1.0a39 and 0.65.4 security releases Today we're releasing two new security patch versions of Datasette: 1.0a39 and 0.65.4 - one for the current alpha series and one for the stable 0.65.x family. These are security fixes which you should apply if you are running a… 36 The Information — AI news-outlet 17d ago Bending Spoons Buys Miro For $1.355 Billion Italian digital conglomerate Bending Spoons is buying whiteboarding software firm Miro at a valuation of $1.355 billion, the latest in a series of purchases the Italian firm has done at knock-down prices. The price for Miro is a huge discount to its peak valuation of $17.5… 16 TechCrunch — AI news-outlet 17d ago Maven Robotics wants to steal your robot deployment deal Maven Robotics emerged from stealth today with a $100 million Series A and active deployments. 29 Hugging Face Daily Papers research 18d ago Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs Abstract This survey examines inference-efficiency techniques for video large language models, analyzing cost reductions across frame sampling, encoding, token compression, and language model stages while identifying evaluation gaps. Generated by thinkingmachines/Inkling-Small… 25 Hugging Face Daily Papers research 18d ago SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Abstract SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of… 20 arXiv — Machine Learning research 18d ago AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning arXiv:2609.05435v1 Announce Type: new Abstract: Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt… 19 arXiv — Machine Learning research 18d ago Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations arXiv:2609.05658v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy… 35 arXiv — Machine Learning research 18d ago Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems arXiv:2609.05933v1 Announce Type: new Abstract: Efficiency is increasingly important for Large Language Model (LLM)-based multi-agent systems (MAS), as larger models and more agents introduce substantial execution costs. Recent methods aim to make MAS cheaper by pruning agents,… 38 arXiv — Machine Learning research 18d ago DataFlex-RL: An Evaluation Platform for RLVR Data Policies arXiv:2609.06107v1 Announce Type: new Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an… 19 arXiv — Machine Learning research 18d ago Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction arXiv:2609.06367v1 Announce Type: new Abstract: LLM-as-a-Judge has emerged as a promising paradigm for evaluating natural language generation. However, the uncertainty associated with such evaluations remains largely unexplored, which limits their reliability in real-world… 30 arXiv — NLP / Computation & Language research 18d ago SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia arXiv:2609.09672v1 Announce Type: new Abstract: The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA)… 22 arXiv — NLP / Computation & Language research 18d ago CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription arXiv:2609.09766v1 Announce Type: new Abstract: Churn models typically identify high-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate. We present CARRE (Counterfactual Action Retrieval and Reason… 33 arXiv — NLP / Computation & Language research 18d ago YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models arXiv:2609.10153v1 Announce Type: new Abstract: Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test… 14 arXiv — NLP / Computation & Language research 18d ago DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs arXiv:2609.10253v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation,… 33 arXiv — NLP / Computation & Language research 18d ago From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges arXiv:2601.08654v3 Announce Type: replace Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the same criteria inconsistently, produce score attributions that are difficult to… 14 r/MachineLearning community 18d ago I trained a 348M model trained from scratch on 22.7B tokens that does 14 digit arithmetic [P] Hello this is my fifth small language model I've made and apart of my third series and it has been a lot of work but it payed off: **348M parameters, 22.7B tokens**, then fine-tuned into a math model that solves arithmetic by *showing the work* — column addition with carries,… 37 TechCrunch — AI news-outlet 18d ago AI research startup Listen Labs scrubbed a $1.5B funding round for Salesforce talks Listen Labs walked away from a signed Series C term sheet from Menlo Ventures, sources say. 22 TechCrunch — AI news-outlet 18d ago Sequoia doubles down on Cymphony as AI agents create new enterprise security risks Cymphony was valued at more than $100 million in a $25 million Series A co-led by Sequoia and SMBC Fin Atlas Beyond Fund. 28 Latent.Space news-outlet 19d ago [AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded Overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5. The most jam packed, feel the AGI day in the history of AI. 9 The Information — AI news-outlet 19d ago Gimlet's Three-Tranche Deal Reveals AI Funding Frenzy Gimlet Labs said late last week it had raised $300 million at a valuation of about $3 billion, in a round led by new investor Andreessen Horowitz. Today we have unreported details on how that round, which boosted its valuation by a whopping 16 times from its last PitchBook… 20 The Information — AI news-outlet 19d ago Gimlet's Three-Tranche Deal Reveals AI Funding Frenzy Gimlet Labs said late last week it had raised $300 million at a valuation of about $3 billion, in a round led by new investor Andreessen Horowitz. Today we have unreported details on how that round, which boosted its valuation by a whopping 16 times from its last PitchBook… 19 TechCrunch — AI news-outlet 19d ago Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market Cognition's valuation multiple is higher than Cursor's was before selling to SpaceX. 5 The Information — AI news-outlet 19d ago Cognition Raised Over $2 Billion at a $48 Billion Valuation Cognition, the artificial intelligence startup behind the Devin coding assistant, has raised more than $2 billion at a $48 billion valuation including the new funding, nearly double its valuation in May, the company said Tuesday. The round shows coding agent startups are… 37 TechCrunch — AI news-outlet 19d ago Mistral raises €3B as sovereign AI becomes big business The French AI lab has raised €3 billion at a €21 billion valuation in a Series D round led by Samsung, Scaleup Europe and PSG Equity. 23 Smol AI News news-outlet 20d ago OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5 **OpenAI** announced a proposed Navier–Stokes proof by an internal model "**significantly more capable than GPT-6 Astra**" using **10,000 agents** over **88 hours** plus **17 hours** of formal verification. The effort highlights the emergence of **massive test-time compute… 19 Hugging Face Daily Papers research 21d ago Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems Abstract The study formalizes multi-agent LLM coordination via bilevel games and stochastic memory reflection, introducing a grounded evaluation gate and SRMA algorithm with convergence guarantees, validated on SWE-bench. Generated by thinkingmachines/Inkling-Small Multi-agent… 4 r/LocalLLaMA community 21d ago Benchmarking calories evaluation with LLMs I wanted a quick calories counter for myself, using LLMs to evaluate the calories from pictures of meals + descriptions. I needed to pick a model so I made a quick benchmark. The setup was: - Nutrition5k photos for photo + calories:… 29 arXiv — Machine Learning research 21d ago BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation arXiv:2609.04292v1 Announce Type: new Abstract: Human mobility predictability concerns the best prediction performance attainable from a given target and input information, but its ground truth is not directly observable on real mobility data. We present BER-PEF, a… 22 arXiv — Machine Learning research 21d ago Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures arXiv:2609.04407v1 Announce Type: new Abstract: Deep neural operators learn mappings between input functions and complete PDE solution fields, enabling forward evaluations of new problem instances orders of magnitude faster than conventional numerical solvers. Attention… 18 arXiv — Machine Learning research 21d ago MomentQuant: an even more minimalist interval method with linear time complexity for time series classification arXiv:2609.05136v1 Announce Type: new Abstract: Time series data is very common in many real-world applications and in numerous domains, with increasing interest for automated information extraction using machine learning. One of these subfields is time series classification,… 21 arXiv — Machine Learning research 21d ago Interface-Induced Trajectory Censoring arXiv:2609.03966v1 Announce Type: cross Abstract: Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's… 38 arXiv — Machine Learning research 21d ago Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys arXiv:2609.04382v1 Announce Type: cross Abstract: We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested. The Trusted Local Node (TLN) sends protected activations to the… 19 arXiv — NLP / Computation & Language research 21d ago You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments arXiv:2609.04384v1 Announce Type: new Abstract: Chinese online comments often convey social meaning through indirect and playful language that is hard to interpret without context. Existing evaluations largely organize items around predefined phenomena or controlled pragmatic… 12 arXiv — NLP / Computation & Language research 21d ago Evaluation of Phonetic Encoding Algorithms on Transcription Datasets arXiv:2609.04391v1 Announce Type: new Abstract: In this work, a novel evaluation scheme built on a generalized variant of the Rand Index measure, namely, the H\"ullermeier-Rifqi Index, is proposed in order to assess how well phonetic encoding algorithms conform to word-based… 28 arXiv — NLP / Computation & Language research 21d ago A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models arXiv:2609.04409v1 Announce Type: new Abstract: Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated… 6 arXiv — NLP / Computation & Language research 21d ago On Epistemic Diversity in Large Language Models arXiv:2609.04835v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A… 17 arXiv — NLP / Computation & Language research 21d ago MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain arXiv:2609.04842v1 Announce Type: new Abstract: Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of… 27 arXiv — NLP / Computation & Language research 21d ago RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents arXiv:2609.04898v1 Announce Type: new Abstract: Repository-scale refactoring requires coding agents to propagate a single change across many interdependent files without altering program behavior, yet to our knowledge no existing harness isolates the design choices that… 12 arXiv — NLP / Computation & Language research 21d ago MoirfEolas and Cr\'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology arXiv:2609.05022v1 Announce Type: new Abstract: This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphological boundaries of the language. We present MoirfEolas, a dataset of over 35,000 Irish words mapped to their respective… 17 arXiv — NLP / Computation & Language research 21d ago Auditing Bias and Safety in Voice AI Customer Care arXiv:2609.04206v1 Announce Type: cross Abstract: Voice AI systems increasingly mediate customer care interactions where caller presentation cues such as accent, affect, fluency, and urgency are available alongside the service request. Existing fairness and safety evaluations… 15 Page 5 of 10 · 500 articles ← Newer Older →