News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 12d ago LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture arXiv:2609.16730v1 Announce Type: cross Abstract: Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation… 33 The Information — AI news-outlet 12d ago AI Agent Startup Instinct in Talks for $10 Billion Valuation The startup behind personal AI assistant Instinct is in talks to raise $1 billion at a valuation of about $10 billion, a person with knowledge of the deal said. The Information reported earlier this month that the startup is looking to raise new funding as the invitation-only… 9 The Information — AI news-outlet 12d ago OpenAI in Early Talks for New Funding Round at $1.2 Trillion Valuation OpenAI has held early conversations with investors about a new funding round that could lift its valuation to $1.2 trillion or higher, according to people familiar with the conversations. The conversations come as OpenAI has pushed off a planned initial public offering to next… 33 TechCrunch — AI news-outlet 12d ago AEO startup Profound hits unicorn valuation, raises $180M Series D 7 months after last round Profound has raised a $180 million Series D at a $1.8 billion valuation, less than seven months after it raised a $96 million Series C. 21 Marcus on AI community 12d ago BREAKING: Secret US AI evaluation framework has been partly revealed Very partly 23 TechCrunch — AI news-outlet 12d ago Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents Their startup, Artificial Intelligence Underwriting Company (AIUC) has raised $40 million in a Series A round led by Ribbit Capital, with participation from First Harmonic. 23 r/LocalLLaMA community 13d ago Dual AMD Radeon AI Pro R9700 or dual NVIDIA or RTX 3090. I’m building a dual GPU box. Originally the goal was 30b at FP8 or 70b at Q4. I planned to run two R9700’s, but I found a pair of NIB RTX 3090’s near me for $1600 each. Discount if I buy both. RTX 40/50 series are off the table. I would love an RTX PRO 6000, but that is also off… 16 arXiv — Machine Learning research 13d ago BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents arXiv:2609.13149v1 Announce Type: new Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an… 32 arXiv — Machine Learning research 13d ago A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion arXiv:2609.13511v1 Announce Type: new Abstract: Large language models are increasingly benchmarked against classical machine learning for network intrusion detection (NIDS), almost always using same-dataset evaluation, and that protocol turns out to be incomplete. Evaluating… 19 arXiv — Machine Learning research 13d ago When Compliance Data Masquerades as Evaluation: Measurement Validity for Deployed AI Systems arXiv:2609.13642v1 Announce Type: new Abstract: We argue that a recurring failure in the evaluation of deployed AI systems occurs when data collected for operational monitoring or regulatory compliance are interpreted as if they were designed for comparative evaluation.… 33 arXiv — Machine Learning research 13d ago JumpStart Your Policy Learning with Lessons from 160,000 Training Runs arXiv:2609.13730v1 Announce Type: new Abstract: Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter… 10 arXiv — Machine Learning research 13d ago Diagnosing Temporal Misalignment in Multichannel Time-Series Classification with Minimum Description Length arXiv:2609.14595v1 Announce Type: new Abstract: Multichannel time-series classification commonly assumes synchronized sensor streams, although latency, clock drift, and preprocessing can introduce relative delays during data collection or after deployment. Existing… 25 arXiv — NLP / Computation & Language research 13d ago Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment arXiv:2609.13454v1 Announce Type: new Abstract: Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future… 19 arXiv — NLP / Computation & Language research 13d ago From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models arXiv:2609.13520v1 Announce Type: new Abstract: While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in… 31 arXiv — NLP / Computation & Language research 13d ago In the Blind: Building Pseudo-References for MT Evaluation arXiv:2609.13611v1 Announce Type: new Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these… 19 arXiv — NLP / Computation & Language research 13d ago When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings arXiv:2609.13824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent… 10 arXiv — NLP / Computation & Language research 13d ago Measuring the Cost of Variety Conflation in Multilingual MT Evaluation: Adding Mozambican Xichangana, Nyanja and Sena to FLORES+ arXiv:2609.13847v1 Announce Type: new Abstract: In this paper, we extend FLORES+ with Portuguese-source evaluation sets for three Mozambican Bantu varieties: Xichangana, Mozambican Nyanja, and Sena. We compare Xichangana with the existing Tsonga reference and Mozambican Nyanja… 10 arXiv — NLP / Computation & Language research 13d ago Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context arXiv:2609.13980v1 Announce Type: new Abstract: Arabic large-language-model (LLM) evaluation has matured around Modern Standard Arabic (MSA): aggregated leaderboards such as the Open Arabic LLM Leaderboard (OALL), HELM Arabic, and BALSAM rank models across dozens of MSA tasks,… 37 arXiv — NLP / Computation & Language research 13d ago Document Topic Alignment Metrics for Evaluating Topic Models of Short-Text Public Health Communications on Social Media arXiv:2609.14256v1 Announce Type: new Abstract: Topic models are widely used to analyze public health-related social media short texts, yet their evaluation remains dominated by metrics that focus entirely on generated topics alone. There is a lack of metrics that quantitatively… 27 arXiv — NLP / Computation & Language research 13d ago E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning arXiv:2609.14302v1 Announce Type: new Abstract: Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not… 6 arXiv — NLP / Computation & Language research 13d ago Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error arXiv:2609.14400v1 Announce Type: new Abstract: Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple… 12 arXiv — NLP / Computation & Language research 13d ago Mind Which Bird You Favour: Parameterizing Adequacy-Fluency Balance in Meta-Evaluation of Machine Translation arXiv:2609.14795v1 Announce Type: new Abstract: There is a tradeoff in machine translation meta-evaluation between prioritizing alignment with adequacy versus fluency. The balance depends on the combination of translation systems in the meta-evaluation dataset. This system set… 31 arXiv — NLP / Computation & Language research 13d ago A primer on evaluation methods for large language models in healthcare arXiv:2609.14819v1 Announce Type: new Abstract: Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine… 26 arXiv — NLP / Computation & Language research 13d ago One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs arXiv:2609.14860v1 Announce Type: new Abstract: Warning: This submission studies stereotypes and biases, and contains toxic and offensive examples, used for illustration purposes only. Fairness benchmarks such as BBQ have become the de facto standard for fairness evaluation… 34 Hugging Face Daily Papers research 13d ago Dream-RSI: Recursive Self-Improvement through Evolving Worlds Abstract Dream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations. Generated by thinkingmachines/Inkling-Small Recursive self-improvement is becoming increasingly vital… 25 The Information — AI news-outlet 13d ago Defense Startup Shield AI in Talks for Valuation of at Least $20 Billion Shield AI, a startup building drones and AI-powered software for the military, is in talks to raise new funds at a valuation of at least $20 billion, according to a person with knowledge of the fundraise. The funding round, if it closes, would represent a roughly 60% increase to… 36 The Information — AI news-outlet 13d ago Why China’s Answer to Surge AI Got a $1 Billion Valuation On Just $30 Million in Orders Unless you’ve been under a rock over the weekend, you’d know that the debate about AI safety took a big leap forward, as Anthropic CEO Dario Amodei on Saturday called for leading AI companies to slow the development of advanced AI, drawing supportive responses from OpenAI CEO… 22 Hugging Face Daily Papers research 13d ago ActionSplice: In-Flight Action Editing for Interactive World Models Abstract ActionSplice introduces counterfactual state transport to splice revised actions into chunk-autoregressive video world models without replaying completed evaluations, improving fidelity and speed. Generated by thinkingmachines/Inkling-Small Chunk-autoregressive video… 38 Hugging Face Daily Papers research 13d ago DataFlex-RL: An Evaluation Platform for RLVR Data Policies Abstract DataFlex-RL evaluates reinforcement learning data policies and finds that uniform sampling matches or exceeds adaptive rollout selection, reweighting, and domain mixing across math, logic, and science benchmarks. Generated by thinkingmachines/Inkling-Small Data policies… 8 arXiv — Machine Learning research 14d ago Fundamental Dynamical Units for Physics-Informed Structural Inference from Perturbation Time-Series in Networked Systems arXiv:2609.11934v1 Announce Type: new Abstract: In networked dynamical systems, the parameter of primary mechanistic interest is signed interaction structure. Recovering this structure from perturbation time-series data is a fundamental identification problem, compounded by… 30 arXiv — Machine Learning research 14d ago On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health arXiv:2609.11961v1 Announce Type: new Abstract: Stress is a pervasive determinant of mental health and a key target for mobile health interventions. On-device language models (ODLMs) offer privacy-preserving inference without cloud dependency, yet their feasibility for health… 35 arXiv — Machine Learning research 14d ago Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration arXiv:2609.12002v1 Announce Type: new Abstract: LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge… 15 arXiv — Machine Learning research 14d ago Certifying Concept Unlearning in Text-to-Image Diffusion Models arXiv:2609.12163v1 Announce Type: new Abstract: Existing evaluations of concept unlearning in text-to-image (T2I) diffusion models primarily rely on attack success rates obtained through automated adversarial prompt search. However, these metrics provide only empirical evidence… 16 arXiv — Machine Learning research 14d ago PLSP (Pre-hoc Liminal Space Profiling): OOD Prediction over Detection -- An Anticipatory Approach for Machine Learning Model Reliability arXiv:2609.12225v1 Announce Type: new Abstract: Out-of-Distribution (OOD) data poses a significant threat to machine learning models, often leading to model failure during deployment. All existing OOD detection methods are post-hoc, relying on evaluation metrics such as accuracy… 5 arXiv — Machine Learning research 14d ago ProactiveBench: Can Streaming Video Models Really Interact Like Humans? arXiv:2609.12658v1 Announce Type: new Abstract: Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query a model at a selected timestamp and therefore do… 26 arXiv — Machine Learning research 14d ago Robust Policy Optimization via Adversarial Importance Sampling arXiv:2609.13044v1 Announce Type: new Abstract: Significant progress has been made in safeguarding deep reinforcement learning (DRL) policies against input perturbations. Developing robust DRL involves three main stages: algorithm design, implementation, and evaluation. In this… 14 arXiv — NLP / Computation & Language research 14d ago GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents arXiv:2609.12191v1 Announce Type: new Abstract: Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the… 30 arXiv — NLP / Computation & Language research 14d ago Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models arXiv:2609.12475v1 Announce Type: new Abstract: Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluation unnecessarily expensive. Although recent benchmark compression methods (BCMs)… 13 arXiv — NLP / Computation & Language research 14d ago Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents arXiv:2609.13117v1 Announce Type: new Abstract: Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The… 38 arXiv — NLP / Computation & Language research 14d ago Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf arXiv:2609.12446v1 Announce Type: cross Abstract: Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing… 17 arXiv — NLP / Computation & Language research 14d ago MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation arXiv:2512.14691v3 Announce Type: replace Abstract: Modern multimodal generative models can synthesize visually compelling images and videos, but it remains unclear whether this visual fluency reflects genuine reasoning: when prompted to generate a solution, can a model preserve… 8 arXiv — NLP / Computation & Language research 14d ago From Bench-to-Bedside: A Review of Clinical Trials in Drug Discovery and Development arXiv:2412.09378v4 Announce Type: replace-cross Abstract: Clinical trials bridge basic research and clinical application, serving as essential steps in drug development. This review examines clinical trial phases (Phase I [safety assessment], Phase II [efficacy evaluation],… 38 Hugging Face Daily Papers research 14d ago Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Abstract Benchmark Radar is a searchable living database and discovery engine for AI evaluation benchmarks that aggregates sources, score histories, and evidence to support benchmark selection and comparison. Generated by thinkingmachines/Inkling-Small Benchmark researchers and… 7 Hugging Face Daily Papers research 14d ago COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization Abstract COBRA-Skills improves LLM agent skill optimization by using contextual-bandit prioritization and evidence-based evolution to cut evaluation costs while maintaining high performance. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents can… 21 The Information — AI news-outlet 16d ago Nvidia May Invest Up to $10 Billion In Anthropic’s IPO Nvidia has discussed investing in Anthropic’s upcoming initial public offering that would raise as much as $100 billion at a valuation of around $2 trillion, Reuters reported Friday. Nvidia could invest as much as $10 billion in Anthropic at the IPO price, according to the… 19 TechCrunch — AI news-outlet 16d ago Mecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training data The round for the two-year-old startup is coming together months after Mecka announced its Series A. 32 arXiv — Machine Learning research 17d ago Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables arXiv:2609.10778v1 Announce Type: new Abstract: Machine learning models can achieve strong test performance while relying on demographic or acquisition-related shortcuts. We propose counterfactual (CF) marginalisation as a test-time evaluation procedure for assessing robustness… 25 arXiv — Machine Learning research 17d ago When does a spectral prior help graph learning? Connectivity-loss estimation under road-network disruptions arXiv:2609.11166v1 Announce Type: new Abstract: Rapid evaluation of many simultaneous road-link disruptions requires a practical compromise between exact spectral recomputation and local approximation. We estimate relative algebraic-connectivity loss after multi-edge deletion… 22 arXiv — Machine Learning research 17d ago CausalArena: Benchmarking Causal Discovery in the Foundation Model Era arXiv:2609.11897v1 Announce Type: new Abstract: Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a… 14 arXiv — Machine Learning research 17d ago A Station-Based Evaluation of Machine Learning-based Weather Forecasting Models in Northern Norway arXiv:2609.10564v1 Announce Type: cross Abstract: Recent machine learning weather prediction (MLWP) models have demonstrated remarkable forecasting skill on global reanalysis-based benchmarks. However, their performance remains unclear in challenging environments such as… 27 Page 4 of 10 · 500 articles ← Newer Older →