News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow TechCrunch — AI news-outlet 1d ago XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation The round is being raised just months after the robot data startup exited from stealth. 6 Hugging Face Daily Papers research 1d ago VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement Abstract VeriPhy verifies generated video by compiling prompts into typed physical obligations, executing frozen expert analyses with provenance tracking, and mapping evidence to auditable three-valued verdicts. Generated by thinkingmachines/Inkling-Small Visual fluency in… 12 GitHub Blog — AI & ML official-blog 2d ago Project HydraFusion: Frontier quality via multi-model orchestration In controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline while reducing estimated workflow cost. Now available as a research preview in GitHub Copilot. The post Project HydraFusion: Frontier quality via… 29 Hugging Face Daily Papers research 2d ago Last Translation Benchmark Abstract The Last Translation Benchmark introduces peer-reviewed, multimodal examples that break leading translation models alongside handcrafted verification rules for reliable, actionable evaluation. Generated by thinkingmachines/Inkling-Small For scientific progress, we need… 22 The Information — AI news-outlet 2d ago Nvidia Discusses $2.5 Billion Investment in Mira Murati’s Thinking Machines Lab Mira Murati’s Thinking Machines Lab is in talks to raise between $5 billion and $6 billion at a pre-money valuation of at least $40 billion, with venture firm Accel in discussions to lead the round, The Information reported . Chipmaking giant Nvidia is expected to contribute… 5 arXiv — Machine Learning research 2d ago Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating arXiv:2609.03900v1 Announce Type: new Abstract: Continual knowledge-updating methods are often declared superior from one final checkpoint and one conventional adapter rank. We show that this can be insufficient to identify the better operating point. Holding a periodic… 27 arXiv — Machine Learning research 2d ago You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring arXiv:2609.03035v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fooled by obfuscation and steganography, motivating… 38 arXiv — NLP / Computation & Language research 2d ago Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation arXiv:2609.02942v1 Announce Type: new Abstract: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants… 12 arXiv — NLP / Computation & Language research 2d ago FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models arXiv:2609.03331v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually… 19 arXiv — NLP / Computation & Language research 2d ago Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks arXiv:2609.03432v1 Announce Type: new Abstract: Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and… 22 arXiv — NLP / Computation & Language research 2d ago Last Translation Benchmark arXiv:2609.04173v1 Announce Type: new Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are… 20 arXiv — NLP / Computation & Language research 2d ago VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis arXiv:2609.03203v1 Announce Type: cross Abstract: Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause,… 16 TechCrunch — AI news-outlet 2d ago Crusoe reportedly raises $3B at a $30B valuation The round came together after the data center developer reportedly secured a $13 billion contract with Jane Street. 32 TechCrunch — AI news-outlet 3d ago Accel reportedly in talks to lead $1B round for Thinking Machines at $40B valuation The high-profile startup's annual revenue run rate stands at at over $100 million. 28 The Information — AI news-outlet 3d ago Thinking Machines Lab In Talks to Raise Billions at Roughly $40 Billion Valuation Thinking Machines Lab, the AI developer led by ex-OpenAI chief technology officer Mira Murati, is in talks to raise at least $1 billion at a valuation of at least $40 billion before the investment, according to a person with knowledge of the fundraise. Existing investor Accel is… 28 arXiv — Machine Learning research 3d ago GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories arXiv:2609.02160v1 Announce Type: new Abstract: Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping… 20 arXiv — Machine Learning research 3d ago Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators arXiv:2609.02304v1 Announce Type: new Abstract: A fundamental quantity in machine learning is the optimal performance achievable by any model on a given task. Estimating this quantity allows us to distinguish the irreducible part of the error from a deficiency of the model,… 30 arXiv — Machine Learning research 3d ago Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights arXiv:2609.02652v1 Announce Type: new Abstract: Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own evaluation protocol. Its kernel decodes one shell; we found no implementation of the multi-shell decoder the rate requires. This paper… 30 arXiv — Machine Learning research 3d ago Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search arXiv:2609.01617v1 Announce Type: cross Abstract: Getting accurate, grounded answers out of large enterprise document repositories is a difficult problem. Dense vector retrieval alone frequently performs poorly on queries that mix technical terminology, vendor-specific acronyms,… 18 arXiv — Machine Learning research 3d ago FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making arXiv:2609.01691v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to make decisions from visual inputs. We introduce FAIRLENS, a benchmark and evaluation framework for measuring both the fairness and the validity of VLM responses in three… 21 arXiv — Machine Learning research 3d ago Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation arXiv:2609.01786v1 Announce Type: cross Abstract: Hyperspectral image classification still relies heavily on random pixel splits within a single scene. The Salinas dataset, randomly split, is among the most widely used datasets for comparing different architectures. However,… 17 arXiv — NLP / Computation & Language research 3d ago VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages arXiv:2609.01788v1 Announce Type: new Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and… 18 arXiv — NLP / Computation & Language research 3d ago Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation arXiv:2609.02163v1 Announce Type: new Abstract: Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic… 30 arXiv — NLP / Computation & Language research 3d ago Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation arXiv:2609.02396v1 Announce Type: new Abstract: Radiology reports are written primarily for clinicians, and their specialized terminology often makes them difficult for patients to interpret. As a result, many patients turn to publicly available Large Language Models (LLMs) to… 26 arXiv — NLP / Computation & Language research 3d ago Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking arXiv:2609.02414v1 Announce Type: new Abstract: Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework… 37 arXiv — NLP / Computation & Language research 3d ago EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction arXiv:2609.02783v1 Announce Type: new Abstract: Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid… 24 arXiv — NLP / Computation & Language research 3d ago EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models arXiv:2609.01611v1 Announce Type: cross Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of… 37 arXiv — NLP / Computation & Language research 3d ago Accurate in space, unreliable in time: how LLMs represent national cultural change arXiv:2609.01902v1 Announce Type: cross Abstract: Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only… 15 arXiv — NLP / Computation & Language research 3d ago Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds arXiv:2609.02302v1 Announce Type: cross Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that… 8 arXiv — NLP / Computation & Language research 3d ago Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model Selection arXiv:2609.02745v1 Announce Type: cross Abstract: Selecting a retrieval model for a production RAG system requires reliable comparative evaluation, but obtaining relevance judgments at scale is expensive and difficult to repeat as new candidate systems arrive. We study pooled… 33 Hugging Face Daily Papers research 3d ago EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Abstract EarlyEval predicts agent outcomes from intermediate behavior to reduce evaluation cost by halting runs early with minimal accuracy loss. Generated by thinkingmachines/Inkling-Small Evaluating LLM agents is essential for guiding their development, yet it has grown… 37 Hugging Face Daily Papers research 3d ago RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests Abstract Real-world coding requests are shorter and more casual than benchmark tasks, and explicitly stating desired behavior and motivation improves LLM software engineering performance. Generated by thinkingmachines/Inkling-Small Coding agents are now commonly evaluated on the… 9 TechCrunch — AI news-outlet 4d ago Wonderful more than doubles its valuation to $5B in under 6 months Wonderful said it will use its $550 million Series C funding to develop products faster, expand its FDE teams, and meet demand for its products. 4 TechCrunch — AI news-outlet 4d ago HiddenLayer nabs $100M as enterprises rush to secure their AI deployments HiddenLayer has raised a $100M Series B from Delta-v Capital, Ten Eleven Ventures, Morgan Stanley, Microsoft's M12, Booz Allen Hamilton, and others. 38 arXiv — Machine Learning research 4d ago Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification arXiv:2609.00093v1 Announce Type: new Abstract: Imbalanced time series classification is often addressed by changing the training distribution, objective, logits, or final threshold. These interventions address important biases, yet leave a representation-level question… 12 arXiv — Machine Learning research 4d ago When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting arXiv:2609.01126v1 Announce Type: new Abstract: Online adaptation can help edge time-series forecasting under distribution drift, but its measured benefit is sensitive to evaluation choices. We study six public multivariate streams, including building-sensor and smart-meter… 16 arXiv — Machine Learning research 4d ago Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data arXiv:2609.01262v1 Announce Type: new Abstract: Tabular deep learning (TDL) leverages neural networks (NN) to extract patterns from tabular data. Traditional TDL methods follow a supervised learning paradigm, where a target feature is explicitly given. In this work, however, we… 22 arXiv — Machine Learning research 4d ago SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers arXiv:2609.01343v1 Announce Type: new Abstract: Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts… 22 arXiv — NLP / Computation & Language research 4d ago trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories arXiv:2609.00038v1 Announce Type: new Abstract: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the… 17 arXiv — NLP / Computation & Language research 4d ago RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving arXiv:2609.00062v1 Announce Type: new Abstract: Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity… 8 arXiv — NLP / Computation & Language research 4d ago Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs arXiv:2609.00184v1 Announce Type: new Abstract: Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on… 14 arXiv — NLP / Computation & Language research 4d ago NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings arXiv:2609.00256v1 Announce Type: new Abstract: LLM-based diagnostic systems achieve high semantic accuracy on benchmarks, but open-ended evaluation on clinically uncommon presentations reveals a systematic gap between headline accuracy and verifiable clinical reliability. We… 29 arXiv — NLP / Computation & Language research 4d ago Toward Workflow-Aware Benchmarking for Healthcare NLP Agents arXiv:2609.00296v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet many evaluations remain limited to static medical… 18 arXiv — NLP / Computation & Language research 4d ago Human-Anchored Factuality Evaluation with Strategic Annotation arXiv:2609.00494v1 Announce Type: new Abstract: LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality evaluation under limited annotation budgets, where… 9 arXiv — NLP / Computation & Language research 4d ago Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents arXiv:2609.00549v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly rely on external skills, yet standard evaluations obscure whether retrieving these skills actually helps. Aggregate metrics often compare retrieved versus non-retrieved tasks,… 8 arXiv — NLP / Computation & Language research 4d ago Investigating Assistant Bias in LLM User Simulators Using a Role Vector arXiv:2609.00608v1 Announce Type: new Abstract: LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task… 24 arXiv — NLP / Computation & Language research 4d ago Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling arXiv:2609.00949v1 Announce Type: new Abstract: Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy.… 36 arXiv — NLP / Computation & Language research 4d ago Disclosure-Gated User Simulation for Companion-Agent Evaluation arXiv:2609.00982v1 Announce Type: new Abstract: Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of… 26 arXiv — NLP / Computation & Language research 4d ago Post-hoc Alignment of LLM-judges to Human Judgment Distribution arXiv:2609.01073v1 Announce Type: new Abstract: The LLM-as-a-judge (LLMaJ) framework offers a cost-effective and reproducible solution for automatic evaluation. However, current evaluation practices typically compare LLMaJ judgments against aggregated ground-truth labels,… 4 arXiv — NLP / Computation & Language research 4d ago Does task decomposition improve automatic NLG evaluation? arXiv:2609.01139v1 Announce Type: new Abstract: The LLM-as-a-judge (LLMaJ) framework has emerged as a promising solution for cheap, reproducible, reference-free Natural Language Generation (NLG) evaluation. Prior work seeks to improve LLMaJ by decomposing evaluation tasks into… 19 Page 1 of 10 · 500 articles Older →