News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — Machine Learning research 3h ago Unifying Generative Models with Path Integrals arXiv:2608.12438v1 Announce Type: new Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. Its… 36 arXiv — Machine Learning research 3h ago When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide arXiv:2608.12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k… 18 arXiv — Machine Learning research 3h ago GENADA: efficient generative time series adversarial attack framework arXiv:2608.12535v1 Announce Type: new Abstract: Deep learning models are widely used for time series analysis in domains such as healthcare, finance, energy systems, and environmental monitoring. However, these models remain vulnerable to adversarial attacks, where small input… 17 arXiv — Machine Learning research 3h ago H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities arXiv:2608.12926v1 Announce Type: new Abstract: Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT)… 31 arXiv — Machine Learning research 3h ago Incremental Evaluation and Training in Relational Deep Learning arXiv:2608.13023v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. However, prevailing RDL evaluation practices rely on static, single-episode dataset… 24 arXiv — Machine Learning research 3h ago A Probe Direction Is a Property of Its Prompt arXiv:2608.13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts… 8 arXiv — Machine Learning research 3h ago Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation arXiv:2608.13337v1 Announce Type: new Abstract: Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent fires at many tokens, and the effect has to be… 14 arXiv — NLP / Computation & Language research 3h ago On Measuring Semantic Preservation in Legal Ontology Learning arXiv:2608.12326v1 Announce Type: new Abstract: Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on… 26 arXiv — NLP / Computation & Language research 3h ago AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement arXiv:2608.12329v1 Announce Type: new Abstract: Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic… 32 arXiv — NLP / Computation & Language research 3h ago Query Timing Produces Opposite Positional Biases Between LLMs and Humans arXiv:2608.12387v1 Announce Type: new Abstract: Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and… 11 arXiv — NLP / Computation & Language research 3h ago BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian arXiv:2608.12894v1 Announce Type: new Abstract: Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian… 35 arXiv — NLP / Computation & Language research 3h ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 3h ago Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability… 20 TechCrunch — AI news-outlet 11h ago Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation. AI is expensive, Ali Ghodsi tells TechCrunch. With so many investors wanting into his latest round, he said yes to more than planned. 18 arXiv — Machine Learning research 1d ago Long-Horizon Forecasting of Complete Financial Statements with Forma arXiv:2608.11327v1 Announce Type: new Abstract: Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm… 6 arXiv — Machine Learning research 1d ago Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits arXiv:2608.11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and… 37 arXiv — Machine Learning research 1d ago Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark arXiv:2608.11423v1 Announce Type: new Abstract: Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across… 5 arXiv — Machine Learning research 1d ago When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits arXiv:2608.11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online… 4 arXiv — Machine Learning research 1d ago Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing arXiv:2608.11704v1 Announce Type: new Abstract: Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are effective for time-series classification but are vulnerable to mislabeled training samples and require numerous DTW computations during inference. We propose… 6 arXiv — Machine Learning research 1d ago JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series arXiv:2608.11801v1 Announce Type: new Abstract: Multivariate time-series anomaly prediction aims to identify whether and when anomalies will occur over a future horizon from historical observations. Existing methods primarily characterize anomalies as deviations in future… 18 arXiv — Machine Learning research 1d ago Towards Truly Unsupervised Evaluation of Feature Selection arXiv:2608.12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method. Most of the methods… 9 arXiv — NLP / Computation & Language research 1d ago TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation arXiv:2608.11236v1 Announce Type: new Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic… 28 arXiv — NLP / Computation & Language research 1d ago Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment arXiv:2608.11528v1 Announce Type: new Abstract: Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree… 23 arXiv — NLP / Computation & Language research 1d ago LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification arXiv:2608.11753v1 Announce Type: new Abstract: Financial text is produced and interpreted within a market environment, yet financial text classifiers almost always receive text alone. We study whether financial time series are useful as an additional input on the task of… 33 arXiv — NLP / Computation & Language research 1d ago GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation arXiv:2608.11787v1 Announce Type: new Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision… 6 arXiv — NLP / Computation & Language research 1d ago When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation arXiv:2608.11843v1 Announce Type: new Abstract: The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name -- a misread name corrupts the historical fact rather than… 28 arXiv — NLP / Computation & Language research 1d ago Benchmarking LLM Judges for Mobile Agent Evaluation arXiv:2608.11434v1 Announce Type: cross Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark… 17 arXiv — NLP / Computation & Language research 1d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused… 25 arXiv — NLP / Computation & Language research 1d ago Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation arXiv:2608.12150v1 Announce Type: cross Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across… 9 Hugging Face Daily Papers research 1d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Abstract ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents… 12 Hugging Face Daily Papers research 1d ago SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Abstract SkillZip compresses self-evolving agent skills by finding a minimal faithful structural explanation that shares repeated rules and procedures while preserving rare exceptions, without requiring evaluation rollouts. Generated by thinkingmachines/Inkling-Small… 26 TechCrunch — AI news-outlet 1d ago AI coding startup Cognition reportedly already in talks to raise at $40B valuation Cognition may be looking to raise another mega round just a few months after raising $1 billion at a $26 billion valuation. 27 TechCrunch — AI news-outlet 1d ago OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Alitmeter Capital. 23 TechCrunch — AI news-outlet 1d ago Lovable confirms new $13.3B valuation, raises another $400M This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch. 24 TechCrunch — AI news-outlet 1d ago Everything announced at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features From the Pixel 11 series and a brand new competitor to Apple’s AirTag, here are all the announcements from the Made by Google 2026 event. 23 TechCrunch — AI news-outlet 1d ago AI code-testing startup Blacksmith’s valuation jumps almost 10x in less than a year Blacksmith says revenue has grown more than tenfold over the past year. 26 arXiv — Machine Learning research 2d ago The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. We reproduce that result by… 18 arXiv — Machine Learning research 2d ago A matched-integrator evaluation of Hamiltonian neural networks on pendulum and Kepler dynamics arXiv:2608.10235v1 Announce Type: new Abstract: Hamiltonian Neural Networks (HNNs) parameterize conservative dynamics through a learned scalar Hamiltonian, providing an architectural prior that is absent from generic vector-field neural networks. We evaluate this prior under a… 34 arXiv — Machine Learning research 2d ago Toward Human Rights Benchmarking for LLMs: A Pilot Methodology arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this… 14 arXiv — Machine Learning research 2d ago Retrieval-Corrected Conformal Prediction for Time Series arXiv:2608.10553v1 Announce Type: new Abstract: Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and… 12 arXiv — NLP / Computation & Language research 2d ago Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization arXiv:2608.10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total… 25 arXiv — Machine Learning research 2d ago Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems arXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operational status of complex dynamical systems. However, collecting such data is often… 23 arXiv — Machine Learning research 2d ago Do AI weather models miss extremes? arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European… 30 arXiv — NLP / Computation & Language research 2d ago The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a… 30 arXiv — NLP / Computation & Language research 2d ago ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct… 5 arXiv — NLP / Computation & Language research 2d ago Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first… 28 arXiv — NLP / Computation & Language research 2d ago VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? arXiv:2608.10875v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs… 20 arXiv — NLP / Computation & Language research 2d ago A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models arXiv:2608.10939v1 Announce Type: new Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and… 23 arXiv — NLP / Computation & Language research 2d ago Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product:… 38 arXiv — NLP / Computation & Language research 2d ago OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that… 31 Page 1 of 10 · 500 articles Older →