News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 7d ago Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation arXiv:2609.20875v1 Announce Type: cross Abstract: Parkinson's disease (PD) often manifests through speech impairments, facilitating accessible, non-invasive, and cost-effective severity assessment for early diagnosis and progression tracking. Despite advances in speech… 31 arXiv — NLP / Computation & Language research 7d ago CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation arXiv:2609.21793v1 Announce Type: cross Abstract: Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine… 35 arXiv — NLP / Computation & Language research 7d ago MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks arXiv:2602.16313v2 Announce Type: replace Abstract: Existing evaluations of agents with memory typically assess memorization and action in isolation. One class of benchmarks evaluates memorization by testing recall of past conversations or text but fails to capture how memory is… 18 arXiv — NLP / Computation & Language research 7d ago Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation arXiv:2604.08797v2 Announce Type: replace Abstract: Stories are key to transmitting values across cultures, but their interpretation varies across linguistic and cultural contexts. Thus, we introduce multilingual story moral generation as a novel culturally grounded evaluation… 11 Vercel — AI dev-tools 7d ago AI Gateway now supports TypeSafe clients and an HTTP API for Jev You can now call Jev from TypeSafe AI through AI Gateway using an existing TypeSafe client or the HTTP API, in addition to the AI SDK. TypeSafe client: Point an existing TypeSafe client at AI Gateway without changing its evaluation calls. HTTP API: Call Jev directly from any… 10 r/MachineLearning community 8d ago Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D] Since OpenAI retired SWE-bench Verified in February (every frontier model tested could reproduce reference fixes for some tasks; underspecified tests rewarded knowing the intended fix), I've been trying to write down precisely what a decontamination report can and can't… 27 Don't Worry About the Vase community 8d ago Anthropic Looks At Some Of Its Alignment Problems Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. 32 TechCrunch — AI news-outlet 9d ago A startup that builds other startups raised $100M, and is all-in on physical AI UP.Labs, now doing business under the name Vantora, is building startups for industrial corporations. 35 The Information — AI news-outlet 9d ago Google’s Gemini Model Hacks Companies During Test Google acknowledged that its Gemini AI model unexpectedly breached the networks of three outside companies during safety evaluations conducted by third-party testing firm Irregular last May, The Wall Street Journal reported. During the exercise, Gemini gained entry to the… 33 r/LocalLLaMA community 9d ago Is HF starting to move against abliterated models? Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models. The announcement lands amid debate for the… 36 TechCrunch — AI news-outlet 9d ago Manus seeks $4B valuation in new $500M fundraise as it resumes independent ops Manus, which earlier this year had to break off a merger with Meta, is in discussions to raise $500M at a $4B valuation. 20 r/LocalLLaMA community 9d ago Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison Hey r/LocalLLaMA , Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark . We wanted to see… 24 r/LocalLLaMA community 9d ago China’s mysterious AI company Naive AI is valued at over $1.4 billion and could release its first open-source llm model as early as this month. According to people familiar with the matter, Naive AI, an AI startup founded in February this year by Tsinghua University professor Dai Jifeng, has completed three funding rounds totaling $400 million, bringing its valuation to more than $1.4 billion and earning it unicorn… 5 The Information — AI news-outlet 9d ago A Tsinghua Professor’s Stealth LLM Startup Hits $1.4 Billion Valuation A secretive Chinese AI model startup founded in February by a Tsinghua University professor is now valued at more than $1.4 billion, after raising $400 million in three funding rounds from investors including Tencent , according to a person with direct knowledge of the matter.… 11 arXiv — Machine Learning research 10d ago Bayesian Optimization with Rich Auxiliary Information via LLMs arXiv:2609.19437v1 Announce Type: new Abstract: Bayesian Optimization (BO) is widely used for optimizing expensive black-box functions, yet many real-world optimization problems contain substantially richer information than function evaluations alone. Examples include training… 9 arXiv — Machine Learning research 10d ago QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization arXiv:2609.20156v1 Announce Type: new Abstract: Ubiquitous time series data across diverse domains enables critical applications in areas such as transportation systems and power grids. Recently, training foundation models on massive datasets to achieve accurate zero-shot… 11 arXiv — Machine Learning research 10d ago When Does Retrieval Help Time-Series Forecasting? arXiv:2609.20193v1 Announce Type: new Abstract: Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry. Published evaluations report consistent gains, and each credits its own mechanism. We show that the benefit belongs instead to the… 29 arXiv — NLP / Computation & Language research 10d ago Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search arXiv:2609.19799v1 Announce Type: new Abstract: LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show… 30 arXiv — NLP / Computation & Language research 10d ago V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering arXiv:2609.19879v1 Announce Type: new Abstract: Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the… 11 arXiv — NLP / Computation & Language research 10d ago KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms arXiv:2609.19916v1 Announce Type: new Abstract: Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary… 32 arXiv — NLP / Computation & Language research 10d ago HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication arXiv:2609.20684v1 Announce Type: new Abstract: Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a… 8 arXiv — NLP / Computation & Language research 10d ago Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations arXiv:2609.20779v1 Announce Type: new Abstract: Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit… 7 arXiv — NLP / Computation & Language research 10d ago What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks arXiv:2609.19182v1 Announce Type: cross Abstract: Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of… 38 arXiv — NLP / Computation & Language research 10d ago EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data arXiv:2609.19523v1 Announce Type: cross Abstract: Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified… 13 arXiv — Machine Learning research 11d ago Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives arXiv:2609.17572v1 Announce Type: new Abstract: Auditing vision-language models (VLMs) for societal bias requires distinguishing direct algorithmic valuation disparities from confounders embedded within archival metadata. In this study, we audit Contrastive Language-Image… 24 arXiv — NLP / Computation & Language research 11d ago Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation arXiv:2609.17544v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library… 33 arXiv — NLP / Computation & Language research 11d ago I code or AI code: A comparative evaluation of AI-rated scores in classroom observations arXiv:2609.18274v1 Announce Type: new Abstract: Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study… 26 arXiv — NLP / Computation & Language research 11d ago DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions arXiv:2609.18649v1 Announce Type: new Abstract: Warning: This paper contains examples of stereotypes and social bias. LLMs are increasingly used in interactive settings by the general public, making the evaluation of model behavior in multi-turn conversational scenarios… 28 arXiv — NLP / Computation & Language research 11d ago LocQE: Principled Domain Adaptation for Localisation Quality Estimation by Leveraging Post-Edits arXiv:2609.18720v1 Announce Type: new Abstract: Learned quality estimation (QE) models such as COMETKiwi are widespread and work well for general machine translation evaluation. However, they are known to struggle on unseen domains, limiting their performance in a real-world… 7 arXiv — NLP / Computation & Language research 11d ago ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts arXiv:2609.18844v1 Announce Type: new Abstract: Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often… 21 arXiv — NLP / Computation & Language research 11d ago Reading Between the Lines: Can LLMs Discover the Question Behind the Text? arXiv:2609.19070v1 Announce Type: new Abstract: This paper introduces ``question archaeology'', a specific evaluation task focused on inferring the single, authentic "genesis question" that motivated the creation of a complete text. Distinct from question generation, which… 16 arXiv — NLP / Computation & Language research 11d ago Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators arXiv:2609.19072v1 Announce Type: new Abstract: Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks… 31 arXiv — NLP / Computation & Language research 11d ago Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation arXiv:2609.19093v1 Announce Type: new Abstract: Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology,… 9 arXiv — NLP / Computation & Language research 11d ago Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations arXiv:2609.19101v1 Announce Type: new Abstract: As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in… 38 The Information — AI news-outlet 11d ago Zipline in Talks to Raise Funds at $20 Billion Valuation Zipline, a startup making drones that can deliver small packages like meals and medical supplies, is in talks to raise new funds at a valuation of roughly $20 billion, according to two people with knowledge of the fundraise. The funding round, if it closes, would more than… 5 Latent.Space news-outlet 11d ago Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC We sit down with AIUC’s CEO on their Series A! 26 r/LocalLLaMA community 11d ago Apple May Return to Server Market With Nvidia Technology Apple is considering offering an AI server built around "M8" series chips and has discussed incorporating Nvidia networking hardware, The Information reports. The system would be sold to outside customers, potentially bringing Apple back into a business it left behind when it… 30 arXiv — Machine Learning research 12d ago A panoramic aerodynamic performance prediction method for turbomachinery cascades using transformer-enhanced neural operator arXiv:2609.16066v1 Announce Type: new Abstract: To enable flexible and rapid aerodynamic performance evaluation in turbomachinery design, this paper proposes a panoramic performance prediction framework. Unlike most previous prediction models that directly predict the objective… 5 arXiv — Machine Learning research 12d ago LLMs as Master Forgers: Generating Synthetic Time Series Data for Manufacturing arXiv:2609.16155v1 Announce Type: new Abstract: This paper presents a novel framework leveraging Large Language Models (LLMs) to generate synthetic time series data for manufacturing processes. Motivated by the scarcity of labeled time-series data in real-world manufacturing… 32 arXiv — NLP / Computation & Language research 12d ago The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting arXiv:2609.16267v1 Announce Type: cross Abstract: Many operational cases are documented more than once, at different workflow stages and for different purposes, yet model evaluations normally select one of these records before model comparison begins. We treat that selection as… 22 arXiv — Machine Learning research 12d ago A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids arXiv:2609.16744v1 Announce Type: new Abstract: The integration of renewable energy sources into the electrical grid introduces complex challenges in fault detection and coordination of grid recovery mechanisms. Traditional relay protection systems, which operate based on static… 10 arXiv — NLP / Computation & Language research 12d ago ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals arXiv:2609.16816v1 Announce Type: cross Abstract: Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading. Such rubrics are reliable only if they reward honest answers over… 35 arXiv — NLP / Computation & Language research 12d ago Comment on arXiv:2607.01233: Survivorship Bias in Published-Paper Baselines for Research-Idea Distributions arXiv:2609.15996v1 Announce Type: new Abstract: Chen, Zhao, and Cohan introduce a valuable distributional evaluation of LLM-generated research ideas. This comment raises a narrower identification concern: their human baseline consists of published papers, whereas the LLM… 35 arXiv — NLP / Computation & Language research 12d ago Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks arXiv:2609.16023v1 Announce Type: new Abstract: Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively, however, is expensive, and rubric-based evaluation has become the… 25 arXiv — NLP / Computation & Language research 12d ago Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits arXiv:2609.16501v1 Announce Type: new Abstract: De-identified r\'esum\'e screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields… 14 arXiv — NLP / Computation & Language research 12d ago Challenges of Auditing: Variability in Outputs of Large Language Models for Health arXiv:2609.16590v1 Announce Type: new Abstract: People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations… 29 arXiv — NLP / Computation & Language research 12d ago Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models arXiv:2609.16739v1 Announce Type: new Abstract: Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain… 18 arXiv — NLP / Computation & Language research 12d ago Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM arXiv:2609.17435v1 Announce Type: new Abstract: We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the… 25 arXiv — NLP / Computation & Language research 12d ago Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models arXiv:2609.16006v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test what a model knows rather than how it behaves when giving open-ended… 5 arXiv — NLP / Computation & Language research 12d ago CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection arXiv:2609.16582v1 Announce Type: cross Abstract: Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation… 25 Page 3 of 10 · 500 articles ← Newer Older →