News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago SlideLab: Audience-Centered Scientific Slide Generation and Evaluation arXiv:2609.30294v1 Announce Type: new Abstract: Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a… 30 arXiv — NLP / Computation & Language research 7h ago Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification arXiv:2609.30467v1 Announce Type: new Abstract: Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical… 16 arXiv — NLP / Computation & Language research 7h ago CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production arXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically… 4 arXiv — NLP / Computation & Language research 7h ago Inquesto Score: A reliability Protocol For Voice Agents arXiv:2609.30514v1 Announce Type: new Abstract: Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto… 22 arXiv — NLP / Computation & Language research 7h ago Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms arXiv:2609.30558v1 Announce Type: new Abstract: Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a… 23 arXiv — NLP / Computation & Language research 7h ago TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding arXiv:2609.30670v1 Announce Type: new Abstract: Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses… 10 arXiv — NLP / Computation & Language research 7h ago Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4 arXiv:2609.30716v1 Announce Type: new Abstract: When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on… 33 arXiv — NLP / Computation & Language research 7h ago Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations arXiv:2609.30867v1 Announce Type: new Abstract: Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model… 8 Vercel — AI dev-tools 1d ago Ember-1 from Fireworks now available on AI Gateway Ember-1 from Fireworks is now available on AI Gateway . Ember-1 is a research preview reasoning model built on Kimi K3 for coding and agentic workflows. Fireworks reports approximately 40% fewer generated tokens than Kimi K3 at comparable quality across its evaluations. For… 13 arXiv — Machine Learning research 3d ago Leakage-Safe Machine Learning for Hydrogen Embrittlement Detection in 316L Stainless Steel: A Region-Held-Out Evaluation of Texture and Deep Features in SEM Micrographs arXiv:2609.28567v1 Announce Type: new Abstract: Scanning electron microscopy (SEM) is routinely used to characterize the microstructural changes caused by hydrogen embrittlement (HE) in structural steels. Machine learning can automate this characterization, but models are often… 6 arXiv — Machine Learning research 3d ago UO-FIE: Combining Exact-Label Supervision with Graded Utility for Factivity Inference arXiv:2609.28605v1 Announce Type: new Abstract: The Factivity Inference Evaluation 2026 (FIE2026) classifies Chinese context-hypothesis pairs into nine ordered factivity intervals. Its evaluation metric rewards both exact predictions and proximity to the correct interval, while… 19 arXiv — Machine Learning research 3d ago fable.intermittent: benchmarking probabilistic forecasting methods for intermittent time series arXiv:2609.28607v1 Announce Type: new Abstract: Intermittent time series are common in spare-parts demand and retail sales. Since the cost of forecast errors is typically asymmetric, decisions such as inventory control require the full predictive distribution rather than a point… 17 arXiv — Machine Learning research 3d ago Learnable Time-Frequency Masks for Explaining Time-Series Classifiers arXiv:2609.29270v1 Announce Type: new Abstract: Time-series explainability remains challenging because discriminative information is often encoded in latent frequency or time-frequency features rather than in the raw signal itself. Existing attribution methods typically operate… 12 arXiv — Machine Learning research 3d ago Neuralized Multi-Wavelet Decomposition for Time Series Classification and Forecasting arXiv:2609.29317v1 Announce Type: new Abstract: Time series analysis is fundamental in domains such as finance, healthcare, and meteorology. Real-world time series often exhibit multiscale characteristics shaped by diverse latent factors, resulting in intricate temporal patterns… 32 arXiv — Machine Learning research 3d ago The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality arXiv:2609.29530v1 Announce Type: new Abstract: Validating a model on a time series asks for three things at once: each training run should use most of the sample (sufficiency), the test sets should together cover most of the sample (coverage), and training data should come… 5 arXiv — Machine Learning research 3d ago Three Ways Classical Test Theory Misleads for LLM Judges arXiv:2609.29709v1 Announce Type: new Abstract: An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without… 14 arXiv — Machine Learning research 3d ago SwitchPFN: Shared Switching Dynamics for Frozen In-Context Time Series Classification arXiv:2609.29814v1 Announce Type: new Abstract: Tabular foundation models (TFMs) provide a promising route to time-series classification, but their effectiveness depends on how sequential data are converted into tabular representations. Existing representations face two… 4 arXiv — NLP / Computation & Language research 3d ago Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks arXiv:2609.28673v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on… 23 arXiv — NLP / Computation & Language research 3d ago COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages arXiv:2609.28826v1 Announce Type: new Abstract: Machine translation (MT) for Indian languages remains constrained by the limited availability of high-quality, Indic-centric parallel corpora and evaluation benchmarks. Existing multilingual resources are largely constructed from… 24 arXiv — NLP / Computation & Language research 3d ago Reasoning Instructions Can Break Answer Decoding in Vision--Language Models arXiv:2609.29278v1 Announce Type: new Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On… 7 arXiv — NLP / Computation & Language research 3d ago Likelihood Ranking doesn't Scale Like Prompting in LLMs arXiv:2609.29390v1 Announce Type: new Abstract: LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still… 21 arXiv — NLP / Computation & Language research 3d ago ModularSQL: A Runtime Guardrail for the Multiplicity Blind Spot in Text-to-SQL arXiv:2609.29573v1 Announce Type: new Abstract: Text-to-SQL systems are increasingly deployed on production databases, where queries that pass benchmark evaluation can still produce results that distort downstream workflows. Standard set-based execution accuracy (Set-EX)… 36 arXiv — NLP / Computation & Language research 3d ago Stochastic Semantic Evidence Graphs: Uncertainty Propagation and Governance for Agentic AI arXiv:2609.29703v1 Announce Type: new Abstract: AI-agent evaluations usually inspect a final answer, yet error may enter through evidence, retrieval, prompting, generation or decision mapping. We introduce a stochastic semantic evidence graph (SSEG), a hierarchical stochastic… 20 arXiv — NLP / Computation & Language research 3d ago TimeBraid: Unifying Time Series and Language for Understanding and Forecasting arXiv:2609.29792v1 Announce Type: new Abstract: We present TimeBraid, a series of unified time-series and language models that align pretrained language models and pretrained time-series foundation models through interleaved global residual attention layers. Each model inherits… 29 arXiv — NLP / Computation & Language research 3d ago Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax arXiv:2609.29848v1 Announce Type: new Abstract: A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a… 13 arXiv — NLP / Computation & Language research 3d ago Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations arXiv:2609.29928v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity… 7 arXiv — NLP / Computation & Language research 3d ago How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure arXiv:2609.30074v1 Announce Type: new Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open… 14 arXiv — NLP / Computation & Language research 3d ago Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots arXiv:2609.29043v1 Announce Type: cross Abstract: General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to interpret diverse natural language commands and generate multi-step action sequences in real home environments. Conventional… 6 The Information — AI news-outlet 3d ago Startup Founded by ex-Tesla Dojo Leaders Nears $10 Billion Valuation DensityAI, an AI chip startup founded just a year ago by former leaders of Tesla’s Dojo supercomputer program, is in advanced talks to raise hundreds of millions of dollars in a round that would value it at $10 billion, according to two people with knowledge of the discussions.… 16 r/MachineLearning community 3d ago NeurIPS Evaluations and Datasets Track notifications are live on OpenReview [D] Mine was accepted with 5,4,3->5,5,3   submitted by   /u/tfburns [link]   [comments] 31 The Information — AI news-outlet 3d ago Jev Fervor Leads to Talk of Big Valuation Boost Silicon Valley has a new AI startup to obsess over—and throw money at. We’re hearing that TypeSafe AI, the startup developing a new kind of model called Jev, has started to talk to investors about raising an enormous round—$1 billion or even more—and that some unnamed investors… 10 TechCrunch — AI news-outlet 3d ago Ando wants to take on Slack with a team messaging app that lets humans and agents work together Ando has raised $20 million in pre-seed and seed funding from investors including Accel, Index Ventures and Emergence. 11 r/LocalLLaMA community 4d ago Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark. I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all… 24 arXiv — Machine Learning research 4d ago Signal2Symbol: Neuro-Symbolic Temporal Reasoning for Explainable Physiological Time-Series Anomaly Detection arXiv:2609.26820v1 Announce Type: new Abstract: Physiological time series such as electrocardiograms (ECG) and electroencephalograms (EEG) exhibit complex temporal structure, substantial acquisition variability, and a strong need for transparent decision-making. Although deep… 4 arXiv — Machine Learning research 4d ago A Leakage-Aware Multimodal Evaluation Framework for Early Intraoperative Acute Kidney Injury Prediction arXiv:2609.26848v1 Announce Type: new Abstract: Postoperative acute kidney injury (AKI) after major non-cardiac surgery carries substantial morbidity, yet early intraoperative risk stratification remains difficult. In this retrospective cohort study, we propose SynerT, a… 35 arXiv — Machine Learning research 4d ago COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation arXiv:2609.26853v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing… 5 arXiv — Machine Learning research 4d ago When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations arXiv:2609.26955v1 Announce Type: new Abstract: Fairness audits in production ML typically occur once, at deployment, on a single domain. Both fail in practice: fairness can shift after retraining or a changing user base, and interventions validated on one dataset are rarely… 7 arXiv — Machine Learning research 4d ago Discrete Diffusion Models via Evolving Variational Autoregressive Networks arXiv:2609.27306v1 Announce Type: new Abstract: Conventional score-based diffusion models learn scores without representing normalized densities, whereas tractable normalized models support both sampling and direct likelihood evaluation. A recent tensor-network approach provides… 5 arXiv — Machine Learning research 4d ago Limiting-Kernel Q($\lambda$): Bridging Short and Long Horizons arXiv:2609.27741v1 Announce Type: new Abstract: In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations relying on $n$-step truncation… 28 arXiv — Machine Learning research 4d ago CAST: Context- and Anomaly Structure-Conditioned Time Series Anomaly Generation arXiv:2609.27825v1 Announce Type: new Abstract: Anomalous time series play a critical role in safety-critical domains, yet they are inherently scarce, heterogeneous, and costly to obtain. Existing time series generation methods predominantly focus on synthesizing normal data,… 16 arXiv — Machine Learning research 4d ago Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel arXiv:2609.27867v1 Announce Type: new Abstract: A forecasting benchmark reports which method won. We show that the answer is set by the evaluator's choices before any model is fitted. We benchmark 24 forecasting methods and one textbook reference, including six 2025-era time… 26 arXiv — Machine Learning research 4d ago Global tree forecasters collapse at the hierarchical aggregate: a five-panel failure characterization arXiv:2609.27912v1 Announce Type: new Abstract: Global forecasting models pool many series and learn one shared function. Gradient-boosted trees are their most common form. We measure a failure of this design that has not, to our knowledge, been documented. Train a global tree… 4 arXiv — NLP / Computation & Language research 4d ago NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task arXiv:2609.27086v1 Announce Type: new Abstract: NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification (NADI) shared task series and the second dedicated to multidialectal Arabic speech processing. This edition comprises five tasks and eight subtasks… 23 arXiv — NLP / Computation & Language research 4d ago Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure arXiv:2609.27176v1 Announce Type: new Abstract: Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a… 28 arXiv — NLP / Computation & Language research 4d ago Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models arXiv:2609.27372v1 Announce Type: new Abstract: Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset,… 28 arXiv — NLP / Computation & Language research 4d ago Cross-Lingual Legal QA for Vietnamese Labour Law: Retrieval, Translation, and Verifier-Guided Correction arXiv:2609.27376v1 Announce Type: new Abstract: Cross-lingual legal question answering must retrieve statutes across languages while preventing unsupported legal claims. We introduce a bilingual evaluation suite of 231 Vietnamese--English question--answer pairs from Vietnamese… 36 arXiv — NLP / Computation & Language research 4d ago Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models arXiv:2609.27510v1 Announce Type: new Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is… 30 arXiv — NLP / Computation & Language research 4d ago MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors arXiv:2609.27590v1 Announce Type: new Abstract: Long-context evaluations often test whether a model can recover distant evidence, but recoverability does not guarantee behavioral influence. We test the prediction that a distant discourse anchor can remain explicitly recoverable… 21 arXiv — NLP / Computation & Language research 4d ago Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond arXiv:2609.27650v1 Announce Type: new Abstract: Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a… 25 arXiv — NLP / Computation & Language research 4d ago Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures arXiv:2609.27773v1 Announce Type: new Abstract: As Large Language Models (LLMs) move from conversational assistants to advanced agentic systems, guardrail failures can convert adversarial intents into harmful executions. However, most guardrail evaluation frameworks focus only… 29 Page 1 of 10 · 500 articles Older →