News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 4d ago From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015-2026 arXiv:2609.27939v1 Announce Type: new Abstract: Natural language processing (NLP) applied to open-ended teaching-evaluation comments (Student Evaluation of Teaching, SET) has tracked the field's technical evolution--from lexicons and conventional classifiers to transformers and… 34 arXiv — NLP / Computation & Language research 4d ago Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing arXiv:2609.28026v1 Announce Type: new Abstract: We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined… 18 arXiv — NLP / Computation & Language research 4d ago Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions arXiv:2609.27749v1 Announce Type: cross Abstract: The surge in AI-assisted generation of educational materials has outpaced our capacity to validate their pedagogical quality. Automated evaluation using Bloom Classifier models is a promising approach to assess educational… 13 r/LocalLLaMA community 4d ago Do y'all remember the snake model evaluation test? Just watched a video from three years ago where Matt Berman tested Bard (yes, remember Bard?) and most of the tests were like summarization or logic tests. If it ever came to coding, it was the snake test, or pong, and those from three years ago could barely get them right. Now… 27 Hacker News — AI on Front Page community 4d ago Linux support is coming to Snapdragon X2 Series Article URL: https://www.qualcomm.com/news/onq/2026/09/snapdragon-summit-agentic-ai-pcs-linux Comments URL: https://news.ycombinator.com/item?id=49823582 Points: 322 # Comments: 138 12 TechCrunch — AI news-outlet 4d ago Ema raises $77M as AI starts eating into enterprise software and services Ema has raised $140 million to date and has more than 50 enterprise customers, including Google and Microsoft. 35 arXiv — Machine Learning research 5d ago Exposing Blind Spots in Deep Imbalanced Regression Evaluation arXiv:2609.25152v1 Announce Type: new Abstract: Deep Imbalanced Regression (DIR) addresses a common failure mode of regression models: target distributions are highly non-uniform, causing models to perform best in densely populated target regions even when reliable performance… 23 arXiv — Machine Learning research 5d ago Multi-Term Fourier Graph Neural Network with Sample Relationship Learning for Enhanced Remaining Useful Life Prediction arXiv:2609.25179v1 Announce Type: new Abstract: Predicting the remaining useful life (RUL) is essential for effective predictive maintenance. Spatio-Temporal Graph Neural Networks (ST-GNNs), which can model both temporal and spatial relationships by representing time series data… 31 arXiv — Machine Learning research 5d ago Spatiotemporal Kronecker Covariance Neural Networks arXiv:2609.25326v1 Announce Type: new Abstract: Multivariate time series contain complex patterns that span across both space and time. While covariance-based statistical tools like spatiotemporal Principal Component Analysis (ST-PCA) help identify these patterns, they are… 25 arXiv — NLP / Computation & Language research 5d ago Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices arXiv:2609.25645v1 Announce Type: cross Abstract: Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval,… 32 arXiv — NLP / Computation & Language research 5d ago Auditing Proxy-Based Validation Across Text Spans arXiv:2609.25808v1 Announce Type: cross Abstract: Evaluation scores are often validated by their agreement with inexpensive proxy labels. When the score and the proxy are computed from the same text span, however, that agreement can arise from surface evidence the two share… 4 arXiv — Machine Learning research 5d ago Protocol before progress: leakage-aware evaluation of AIS trajectory prediction arXiv:2609.25827v1 Announce Type: new Abstract: Reported gains in vessel-trajectory prediction from Automatic Identification System (AIS) data are credited to new architectures, but the evaluation protocol is rarely measured as a source of error reduction. We build a… 30 arXiv — Machine Learning research 5d ago In-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization arXiv:2609.25836v1 Announce Type: new Abstract: Multi-task optimization (MTO) addresses a set of optimization tasks simultaneously, often suffering from inaccurate inter-task relationship estimation under limited evaluation budgets, leading to negative transfer. This paper… 26 arXiv — Machine Learning research 5d ago Margin-Drop Coordinates for Cross-Budget Robustness Evaluation arXiv:2609.26081v1 Announce Type: new Abstract: Fixed-budget robustness evaluation can select the wrong frozen vision encoder. An encoder that survives a shallow attack may lose most of that robustness when the same evaluation is strengthened. We ask whether the shallow… 21 arXiv — Machine Learning research 5d ago Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting arXiv:2609.26288v1 Announce Type: new Abstract: Comparative studies of predictive models often end by ranking candidate models, yet these rankings depend on evaluation protocols whose influence is rarely treated as a source of uncertainty. We formalize this problem as… 29 arXiv — NLP / Computation & Language research 5d ago Understanding Reliability in LLM-based Human Behavior Simulation arXiv:2609.25066v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end… 8 arXiv — NLP / Computation & Language research 5d ago FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing arXiv:2609.25298v1 Announce Type: new Abstract: Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally… 32 arXiv — NLP / Computation & Language research 5d ago Calibration as a First-Class Criterion in LLM Evaluation arXiv:2609.26489v1 Announce Type: new Abstract: Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this… 9 arXiv — NLP / Computation & Language research 5d ago Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation arXiv:2609.26693v1 Announce Type: new Abstract: A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that… 36 TechCrunch — AI news-outlet 5d ago Snorkel AI triples valuation to $3.5B as demand for AI training data booms The seven-year-old startup has raised a $350 million Series E to fuel its data-as-a-service approach. 33 arXiv — Machine Learning research 6d ago ZoAQ: Adaptive Zeroth-Order Querying via Query-Reuse Coupling arXiv:2609.22115v1 Announce Type: new Abstract: Zeroth-order optimization (ZOO) estimates updates from function evaluations, making perturbation queries a primary cost. Fixed budgets spend the same number of queries at every step, while adaptive controllers may offset their… 13 arXiv — Machine Learning research 6d ago Rank Portability Does Not Imply Feasibility Portability: Target-Specific Evaluation of Joint Hardware Constraints arXiv:2609.22122v1 Announce Type: new Abstract: Cross-device hardware evaluation often assumes that if architecture rankings transfer across devices, a proxy device can support target-side model selection. We stress-test this assumption for joint latency-energy feasibility… 5 arXiv — Machine Learning research 6d ago SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs arXiv:2609.22153v1 Announce Type: new Abstract: Methods for addressing safety drift in fine-tuned Large Language Models (LLMs) are scattered across incompatible implementations, lifecycle stages, and evaluation protocols, making them difficult to adopt and compare. We introduce… 29 arXiv — Machine Learning research 6d ago A Synthetic Multivariate Refrigerator Time-Series Dataset for Predictive Maintenance arXiv:2609.22229v1 Announce Type: new Abstract: We generated synthetic multivariate time series for 27 refrigerators with a simplified physicsinspired simulator at one-minute resolution. The simulator includes ambient-temperature variation, door use, thermostat and compressor… 20 arXiv — Machine Learning research 6d ago Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints arXiv:2609.22819v1 Announce Type: new Abstract: Counterfactual tool evaluation must distinguish authority, historical support, and what a comparison actually estimates. We study these distinctions with eleven executable enterprise-inspired tools, exact-propensity logs, and real… 30 arXiv — NLP / Computation & Language research 6d ago Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models arXiv:2609.22097v1 Announce Type: new Abstract: The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the… 38 arXiv — NLP / Computation & Language research 6d ago DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation arXiv:2609.22104v1 Announce Type: new Abstract: As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottleneck from idea generation to idea evaluation. Existing evaluators mainly rely on… 11 arXiv — NLP / Computation & Language research 6d ago Evaluation Awareness Shifts from Format to Context with Model Scale arXiv:2609.22119v1 Announce Type: new Abstract: Evaluation awareness poses an unprecedented threat to model evaluation, but the mechanisms by which models detect it remain unknown. This study focuses on determining this and identifying contrasting mechanisms between smaller and… 23 arXiv — NLP / Computation & Language research 6d ago Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation arXiv:2609.22127v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to draft contractual language, yet conventional accuracy or preference-based evaluations are poorly matched to legal drafting. A clause may be fluent and stylistically polished… 15 arXiv — NLP / Computation & Language research 6d ago Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization arXiv:2609.22149v1 Announce Type: new Abstract: Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after tabular--text pairings are disrupted. We present a projection-based evaluator for… 20 arXiv — NLP / Computation & Language research 6d ago Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models arXiv:2609.22152v1 Announce Type: new Abstract: Imagination performs as a high-level function of large language models (LLMs) which determines the potential of how an LLM creates unseen or creative content. While existing works have built a rich family of creativity benchmarks… 14 arXiv — NLP / Computation & Language research 6d ago EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems arXiv:2609.22231v1 Announce Type: new Abstract: Long-horizon interactions with LLM-based assistants require memory systems that preserve and update user states, preferences, and interaction histories. Existing evaluations report end-to-end QA accuracy and cannot determine… 13 arXiv — NLP / Computation & Language research 6d ago Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations arXiv:2609.22255v1 Announce Type: new Abstract: Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a… 30 arXiv — NLP / Computation & Language research 6d ago Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation arXiv:2609.22603v1 Announce Type: new Abstract: Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while… 27 arXiv — NLP / Computation & Language research 6d ago Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments arXiv:2609.22705v1 Announce Type: new Abstract: Online discourse about urban issues - walkability, cycling infrastructure, public transit, housing density, and street safety - is voluminous but unstructured, and existing city-evaluation tools capture none of it. We present a… 24 arXiv — NLP / Computation & Language research 6d ago Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation arXiv:2609.22793v1 Announce Type: new Abstract: LLM-based machine translation evaluation can closely match human judgments, but in practice it remains largely diagnostic, with the signals rarely translating into direct quality improvements under real production constraints. We… 7 arXiv — NLP / Computation & Language research 6d ago LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage arXiv:2609.22904v1 Announce Type: new Abstract: Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to… 13 OpenAI official-blog 7d ago Building standards for the next phase of AI OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety. 37 arXiv — Machine Learning research 7d ago When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation arXiv:2609.20942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer… 27 arXiv — Machine Learning research 7d ago Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization arXiv:2609.21197v1 Announce Type: new Abstract: Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker… 16 arXiv — NLP / Computation & Language research 7d ago FairLMs: A Turnkey Library for Fairness in Language Models arXiv:2609.21296v1 Announce Type: cross Abstract: Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing tools offer complementary functionality through different interfaces, so… 4 arXiv — Machine Learning research 7d ago Efficient Architecture Search under Leave-One-Subject-Out Evaluation arXiv:2609.21457v1 Announce Type: new Abstract: Deep neural architectures are widely used for signal processing in automated pain assessment systems. However, architecture design has remained largely a manual task despite the potential efficiency benefits of Neural Architecture… 9 arXiv — NLP / Computation & Language research 7d ago HERMES: Contrast-Aware Knowledge Graph Reasoning from Clinical Notes for Patient Outcome Prediction arXiv:2609.20825v1 Announce Type: new Abstract: Clinical predictive models often rely on structured Electronic Health Record data, such as time-series and procedure codes. While recent approaches have begun leveraging unstructured clinical notes, they typically encode them as… 20 arXiv — NLP / Computation & Language research 7d ago From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators arXiv:2609.20827v1 Announce Type: new Abstract: Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not… 17 arXiv — NLP / Computation & Language research 7d ago SAGE: Schema-Guided LLMs for Grant Review arXiv:2609.20829v1 Announce Type: new Abstract: Grant reviewers must apply detailed criteria to application forms, budgets, and supporting documents while producing assessments that colleagues can inspect. We present SAGE, Schema-Guided Aspect-Based Grant Evaluation, a system… 12 arXiv — NLP / Computation & Language research 7d ago TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar arXiv:2609.20832v1 Announce Type: new Abstract: We introduce TatBLiMP, the first benchmark of linguistic minimal pairs for Tatar (tt, ISO 639-3 tat), a Qypchaq Turkic language written in Cyrillic. To our knowledge it is the first grammaticality evaluation for Tatar language… 28 arXiv — NLP / Computation & Language research 7d ago MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs arXiv:2609.20850v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related… 29 arXiv — NLP / Computation & Language research 7d ago Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction arXiv:2609.21231v1 Announce Type: new Abstract: Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving… 17 arXiv — NLP / Computation & Language research 7d ago Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations arXiv:2609.21490v1 Announce Type: new Abstract: Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that… 19 arXiv — NLP / Computation & Language research 7d ago Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER arXiv:2609.21663v1 Announce Type: new Abstract: Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question:… 11 Page 2 of 10 · 500 articles ← Newer Older →