News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — Machine Learning research 15d ago Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark arXiv:2607.26993v1 Announce Type: new Abstract: Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pretrained… 22 arXiv — Machine Learning research 15d ago BayesAME: Bayesian Active Model Evaluation arXiv:2607.27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items,… 38 arXiv — Machine Learning research 15d ago Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark arXiv:2607.27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal… 15 arXiv — Machine Learning research 15d ago Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes arXiv:2607.27188v1 Announce Type: new Abstract: Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation,… 30 arXiv — Machine Learning research 15d ago When benchmark inferences do not compose: Projectibility in AI evaluation arXiv:2607.26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and… 34 arXiv — NLP / Computation & Language research 15d ago Position: Evaluation Scores Are Perishable Knowledge Claims arXiv:2607.26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via… 10 arXiv — NLP / Computation & Language research 15d ago When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses arXiv:2607.26348v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails,… 23 arXiv — NLP / Computation & Language research 15d ago Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification arXiv:2607.26397v1 Announce Type: new Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC… 38 arXiv — NLP / Computation & Language research 15d ago ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models arXiv:2607.26455v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing… 37 arXiv — NLP / Computation & Language research 15d ago Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms arXiv:2607.26497v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost… 35 arXiv — NLP / Computation & Language research 15d ago Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text arXiv:2607.26751v1 Announce Type: new Abstract: State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent… 33 arXiv — NLP / Computation & Language research 15d ago Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? arXiv:2607.26952v1 Announce Type: new Abstract: We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers… 33 arXiv — NLP / Computation & Language research 15d ago DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search arXiv:2607.27178v1 Announce Type: new Abstract: State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to… 21 arXiv — NLP / Computation & Language research 15d ago APEX-Accounting arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions,… 30 arXiv — NLP / Computation & Language research 15d ago SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.… 32 arXiv — NLP / Computation & Language research 15d ago Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data arXiv:2607.27056v1 Announce Type: cross Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also… 24 arXiv — NLP / Computation & Language research 15d ago OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding arXiv:2607.27155v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a… 15 arXiv — NLP / Computation & Language research 15d ago SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch arXiv:2607.27167v1 Announce Type: cross Abstract: LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap:… 35 arXiv — NLP / Computation & Language research 15d ago ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models. Despite growing global AI adoption, no systematic evaluation exists for ML… 30 arXiv — NLP / Computation & Language research 15d ago Ensembling LLM-Induced Decision Trees for Explainable and Robust Error Detection arXiv:2512.07246v3 Announce Type: replace Abstract: Error detection (ED), which aims to identify incorrect or inconsistent cell values in tabular data, is important for ensuring data quality. Recent state-of-the-art ED methods leverage the pre-trained knowledge and semantic… 21 arXiv — NLP / Computation & Language research 15d ago $\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials arXiv:2601.06300v2 Announce Type: replace Abstract: Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \textit{eligibility criteria amendment… 15 arXiv — NLP / Computation & Language research 15d ago CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators arXiv:2605.08334v2 Announce Type: replace Abstract: We present CustomerSim, an environment and benchmark to evaluate the extent to which Multimodal Large Language Models (MLLMs) can simulate realistic, persona-driven customer behavior in chat-based retail environments. While… 27 r/MachineLearning community 15d ago AI Security Leaderboard: benchmarking model robustness [P] We developed a leaderboard ranking frontier model security. There's no shortage of model capability rankings, but we didn't find anything comparable for model security. Yet security is becoming increasingly critical to deployment decisions: from the USG making developers pull… 19 r/LocalLLaMA community 15d ago Budget Inference: A GPU for dense models vs. More RAM for MoE models? Hi all, I’m building a budget inference machine primarily for personal use (chat/assistant tasks, possibly some RAG). I'm torn between two hardware paths and would love input from anyone who has actually benchmarked these setups. The Dilemma: Option A (GPU for dense models): Buy… 35 Ars Technica — AI news-outlet 15d ago Elon Musk’s xAI is trying to sue its way out of a Grok reckoning Musk defends Grok, says Minnesota's nudifying app ban is unconstitutional. 16 OpenAI official-blog 15d ago How enabling two settings tripled our scores on the ARC-AGI-3 benchmark How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction. 6 Hugging Face Daily Papers research 16d ago GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels Abstract Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definition focuses on tumor subregions and does not systematically represent coexisting white matter hyperintensities (WMH). In joint segmentation settings,… 38 Hugging Face Daily Papers research 16d ago PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Abstract We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with… 34 Hugging Face Daily Papers research 16d ago Parallel Decoding Distillation for Fast Image and Video Generation Abstract Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill… 31 r/LocalLLaMA community 16d ago Is Laguna s2.1 fixed? Laguna s2.1 launched about a week ago , the benchmark's that they advertised were crazy good . But It was a mess , looping issues, tool ussage problems, not performing near the advertised benchmark. So I was gonna ask kindly if anyone used it with the updates that they gave, and… 24 Hugging Face Daily Papers research 16d ago Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking Abstract Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without… 28 arXiv — NLP / Computation & Language research 16d ago MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To… 33 arXiv — NLP / Computation & Language research 16d ago Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recognized languages. Large language models (LLMs) are now being deployed on a massive… 20 arXiv — NLP / Computation & Language research 16d ago WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing arXiv:2607.25765v1 Announce Type: new Abstract: Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs for file relationships. Existing benchmarks typically evaluate retrieval or… 6 arXiv — NLP / Computation & Language research 16d ago Shieldstral arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety… 19 arXiv — NLP / Computation & Language research 16d ago Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections arXiv:2607.25953v1 Announce Type: new Abstract: As LLMs increasingly mediate the political information citizens rely on, there is still no standardized way to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded benchmark for evaluating LLMs as… 17 arXiv — NLP / Computation & Language research 16d ago Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines arXiv:2607.25356v1 Announce Type: cross Abstract: Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundational for data-centric AI pipelines, yet exhaustive scans over millions of rows… 36 arXiv — NLP / Computation & Language research 16d ago HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following arXiv:2607.25398v1 Announce Type: cross Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing… 6 arXiv — NLP / Computation & Language research 16d ago PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents arXiv:2607.25485v1 Announce Type: cross Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting… 7 arXiv — NLP / Computation & Language research 16d ago Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than… 34 arXiv — NLP / Computation & Language research 16d ago RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement arXiv:2607.25886v1 Announce Type: cross Abstract: Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and validating training-data strategies, and learning… 15 Hugging Face Daily Papers research 16d ago Shieldstral Abstract We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content… 14 Hugging Face Daily Papers research 16d ago Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Abstract Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge… 31 Hugging Face Daily Papers research 16d ago Towards Robust Reinforcement Learning for Small-Scale Language Model Agents Abstract The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen… 34 Hugging Face Daily Papers research 16d ago FilmBench: A Film-Grade Benchmark for Cinematic Video Generation Abstract Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally,… 17 r/LocalLLaMA community 16d ago Gemma 4 26B/31B Q4 QAT vs Q4/Q5/Q6/Q8 What are your experiences with Gemma 4's QAT versions compared to their regular ones? So far I have mostly heard about regressions, but if you have a different experience or even benchmarks that are in favor of QAT, this is the thread to share them. Please share your positive or… 14 r/LocalLLaMA community 16d ago [PAPER] GPQA, MMLU-Pro, and MMMU-Pro were audited for broken questions, and up to 12% of them had to be removed. New drop in clean versions released I was very curious why all the models were topping out on GPQA-Diamond around 92 or 93% ( AA ) and spent the last few weeks pouring over GPQA (Diamond and Extended), and then expanded to auditing MMLU-Pro and MMMU-Pro. It was quite frankly shocking just how many questions were… 9 r/MachineLearning community 16d ago Might need math+code benchmark for frontier model(LLMs Silently Replace Math)[D] Hello guys. I found some problems in current frontier models. And want to share. # math_code_hallucination > Record of a failure caused by combining mathematics and code in a single prompt. --- ## Case 1 ### Initial prompt ( `p0` ) If you enter the following prompt: ```python… 35 r/LocalLLaMA community 16d ago SWE-rebench Multilingual Update (Go, Java, Python, Rust, TS). Evaluated: GLM-5.2, DeepSeek-V4 Pro, Qwen3.6-27B and others Hi everyone! We’ve just released a major update to the leaderboard! We are expanding beyond Python with a new multilingual slice featuring real-world software engineering tasks across 5 languages . Open-weight models: Model Pass@1 Pass@5 Pass all 5 GLM-5.2 [high] 62,9% (± 1.19%)… 33 Hugging Face Daily Papers research 17d ago Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels Abstract Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the… 13 Page 7 of 10 · 500 articles ← Newer Older →