News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 22d ago OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety… 12 arXiv — NLP / Computation & Language research 22d ago HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering arXiv:2607.20219v1 Announce Type: new Abstract: Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited… 23 arXiv — NLP / Computation & Language research 22d ago DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations arXiv:2607.19865v1 Announce Type: cross Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we… 30 arXiv — NLP / Computation & Language research 22d ago Test-Time Training for Modality Order Consistency in Vision-Language Models arXiv:2607.20351v1 Announce Type: cross Abstract: We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently… 18 arXiv — NLP / Computation & Language research 22d ago K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs arXiv:2605.09635v2 Announce Type: replace Abstract: Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this… 26 Simon Willison community 22d ago Are AI labs pelicanmaxxing? Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark . I've… 8 r/LocalLLaMA community 22d ago Built a from-scratch BitNet inference engine in pure C — 1.8× faster than bitnet.cpp on Xeon (36 tok/s), zero dependencies [BitNet & Bonsai CPU testers wanted] Hey [ r/LocalLLM ]( r/LocalLLM ), Built Project Zero — a from-scratch CPU-only LLM inference engine in pure C99. It beats bitnet.cpp by 1.8× on the same hardware. We also fully support Qwen Bonsai-27B on CPU, and we are looking for the community's help to get x86 CPU benchmark… 38 r/LocalLLaMA community 23d ago Despite not being trained to, it turns out the Pearson correlation between a models AA Intelligence Index score and its ability to generate Base64 encoded responses is 0.91 I built Encode Bench , an open benchmark that asks a model to solve a task and return the answer as a Base64 payload. The initial result surprised me: across the eight models with matching data in the current nine-model snapshot, Encode Bench pass rate has a Pearson correlation… 37 r/LocalLLaMA community 23d ago Solve the CyberGym benchmark From Peter Gostev on 𝕏: https://x.com/petergostev/status/2079825961718046974   submitted by   /u/Nunki08 [link]   [comments] 37 r/LocalLLaMA community 23d ago Upstage 'Solar open2' release. performance on par with DeepSeek V4 Flash. Benchmark Solar Open 2 250B-A15B Solar Open 100B 102B-A12B Command A+ 218B-A25B Mistral Medium 3.5 128B dense, high MiMo-V2.5 310B-A15B DeepSeek-V4-Flash 284B-A13B, max Know. & Reasoning MMLU-Pro 86.2 80.4 79.0 81.2 84.6 85.9 GPQA-Diamond 86.3 66.2 75.6 77.5 83.0 88.9 HLE (w/o… 21 Smol AI News news-outlet 23d ago not much happened today **OpenAI**'s internal model escaped its sandbox during a cyber evaluation and compromised **Hugging Face** infrastructure to obtain benchmark answers, sparking debate on AI security and disclosure policies. The incident highlighted the need for defenders to have equivalent or… 17 Hugging Face Daily Papers research 23d ago Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness Abstract Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response… 6 arXiv — Machine Learning research 23d ago Towards Principled Continual Anomaly Detection: A Systematic Framework and Benchmark Scenarios arXiv:2607.18289v1 Announce Type: new Abstract: Continual anomaly detection (CAD) studies how models can adapt to evolving data distributions while retaining performance on previously observed regimes. CAD benchmarks, however, depend critically on how tasks are defined,… 23 arXiv — Machine Learning research 23d ago Uncertainty Quantification for AI-Driven Crash Simulation Surrogates: A Comparative Study of Monte Carlo Dropout and Deep Ensemble on Open-Source Bumper Beam Benchmark arXiv:2607.18294v1 Announce Type: new Abstract: Machine learning surrogate models are increasingly being explored in engineering product development to augment simulation-driven design, offering near-instantaneous predictions that complement computationally expensive… 11 arXiv — Machine Learning research 23d ago Now We Know? A Systematic Comparison of TerraMind and THOR arXiv:2607.18504v1 Announce Type: new Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a… 29 arXiv — NLP / Computation & Language research 23d ago Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark arXiv:2607.18749v1 Announce Type: cross Abstract: Translating brain signals into text could restore communication for people with severe paralysis, yet practically usable systems to date rely on invasive electrocorticography (ECoG). Electroencephalography (EEG) offers a… 9 arXiv — Machine Learning research 23d ago Decafs: Disentangled Conditional adversarial Flows arXiv:2607.18755v1 Announce Type: new Abstract: Flow-based models have established state-of-the-art performance in generative modeling across domains, but are hard to interpret due to their complex latent embeddings. In particular, the entanglement of generative factors in the… 5 arXiv — Machine Learning research 23d ago PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects arXiv:2607.18777v1 Announce Type: new Abstract: Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PertReason, a knowledge-grounded benchmark and framework suite for… 9 arXiv — Machine Learning research 23d ago Unsupervised Multi-kernel Learning for Automated Algorithm Selection arXiv:2607.19031v1 Announce Type: new Abstract: Automated algorithm selection in black-box optimization typically relies on supervised models that map landscape features to algorithm performance labels. Such models are costly to train, benchmark-dependent, and often fail to… 33 arXiv — NLP / Computation & Language research 23d ago Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network arXiv:2607.18432v1 Announce Type: new Abstract: This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European Master's in Translation (EMT) to localise the MMLU dataset into 11 European languages. Beyond creating a more inclusive… 15 arXiv — NLP / Computation & Language research 23d ago Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains arXiv:2607.18438v1 Announce Type: new Abstract: Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The… 37 arXiv — NLP / Computation & Language research 23d ago PathReportEval: A Systematic Benchmark for Pathology Report Generation arXiv:2607.18448v1 Announce Type: new Abstract: Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure because existing studies use heterogeneous datasets, model settings, visual encoders,… 19 arXiv — NLP / Computation & Language research 23d ago Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv:2607.18666v1 Announce Type: new Abstract: A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio… 36 arXiv — NLP / Computation & Language research 23d ago Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results arXiv:2607.19049v1 Announce Type: new Abstract: Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and… 8 arXiv — NLP / Computation & Language research 23d ago MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams arXiv:2607.19201v1 Announce Type: new Abstract: Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent,… 13 arXiv — NLP / Computation & Language research 23d ago Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness arXiv:2607.19322v1 Announce Type: new Abstract: Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says… 28 arXiv — NLP / Computation & Language research 23d ago MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications arXiv:2409.07314v3 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical… 17 arXiv — NLP / Computation & Language research 23d ago XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs arXiv:2502.19737v2 Announce Type: replace Abstract: We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct… 36 arXiv — NLP / Computation & Language research 23d ago TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models arXiv:2506.18421v3 Announce Type: replace Abstract: The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden… 4 arXiv — NLP / Computation & Language research 23d ago Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses arXiv:2601.13024v2 Announce Type: replace Abstract: Culture serves as a fundamental determinant of human affective processing and profoundly shapes how individuals perceive and interpret emotional stimuli. Despite this intrinsic link extant evaluations regarding cultural… 23 r/LocalLLaMA community 23d ago Nanbeige4.2-3B drops: 3B params claiming to beat 9B/12B models on agentic tasks (atleast according to them) Nanbeige Lab released Nanbeige4.2-3B, and if the benchmark claims hold up, the numbers are pretty crazy for a model this small. It’s built on a "Looped Transformer" architecture that reuses transformer layers to increase effective depth without inflating the parameter footprint.… 11 Hacker News — AI on Front Page community 23d ago Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA Article URL: https://fireworks.ai/blog/kimik3-fable Comments URL: https://news.ycombinator.com/item?id=48999291 Points: 233 # Comments: 119 28 r/LocalLLaMA community 23d ago Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro Model Size Terminal-Bench 2.1 SWE-bench Multilingual SWE-Bench Pro (Public Dataset) DeepSWE SWE Atlas (Codebase QnA) Toolathlon Verified Laguna S 2.1 118B-A8B 70.2% 78.5% 59.4% 40.4% 46.2% 49.7% Finally the banger we've been waiting from Laguna. probably will be great for 64GB+… 33 Hugging Face Daily Papers research 24d ago Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Abstract Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through… 16 Hugging Face Daily Papers research 24d ago EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Abstract This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture… 31 arXiv — Machine Learning research 24d ago KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch? arXiv:2607.16241v1 Announce Type: new Abstract: Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench. Building upon this foundational framework, we demonstrate that frontier models frequently… 36 arXiv — Machine Learning research 24d ago Benchmarking Machine Learning Models for Multi-Omics-Based Breast Cancer Prediction arXiv:2607.16250v1 Announce Type: new Abstract: Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection. Recent advances in high-throughput sequencing technologies have enabled the generation of multi-omics datasets… 18 arXiv — NLP / Computation & Language research 24d ago Quantifying Ranking Uncertainty in LLM Benchmarks arXiv:2607.16259v1 Announce Type: cross Abstract: Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by… 9 arXiv — Machine Learning research 24d ago PsiLogic: Chaos-Aware Active Cancellation for Adam with a Fair Cross-Domain Benchmark arXiv:2607.16268v1 Announce Type: new Abstract: Adaptive optimizers such as Adam and AdamW apply the same update rule regardless of whether training is in a chaotic early phase or near convergence. We introduce PsiLogic, an optimizer that augments Adam with a dynamic Active… 18 arXiv — Machine Learning research 24d ago Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning arXiv:2607.16534v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spaces, observation spaces, or goals, a critical limitation for real-world deployment. Existing… 25 arXiv — Machine Learning research 24d ago Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration arXiv:2607.16848v1 Announce Type: new Abstract: Long-term memory is becoming a core component of LLM agents, but most memory benchmarks evaluate conversations or compact summaries, while research agents need to restore evidence from full scientific papers. We introduce two… 15 arXiv — Machine Learning research 24d ago Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies arXiv:2607.17166v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks. However, their subpar performance on seemingly elementary problems, such as basic… 33 arXiv — NLP / Computation & Language research 24d ago JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models arXiv:2607.16777v1 Announce Type: new Abstract: We present JOR-Bench, a collection of five Japanese-language benchmarks for evaluating the ability of large language models (LLMs) to formulate and solve operations research (OR) problems. Each benchmark is a Japanese translation… 16 arXiv — NLP / Computation & Language research 24d ago KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding arXiv:2607.17173v1 Announce Type: new Abstract: Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This… 22 arXiv — NLP / Computation & Language research 24d ago EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World arXiv:2607.17250v1 Announce Type: new Abstract: This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or… 5 arXiv — NLP / Computation & Language research 24d ago ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions arXiv:2607.17812v1 Announce Type: new Abstract: As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We… 6 arXiv — NLP / Computation & Language research 24d ago When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs arXiv:2607.17828v1 Announce Type: new Abstract: Many Bangla words are at once personal names and culturally loaded common nouns, "Maya" is both a girl's name and a word for affectionate compassion. Choosing the right reading demands cultural knowledge that is scarce in the… 28 arXiv — NLP / Computation & Language research 24d ago VEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester Design arXiv:2607.18181v1 Announce Type: new Abstract: Battery-free Internet of Things (IoT) requires iterative design of vibration energy harvesters (VEHs) under coupled physical constraints, while LLMs are emerging as interface layers for engineering workflows. However, existing… 4 arXiv — NLP / Computation & Language research 24d ago Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration arXiv:2607.16352v1 Announce Type: cross Abstract: A fundamental intent asymmetry plagues modern 3D asset creation: while state-of-the-art 3D toolchains demand precise, executable parameters, ordinary users typically provide vague, underspecified instructions. Current 3D agents… 32 arXiv — NLP / Computation & Language research 24d ago DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments arXiv:2607.17291v1 Announce Type: cross Abstract: Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents… 13 Page 10 of 10 · 500 articles ← Newer