News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow Hugging Face Daily Papers research 4d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination Abstract Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability,… 14 Hugging Face Daily Papers research 4d ago Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events Abstract Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is… 7 Hugging Face Daily Papers research 4d ago PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say Abstract LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Existing privacy benchmarks audit what the agent's response or outgoing… 9 arXiv — Machine Learning research 4d ago Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks arXiv:2608.07335v1 Announce Type: new Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q-Network (PQN) algorithm achieves stable off-policy learning without relying on… 9 arXiv — Machine Learning research 4d ago UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys arXiv:2608.06404v1 Announce Type: cross Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic… 36 arXiv — NLP / Computation & Language research 4d ago Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events arXiv:2608.06485v1 Announce Type: new Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key… 21 arXiv — NLP / Computation & Language research 4d ago Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand arXiv:2608.06506v1 Announce Type: new Abstract: Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while… 16 arXiv — NLP / Computation & Language research 4d ago TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade arXiv:2608.06549v1 Announce Type: new Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where… 11 arXiv — NLP / Computation & Language research 4d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination arXiv:2608.07341v1 Announce Type: new Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and… 7 arXiv — NLP / Computation & Language research 4d ago LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering arXiv:2608.07370v1 Announce Type: new Abstract: Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent… 15 arXiv — NLP / Computation & Language research 4d ago StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a… 5 arXiv — NLP / Computation & Language research 4d ago How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social… 14 arXiv — NLP / Computation & Language research 4d ago GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks arXiv:2608.07411v1 Announce Type: cross Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a… 36 arXiv — NLP / Computation & Language research 4d ago SABRE: Scalable and Automated Benchmarking of VLMs under Stress arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and… 22 arXiv — NLP / Computation & Language research 4d ago Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages arXiv:2605.02608v2 Announce Type: replace Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers---the… 32 r/LocalLLaMA community 4d ago DeepSeek v4 Flash 0731 locally on CPU After seeing the benchmark results for the full release of DS v4 Flash 0731, I replaced my 2 x 16GB DDR4 ram sticks with 2 x 32GB DDR4 ram sticks to get a max supported of 128 GB RAM, in hope to be able to run GLM 5.2 equivalent model locally i.e. DS v4 Flash 0731 I also have… 25 r/LocalLLaMA community 5d ago Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local) Howdy - I posted a benchmark here - https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/ This was using the hosted API - since then I've been playing around with quants Here is the lastest benchmark -… 20 r/LocalLLaMA community 5d ago any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on? I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 million tokens?   submitted by   /u/nomorebuttsplz [link]  … 35 r/LocalLLaMA community 6d ago DeepSeek V4 Flash 0731 appreciation post I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinely amazed at what it can handle. I can throw a two-hour coding session at it,… 12 r/LocalLLaMA community 6d ago Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks? (I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence benchmarks. And those flaws render it useless unfortunetely for anything else… 21 r/LocalLLaMA community 6d ago DeepSeek V4 Flash 0731 - ARC-AGI Results   submitted by   /u/johnnyApplePRNG [link]   [comments] 20 r/LocalLLaMA community 6d ago LFM2.5-2.6B model+KV cache quantization report LFM2.5-2.6B is a new tiny model by LiquidAI, with benchmarks that put it head to head with much larger models. I've run llama-perplexity on many model GGUF quants, crossed with many KV cache quants, to understand the model's best overall quantization for any given amount of… 32 r/LocalLLaMA community 6d ago A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s I was going through the current llama.cpp CPU PRs and #26348 stood out because this isn't the usual +5% kernel optimization. It adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, and the author's controlled CPU-only benchmarks show roughly 3–3.6x higher throughput… 24 Hugging Face Daily Papers research 6d ago DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces Abstract Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended… 31 r/LocalLLaMA community 6d ago LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks. LabyrinthBench measures the thing that actually kills long agent runs — whether a model can still use what it learned twenty turns ago — deterministically, with no LLM judge, on your own hardware, with a swappable harness for testing whatever context-management strategy you… 36 Hugging Face Daily Papers research 6d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Abstract We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish… 5 r/LocalLLaMA community 6d ago Gemma 4 QAT could be improved further by Google aligning the QAT model to modern q4_k instead of q4_0 Hello, For the past few days I have been benchmarking Gemma 4 26b QAT UD Q4_K_XL extensively versus Bartowski's Q4_K_L. While QAT is certainly very effective and reducing memory consumption versus the highest q4 quant from him, I also have noticed some regressions in my own… 4 Hugging Face Daily Papers research 7d ago GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Abstract Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce… 20 Hugging Face Daily Papers research 7d ago Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains Abstract Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation… 5 arXiv — Machine Learning research 7d ago MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification arXiv:2608.05196v1 Announce Type: new Abstract: Multiple sclerosis (MS) is diagnosed through clinical assessment, magnetic resonance imaging, laboratory evidence when appropriate, and exclusion of better explanations. Blood RNA expression data may contain disease associated… 14 arXiv — Machine Learning research 7d ago SEAM: Global consistency beyond local accuracy in scientific machine learning arXiv:2608.05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction. Yet such local checks cannot establish whether the resulting explanations can be assembled… 12 arXiv — NLP / Computation & Language research 7d ago GROM: Gradient-Free Rapid One-Shot Machine Unlearning arXiv:2608.05783v1 Announce Type: cross Abstract: Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state-of-the-art approaches primarily rely on iterative, training-time unlearning… 27 arXiv — Machine Learning research 7d ago Is Self-Pretraining really useful to improve diagnosis in medical Time Series? arXiv:2608.06122v1 Announce Type: new Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical… 23 arXiv — NLP / Computation & Language research 7d ago PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing… 24 arXiv — NLP / Computation & Language research 7d ago Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles… 38 arXiv — Machine Learning research 7d ago Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks arXiv:2608.05266v1 Announce Type: cross Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure… 38 arXiv — NLP / Computation & Language research 7d ago M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring… 25 arXiv — NLP / Computation & Language research 7d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark arXiv:2608.05850v1 Announce Type: new Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have… 23 arXiv — NLP / Computation & Language research 7d ago Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents arXiv:2608.06312v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards,… 17 arXiv — NLP / Computation & Language research 7d ago Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents arXiv:2608.06329v1 Announce Type: new Abstract: Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited… 24 arXiv — NLP / Computation & Language research 7d ago EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents arXiv:2608.05519v1 Announce Type: cross Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human… 7 arXiv — NLP / Computation & Language research 7d ago From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a… 33 r/LocalLLaMA community 7d ago KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0 , fork of llama.cpp with more KV cache quantization options. Models: Qwen 3.6 27B Q5_K_S 64k context, Gemma 4 31B Q5_K_S 16k context Standard quants, extended:… 9 Hugging Face Daily Papers research 7d ago SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models Abstract Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a… 14 r/LocalLLaMA community 7d ago How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a benchmarking issue? EDIT: The contribution of this benchmark to the… 28 r/MachineLearning community 7d ago The current state of language models and human preference based rankings [R] "Arena ai" has been a great success in producing a human preference based ranking, additional to other more objective benchmarks. However, this (probably) had also played a role in the syncopancy crisis and the general tendency of some models to tilt towards overformatting to… 27 Hugging Face Daily Papers research 8d ago GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Abstract Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not… 17 Hugging Face Daily Papers research 8d ago AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Abstract While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on… 15 Hugging Face Daily Papers research 8d ago Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Abstract Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming… 25 arXiv — Machine Learning research 8d ago Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification arXiv:2608.04702v1 Announce Type: new Abstract: Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize beyond the evaluated datasets. In this work, we move beyond rankings by… 29 Page 3 of 10 · 500 articles ← Newer Older →