News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — Machine Learning research 3h ago When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide arXiv:2608.12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k… 18 arXiv — Machine Learning research 3h ago CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility arXiv:2608.12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the… 38 arXiv — Machine Learning research 3h ago Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity arXiv:2608.13197v1 Announce Type: new Abstract: Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention. However, real-world falls are extremely rare: collecting 100 of them requires… 10 arXiv — Machine Learning research 3h ago Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks arXiv:2608.13296v1 Announce Type: new Abstract: Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back even to the 1970s. This causes a risk of biasing the development of global optimization… 24 arXiv — NLP / Computation & Language research 3h ago Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition arXiv:2608.12327v1 Announce Type: new Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium,… 34 arXiv — NLP / Computation & Language research 3h ago Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents arXiv:2608.12342v1 Announce Type: new Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial… 7 arXiv — NLP / Computation & Language research 3h ago Vision-Language Models are Fragile Multilingual Associators arXiv:2608.12333v1 Announce Type: new Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark… 37 arXiv — NLP / Computation & Language research 3h ago Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification arXiv:2608.12340v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), generative data augmentation has attracted considerable attention for imbalanced text classification in natural language processing. However, no empirical benchmark to… 7 arXiv — NLP / Computation & Language research 3h ago The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models arXiv:2608.12341v1 Announce Type: new Abstract: Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or… 11 arXiv — NLP / Computation & Language research 3h ago Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark arXiv:2608.12343v1 Announce Type: new Abstract: AI chatbots are widely used by students as knowledge sources, yet LLM benchmarks rarely assess interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exams (Matura) in history - three… 12 arXiv — NLP / Computation & Language research 3h ago Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models arXiv:2608.12391v1 Announce Type: new Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input… 36 arXiv — NLP / Computation & Language research 3h ago Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection arXiv:2608.12652v1 Announce Type: new Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or… 7 arXiv — NLP / Computation & Language research 3h ago BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian arXiv:2608.12894v1 Announce Type: new Abstract: Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian… 35 arXiv — NLP / Computation & Language research 3h ago LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation arXiv:2608.13136v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas.… 32 arXiv — NLP / Computation & Language research 3h ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 3h ago Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability… 20 arXiv — NLP / Computation & Language research 3h ago When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models arXiv:2608.12324v1 Announce Type: cross Abstract: People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about… 37 arXiv — NLP / Computation & Language research 3h ago Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists arXiv:2608.12345v1 Announce Type: cross Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct… 9 arXiv — NLP / Computation & Language research 3h ago SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries arXiv:2608.12654v1 Announce Type: cross Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy… 38 Hugging Face Daily Papers research 19h ago AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research Abstract The benchmark evaluates autonomous coding agents on open-ended world-model research by having them iteratively improve a starter model across game environments using a shared structured-state format. Generated by thinkingmachines/Inkling-Small World modeling is an… 20 Smol AI News news-outlet 1d ago not much happened today **Google** rapidly released **Gemini 3.7 Flash** just three weeks after 3.6 Flash, targeting coding, web development, knowledge work, and agentic workflows with a 50% introductory price cut and improved benchmark scores like **DeepSWE 65.3%** and **Code Arena Elo 1588**. The… 17 arXiv — Machine Learning research 1d ago Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark arXiv:2608.11423v1 Announce Type: new Abstract: Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across… 5 arXiv — NLP / Computation & Language research 1d ago Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs arXiv:2608.11232v1 Announce Type: new Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a… 18 arXiv — NLP / Computation & Language research 1d ago CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models arXiv:2608.11534v1 Announce Type: new Abstract: In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that… 10 arXiv — NLP / Computation & Language research 1d ago The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv:2608.11694v1 Announce Type: new Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while… 11 arXiv — NLP / Computation & Language research 1d ago Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory… 15 arXiv — NLP / Computation & Language research 1d ago LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence arXiv:2608.11922v1 Announce Type: new Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token… 8 arXiv — NLP / Computation & Language research 1d ago Accuracy and Order Sensitivity Diverge Under Label-Free Strategies arXiv:2608.11947v1 Announce Type: new Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test… 35 arXiv — NLP / Computation & Language research 1d ago Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existing approaches to building SLMs typically… 34 arXiv — NLP / Computation & Language research 1d ago A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench arXiv:2608.12138v1 Announce Type: new Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed… 10 arXiv — NLP / Computation & Language research 1d ago When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs arXiv:2608.11403v1 Announce Type: cross Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer. On the full GPQA Diamond benchmark (198 graduate-level science questions),… 30 arXiv — NLP / Computation & Language research 1d ago Benchmarking LLM Judges for Mobile Agent Evaluation arXiv:2608.11434v1 Announce Type: cross Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark… 17 arXiv — NLP / Computation & Language research 1d ago FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents arXiv:2608.11683v1 Announce Type: cross Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that… 25 arXiv — NLP / Computation & Language research 1d ago VICBench: A Multi-Language Benchmark for Code Vulnerability Detection arXiv:2608.12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the… 7 Hugging Face Daily Papers research 1d ago From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection Abstract A closed-loop framework combining physics-based video synthesis, diffusion-based video dereflection, and a new benchmark achieves state-of-the-art video reflection removal with fast inference. Generated by thinkingmachines/Inkling-Small Videos captured through glass… 29 Hugging Face Daily Papers research 1d ago MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Abstract Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only… 21 r/LocalLLaMA community 1d ago I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM. Hey guys, Just finished benchmarking DeepSeek V4 Flash 284B + DSpark on a single RTX PRO 6000 96GB . Short version: DSpark: ~15–17% faster generation on my coding workload On this setup, the DSpark drafter was faster in system RAM than VRAM q8_0 KV cache: 256K → 768K context… 18 r/LocalLLaMA community 1d ago LFM2.5-VL-3B recognizes Steve from Minecraft running locally on an iPhone 17 Liquid AI put out LFM2.5-VL-3B today, which is a 3.1B vision model that weighs roughly 2GB and fits well on a phone Benchmarks are benchmarks so I tried something sillier. Took a photo of a little Steve toy I have, gave it to the model and asked it what it was looking at It… 29 Hacker News — AI on Front Page community 1d ago Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index Article URL: https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis Comments URL: https://news.ycombinator.com/item?id=49275385 Points: 236 # Comments: 229 23 r/LocalLLaMA community 1d ago DeepSeek V4-Pro-0813 Benchmarks   submitted by   /u/MagicZhang [link]   [comments] 14 r/LocalLLaMA community 1d ago Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show Link to the article: KV Cache Quantization on Gemma 4 31B: Non-QAT vs QAT KLD benchmarks with BeeLlama.cpp v0.4.3 , fork of llama.cpp with more KV cache quantization options, comparing Gemma Q4_0 non-QAT vs Gemma Q4_0 QAT. Long story short: QAT is much more friendly to KV cache… 4 Hugging Face Daily Papers research 2d ago 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents Abstract A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning. Generated by thinkingmachines/Inkling-Small We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of… 8 Hugging Face Daily Papers research 2d ago Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Abstract The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models. Generated by thinkingmachines/Inkling-Small The rapid advancement of Large Language… 16 arXiv — Machine Learning research 2d ago Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification arXiv:2608.10007v1 Announce Type: new Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and… 30 arXiv — Machine Learning research 2d ago UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark… 29 arXiv — Machine Learning research 2d ago Toward Human Rights Benchmarking for LLMs: A Pilot Methodology arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this… 14 arXiv — Machine Learning research 2d ago Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting arXiv:2608.10891v1 Announce Type: new Abstract: Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time series generation has been developed primarily for data augmentation, where… 6 arXiv — Machine Learning research 2d ago Derivative Computation in PINNs: Automatic Differentiation, Finite Differences and Beyond arXiv:2608.11020v1 Announce Type: new Abstract: We systematically investigate finite-difference (FD) derivative computation in Physics-Informed Neural Networks (PINNs) as an alternative to automatic differentiation (AD). On three benchmark PDEs we show that, with a properly… 31 arXiv — NLP / Computation & Language research 2d ago Mapping and Measuring the Behavioral Evolution of Large Language Models arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using… 27 arXiv — Machine Learning research 2d ago Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives arXiv:2608.11093v1 Announce Type: new Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and… 35 Page 1 of 10 · 500 articles Older →