News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 17d ago Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail… 21 arXiv — Machine Learning research 17d ago Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions arXiv:2607.22931v1 Announce Type: new Abstract: Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursive Least Squares (RLS) and have achieved the state-of-the-art results compared to… 13 arXiv — Machine Learning research 17d ago Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram arXiv:2607.22980v1 Announce Type: new Abstract: Brain-computer interfaces (BCIs) have long sought calibration-free operation, but classifiers are typically benchmarked by discrimination alone, blind to whether predicted probabilities are well calibrated - a meaningful gap given… 29 arXiv — Machine Learning research 17d ago ParasGB: A Graph Benchmark Suite for Parasitic Estimation on AMS Circuits arXiv:2607.23225v1 Announce Type: new Abstract: As chip manufacturing processes advance to deep submicron nodes, parasitic interconnect effects increasingly dominate the performance of analog and mixed-signal (AMS) circuits and often lead to costly layout iterations. This makes… 10 arXiv — Machine Learning research 17d ago Harmonized Interpretable ECG Waveform Features for Robust Cross-Dataset Clinical Prediction arXiv:2607.23412v1 Announce Type: new Abstract: Electrocardiograms (ECGs) are widely used for cardiovascular risk prediction, yet models often fail to transfer across hospitals because of protocol, population, and measurement differences. We benchmark cross-dataset… 21 arXiv — NLP / Computation & Language research 17d ago Learning When to Reason for Text-to-SQL via SFT and DPO arXiv:2607.22622v1 Announce Type: new Abstract: Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of… 36 arXiv — NLP / Computation & Language research 17d ago PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs arXiv:2607.22859v1 Announce Type: new Abstract: Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progress in high resource languages, Bengali remains underexplored due to the… 24 arXiv — NLP / Computation & Language research 17d ago ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation arXiv:2607.23058v1 Announce Type: new Abstract: Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical… 33 arXiv — NLP / Computation & Language research 17d ago Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking arXiv:2607.23514v1 Announce Type: new Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using… 28 arXiv — NLP / Computation & Language research 17d ago Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages arXiv:2607.23808v1 Announce Type: new Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field… 5 arXiv — NLP / Computation & Language research 17d ago Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance arXiv:2607.23813v1 Announce Type: new Abstract: We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i)… 12 arXiv — NLP / Computation & Language research 17d ago StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting arXiv:2607.24191v1 Announce Type: new Abstract: Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly… 36 arXiv — NLP / Computation & Language research 17d ago Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets arXiv:2607.24268v1 Announce Type: new Abstract: Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation… 6 arXiv — NLP / Computation & Language research 17d ago INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models arXiv:2607.24273v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This… 32 arXiv — NLP / Computation & Language research 17d ago Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory arXiv:2607.24368v1 Announce Type: new Abstract: Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the… 31 arXiv — NLP / Computation & Language research 17d ago Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy arXiv:2607.22554v1 Announce Type: cross Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study… 32 Hugging Face Daily Papers research 17d ago DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification Abstract Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may… 29 Hacker News — AI on Front Page community 17d ago Benchmarking Opus 5 on SlopCodeBench Article URL: https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarking-opus-5-on-slop-code-bench.md Comments URL: https://news.ycombinator.com/item?id=49076391 Points: 213 # Comments: 52 12 r/MachineLearning community 17d ago Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R] I ran a solo evaluation project benchmarking six current frontier models: GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro, Gemini Flash, and Grok 4.3. I tested tham across 8 established bias/fairness datasets (WinoBias, BBQ Race/Ethnicity, SeeGULL, OpinionsQA, cajcodes… 26 Hugging Face Daily Papers research 18d ago SceneActBench: Can Agents Act on the 3D Scenes They See? Abstract Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present… 36 r/LocalLLaMA community 18d ago Qwen3.6-27B speculative decoding gets better on heavier quants I finished the speed leg of my spec-decode benchmarking for Qwen3.6-27B, main algorithms across quants. Overall: the heavier the quant, the more spec-decode buys you (10 of 10 speculative configs rank Q8 > Q6 > Q4 by multiplier). Acceptance is quant-independent at matched depth,… 33 Hugging Face Daily Papers research 18d ago DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Abstract The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM-driven data… 5 Smol AI News news-outlet 18d ago not much happened today **Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with… 33 r/LocalLLaMA community 18d ago What local model do you still use after the hype wore off? Every time a new model is released, I tend to check it out. The benchmarks, readme, or whatever seem pretty convincing, so I download it, test it for a few hours, and then I just go back to the same couple of ones I already had. Curious what models people here have actually kept… 20 arXiv — Machine Learning research 18d ago Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under… 32 arXiv — Machine Learning research 18d ago Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning arXiv:2607.21637v1 Announce Type: new Abstract: This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments. Various sampling methods are used to bound a population-based search and… 24 arXiv — Machine Learning research 18d ago Searching the Space of Feed-Forward Neural-Network Weight-Update Rules with Fixed Depth Symbolic Regression arXiv:2607.21855v1 Announce Type: new Abstract: We investigate whether symbolic regression can discover explicit neural network weight-update rules that outperform standard hand-designed optimizers on small symbolic regression benchmarks. Candidate update rules are represented… 12 arXiv — Machine Learning research 18d ago Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study arXiv:2607.21866v1 Announce Type: new Abstract: Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We present a distributed classroom-scale replication:… 24 arXiv — Machine Learning research 18d ago CEL: Comprehensive Counterfactual Explanations Library and Benchmark arXiv:2607.22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods… 38 arXiv — Machine Learning research 18d ago SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text arXiv:2607.21610v1 Announce Type: cross Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and… 12 arXiv — NLP / Computation & Language research 18d ago From Seasonality to Semantics: Benchmarking a Hybrid Probabilistic Forecasting System for Roadblocks in Bolivia arXiv:2607.21785v1 Announce Type: cross Abstract: Roadblocks in Bolivia are a social conflict phenomenon with devastating economic impacts, estimated at losses equivalent to 4% of the national Gross Domestic Product. Despite their recurrence and impact, there is a lack of local… 8 arXiv — Machine Learning research 18d ago Probing Speaker Identity Sensitivity in Audio Deepfake Detectors arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate… 31 arXiv — NLP / Computation & Language research 18d ago A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models arXiv:2607.21632v1 Announce Type: new Abstract: Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone… 25 arXiv — NLP / Computation & Language research 18d ago Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark arXiv:2607.21685v1 Announce Type: new Abstract: A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings… 13 arXiv — NLP / Computation & Language research 18d ago Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms arXiv:2607.21780v1 Announce Type: new Abstract: Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into their constituent documents is difficult, especially for low-resource languages.… 31 arXiv — NLP / Computation & Language research 18d ago Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contamination problems -- and they overwhelmingly measure short interaction histories. We… 18 arXiv — NLP / Computation & Language research 18d ago Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination arXiv:2607.22067v1 Announce Type: new Abstract: The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on… 26 arXiv — NLP / Computation & Language research 18d ago From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models arXiv:2607.22182v1 Announce Type: new Abstract: Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities… 30 arXiv — NLP / Computation & Language research 18d ago DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write… 17 arXiv — NLP / Computation & Language research 18d ago InteractComp: Evaluating Search Agents With Ambiguous Queries arXiv:2510.24668v2 Announce Type: replace Abstract: Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a… 13 arXiv — NLP / Computation & Language research 18d ago LMEB: Long-horizon Memory Embedding Benchmark arXiv:2603.12572v5 Announce Type: replace Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchmarks, which narrowly focus on traditional passage retrieval and fail to assess… 27 arXiv — NLP / Computation & Language research 18d ago WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47… 25 arXiv — NLP / Computation & Language research 18d ago Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models arXiv:2605.17770v4 Announce Type: replace-cross Abstract: The advancement of Large Reasoning Models (LRMs) has catalyzed a paradigm shift from reactive ``fast thinking'' text generation to systematic, step-by-step ``slow thinking'' reasoning, unlocking state-of-the-art… 17 Vercel — AI dev-tools 18d ago DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities Last week, OpenAI evaluated two models on an exploit benchmark within an isolated sandbox. Guardrails were reduced for testing, and the models found a vulnerability in their environment, accessed the internet, and reached Hugging Face's production database. No human directed the… 30 r/LocalLLaMA community 18d ago Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Claude code (with DS in CLIProxyAPI ) takes nearly 4 times longer than the fastest… 38 r/LocalLLaMA community 18d ago BeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), and more. BeeLLama v0.4.1 is here, building up on top of v0.4.0 feature set, now… 37 r/LocalLLaMA community 19d ago 23 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most broken This is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet. We also have a new abliterlitics discord , feel free to jump on and roast my choice of benchmarks! Or just chat and hang out. This is similar… 32 r/MachineLearning community 19d ago We compared different LLMs on IMO 2026 [R] There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math problems are quite a good proxy for general intelligence capability - These are… 30 r/LocalLLaMA community 19d ago I run 35B–480B coding models on my 36 GB MacBook by streaming MoE experts from SSD — self-contained app, and I publish the benchmarks that *failed* too I got tired of "your Mac can't run that" so I forked llama.cpp to stream a MoE model's expert weights from SSD instead of forcing the whole thing into RAM. A MoE only fires a few experts per token, so most weights sit idle — Slipstream keeps the always-needed weights resident… 12 r/LocalLLaMA community 19d ago Benchmarks: TensorSharp vs. llama.cpp Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like Gemma4, DiffusionGemma, Qwen3.6 with multi-modal (image, vision, audio), Qwen… 38 Page 8 of 10 · 500 articles ← Newer Older →