News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 11d ago DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain arXiv:2605.07699v2 Announce Type: replace Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations. Despite… 14 arXiv — NLP / Computation & Language research 11d ago The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence arXiv:2606.21008v2 Announce Type: replace Abstract: The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs. No content is given in advance; the contestants create all of it -- a new kind of… 34 r/LocalLLaMA community 11d ago DeepSeek-V4-Flash-0731: surpasses Fable-5, Sol & Kimi-K3 on Chess Benchmark   submitted by   /u/mrwang89 [link]   [comments] 4 r/LocalLLaMA community 11d ago DSpark Benchmark Result on Deepseek v4 Flash 0731 TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF DSpark draft model from:… 31 r/MachineLearning community 11d ago Kimi K3 Deep Dive — Architecture, Training & Benchmarks of the 2.78-Trillion-Parameter Open-Weight Model [D] Hi everyone! 👋 I wrote an extensive technical deep-dive into Moonshot AI's Kimi K3 , analyzing its architectural innovations, training stability tricks, and benchmark performance. The blog covers: Kimi Delta Attention (KDA) Attention Residuals Stable LatentMoE Quantile… 4 r/LocalLLaMA community 11d ago Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes. I let Gemma4-31b run on my laptop for like almost a day using a heavily altered pi to do a deep dive on our beloved Llama tangentially related Subreddit, and this was the conclusion. Feels pretty accurate. Kind funny to let a small LLM loose and see what happens. Next target I'm… 30 r/MachineLearning community 12d ago [R] CausalVLBench: Benchmarking Visual Causal Reasoning in Large VLMs.   submitted by   /u/moschles [link]   [comments] 5 r/LocalLLaMA community 12d ago Why are almost all new benchmarks and leaderboards coding focused? I know in in this community LLM's are generally used for coding but there are other usecases besides coding and those usecases should be tested too. I also know benchmarks can sometimes be benchmaxxed and the model can still turn out shit but it can give a good outline on how a… 35 r/LocalLLaMA community 12d ago A collection of small domain-specific benchmarks for local models (30+ and growing) Hello fellow local AI people! I took "you must create your own benchmarks" literally, and built a website for this. How does the end result look like Let's say I want to know which model has most common sense in its responses, I did everything including evaluating responses (see… 34 TechCrunch — AI news-outlet 12d ago Judge denies xAI’s request to block Minnesota ban on ‘nudify’ apps Despite a lawsuit from xAI, a Minnesota ban on apps that allow users to “nudify” images can move forward. 31 r/LocalLLaMA community 12d ago I've had ling-3.0-flash and glm-5.2 both in my executor slot for a few weeks. They don't split the way the benchmarks predict Same harness, same task set, same agent scaffold, the only thing I swapped was the executor. Not a proper benchmark, no clean tok/s numbers, this is a workflow read not a leaderboard. glm-5.2 is the better model and it shows on anything that needs an actual decision. When the… 22 r/LocalLLaMA community 12d ago DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s. Thanks to the community help I finally launched this llm. LM Studio refused to load weight onto second GPU but Unsloth Studio did so everything was done in there. Not a proper benchmark (used PC in parallel as well) but it gives an idea of the performance from dual 3060 with… 16 r/LocalLLaMA community 12d ago Gemma4 (31B, bf16) constantly fails to edit files due to mismatches in original text - just me? I've been trying to find a good model to run locally, and in the benchmarks I can Gemma4 does well. However, whenever I give it anything that involves writing any code, it sits in loops trying to edit files sending the wrong original content (usually it messes up something like… 37 r/MachineLearning community 13d ago VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P] While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics.… 24 r/LocalLLaMA community 13d ago DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026 March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models available to run locally on <8K USD (us prices - just guestimating/not exact)… 19 r/LocalLLaMA community 13d ago Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper https://preview.redd.it/h7zv5tb3tmgh1.png?width=2854&format=png&auto=webp&s=507380e8f862c18f10f7c5c84da9e8d1c59139b0 Deepseek's new flash model is unexpectedly cheap and high-performing across useful benchmarks. It's priced at $0.09 / $0.18 per 1M. Truly "intelligence too cheap… 23 r/LocalLLaMA community 13d ago Deepseek V4 Flash on SlopCodeBench While waiting for some of the quants to drop, I load the API with $50 and ran it on SlopCodeBench Just vibe reading the results it seems like Opus 4.8 < Deepseek < Opus 5 https://github.com/michaelasper/benchmarks/blob/main/deepseek-v4-flash-on-slop-code-bench.md I was mostly… 12 r/LocalLLaMA community 13d ago SenseNova U1.5 Lite preview just dropped SenseNova released U1.5-Lite-Preview Benchmarks: Qwen-Image-Bench from 47.14 to 55.20. ImgEdit-Bench from 3.90 to 4.37. GEdit-Bench-en from 7.47 to 8.17. Key updates: 4K native generation with better texture, material, and lighting detail Improved Chinese and English text… 15 Hugging Face Daily Papers research 14d ago See2Think: Do Multimodal Models Really Use Intermediate Visual States? Abstract Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage… 9 r/LocalLLaMA community 14d ago DeepSeek-V4-Flash-0731 now far surpassing the DeepSeek-V4-Pro-Preview in benchmarks   submitted by   /u/SnooBunnies8392 [link]   [comments] 28 r/LocalLLaMA community 14d ago DeepSeek v4 Flash has a nice bump in Capability DeepSeek V4 Flash: Preview → 2026-07-31 Benchmark Preview 0731 Δ Terminal Bench* 56.9 82.7 +25.8 Toolathlon 51.8 70.3 +18.5 NL2Repo — 54.2 new Cybergym — 76.7 new DeepSWE — 54.4 new Agent Last Exam — 25.2 new Automation Bench — 25.1 new DSBench-FullStack — 68.7 new DSBench-Hard… 28 arXiv — Machine Learning research 14d ago DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series arXiv:2607.27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy… 29 arXiv — Machine Learning research 14d ago PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective arXiv:2607.27265v1 Announce Type: new Abstract: Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side Platform (DSP) bidding for advertisers, and Ad Exchange conducting auctions… 22 arXiv — Machine Learning research 14d ago Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance arXiv:2607.27283v1 Announce Type: new Abstract: Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary… 18 arXiv — Machine Learning research 14d ago ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders arXiv:2607.27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically… 8 arXiv — Machine Learning research 14d ago ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute… 18 arXiv — Machine Learning research 14d ago Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction arXiv:2607.28079v1 Announce Type: new Abstract: Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However, existing benchmarks often suffer from limited task diversity, fragmented… 24 arXiv — Machine Learning research 14d ago HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks arXiv:2607.28301v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race detection and benchmark question answering. However, knowledge alone does not… 23 arXiv — NLP / Computation & Language research 14d ago LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer… 36 arXiv — NLP / Computation & Language research 14d ago AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes arXiv:2607.27393v1 Announce Type: new Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has… 4 arXiv — NLP / Computation & Language research 14d ago Benchmarking LLM Competence on Logical Inference over Probability Operators arXiv:2607.27405v1 Announce Type: new Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes… 31 arXiv — NLP / Computation & Language research 14d ago From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved… 14 arXiv — NLP / Computation & Language research 14d ago RepBench: Compiling Benchmarks into Capability Representations for Large Language Models arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting measurements are difficult to compare or reproduce and… 28 arXiv — NLP / Computation & Language research 14d ago Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and… 26 arXiv — NLP / Computation & Language research 14d ago MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek arXiv:2607.28274v1 Announce Type: new Abstract: Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological… 38 arXiv — NLP / Computation & Language research 14d ago Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments arXiv:2607.28591v1 Announce Type: cross Abstract: Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable… 15 Hugging Face Daily Papers research 14d ago MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Abstract Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities,… 12 Hugging Face Daily Papers research 14d ago BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Abstract Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we… 29 r/LocalLLaMA community 14d ago Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus) Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without compromising others or making things bulky I’ve been noticing a massive… 14 Hugging Face Daily Papers research 14d ago SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch Abstract LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation… 18 r/LocalLLaMA community 14d ago Nanbeige4.2-3B: I'm not impressed I've tested Nanbeige-4.2-3B. On paper, the benchmarks promise it blows away Qwen3.5-9B and Gemma4-12B. My goal was to have something very light and fast to replace Qwen3.6-35B (or finetunes thereof) for simple and straightforward coding tasks. In the past I tried downgrading… 35 r/LocalLLaMA community 15d ago unsloth/Qwen3.6-27B-NVFP4 vs. Intel/Qwen3.6-27B-int4-AutoRound vs. nvidia/Qwen3.6-27B-NVFP4 -- which one to choose? Are there any benchmarks on these 4 bit quants, like how Artificial Analysis runs a slew of various benchmarks? If not, how can I run one (5x over for consistency) on them? I'm also very interested in hallucinations, as community discussions seem to point them out.  … 15 r/LocalLLaMA community 15d ago Benchmarked: MindControl for Llama.cpp I recently shared the original MindControl PoC (and on github ) - sampler-level guided reasoning budgets for llama.cpp, nudging the model with self-aware statements about its own thinking budget instead of just hard-truncating it. We received some great feedback, and the most… 27 Hugging Face Daily Papers research 15d ago OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding Abstract Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce… 14 r/LocalLLaMA community 15d ago 4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s I've been benchmarking a two-card box for a few weeks and I still can't quite get over some of these numbers, so I'm dumping them here. Box: RTX 4090 (24GB) + RTX 5060 Ti (16GB), i9-13900K, 64GB DDR5. WSL2 with 47GB allocated to the VM, CUDA 12.8 (12.8 specifically,13.1… 21 Hugging Face Daily Papers research 15d ago SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response Abstract Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks… 36 Smol AI News news-outlet 15d ago not much happened today **OpenAI** aggressively cut prices for **GPT-5.6 Luna** by 80% and **Terra** by 20%, introducing a faster **Sol Fast** tier with up to 2.5× lower latency at double the price, improving agent workflow costs by roughly 10×. The **ARC-AGI-3** debate highlighted that the complete… 14 arXiv — Machine Learning research 15d ago Entity Resolution in Practice: Lessons from a Self-Serve Pipeline arXiv:2607.26298v1 Announce Type: new Abstract: We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged that are absent from existing ER literature. (1) No single matching algorithm wins everywhere… 8 arXiv — Machine Learning research 15d ago Benchmarking ConvLSTM for One-Day-Ahead IMDAA Rainfall-Field Prediction across Four Indian Cities arXiv:2607.26581v1 Announce Type: new Abstract: Convolutional long short-term memory networks (ConvLSTMs) are widely used for precipitation forecasting, but most evidence for their performance comes from dense, high-frequency radar sequences. This study tests whether… 9 arXiv — Machine Learning research 15d ago Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models arXiv:2607.26922v1 Announce Type: new Abstract: Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial models. In this study, we investigate Parishad, a structured multi-agent system… 22 Page 6 of 10 · 500 articles ← Newer Older →