News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow r/LocalLLaMA community 9d ago inclusionAI/Ling-3.0-flash · Hugging Face The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing. Discussion on the benchmarks are here:… 12 r/LocalLLaMA community 9d ago Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the benchmark because it's quick to run, is pretty "real-world" and requires good… 37 Hugging Face Daily Papers research 9d ago MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Abstract Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across… 13 Hugging Face Daily Papers research 9d ago ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures Abstract Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset,… 6 Hugging Face Daily Papers research 9d ago GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation Abstract Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating them on specific tasks requires large, high-quality, multi-modal benchmarks that measure how well such models extract value from data. Concerning flood… 19 arXiv — Machine Learning research 10d ago Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark arXiv:2608.00106v1 Announce Type: new Abstract: Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. A controller may answer directly, decompose a request, retrieve evidence, execute code, delegate to a… 17 arXiv — Machine Learning research 10d ago Rethinking Pretraining for Specialized Design Data: Evidence from the JONES-19 Cultural Design Dataset arXiv:2608.00135v1 Announce Type: new Abstract: Design and architectural archives encode expert human knowledge in graphical formats, providing a critical testbed for design-inspired Machine Learning (ML) challenges absent with typical computer vision benchmarks. Building on… 33 arXiv — Machine Learning research 10d ago UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about… 33 arXiv — Machine Learning research 10d ago Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard arXiv:2608.01575v1 Announce Type: new Abstract: Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning… 20 arXiv — NLP / Computation & Language research 10d ago AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents arXiv:2608.00009v1 Announce Type: new Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns. We present AgentMemBench, a unified, reproducible benchmark… 32 arXiv — NLP / Computation & Language research 10d ago Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams arXiv:2608.00012v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing… 26 arXiv — NLP / Computation & Language research 10d ago XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding arXiv:2608.00036v1 Announce Type: new Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing… 23 arXiv — NLP / Computation & Language research 10d ago CurveShift: Is Agent Progress Scalar? Separating Level from Shape arXiv:2608.00355v1 Announce Type: new Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do… 6 arXiv — NLP / Computation & Language research 10d ago TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional… 38 arXiv — NLP / Computation & Language research 10d ago ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with… 19 arXiv — NLP / Computation & Language research 10d ago CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models arXiv:2608.01292v1 Announce Type: new Abstract: Legal reasoning is inherently jurisdiction-dependent: the same facts can call for different legal rules and yield different conclusions across legal systems. Yet existing benchmarks rarely evaluate whether large language models… 34 arXiv — NLP / Computation & Language research 10d ago LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning arXiv:2608.01328v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly… 10 arXiv — NLP / Computation & Language research 10d ago Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wording on this spend has not been measured systematically. We present a… 36 arXiv — NLP / Computation & Language research 10d ago Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important… 7 Hugging Face Daily Papers research 10d ago WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Abstract Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from… 6 Hugging Face Daily Papers research 10d ago SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Abstract Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to… 38 Hugging Face Daily Papers research 10d ago ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step Abstract To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expose semantic tool schemas in static… 24 r/MachineLearning community 10d ago I created an autonomous boxing benchmark [D] I created an AI boxing match to test the decision speed, adaptability and strategy. I fed the LLMs with data about the current match and if they have vision, they will get even more data. The match has street rules, anything goes and an AI is not defeated until the ref counts to… 24 r/LocalLLaMA community 10d ago 'I ran my own benchmarks on it' seems to be pretty common comment around here. How about dedicating a thread for this and sharing? Of course, the concern is that in the end, this thread will be fed into the models' training data, but I feel benchmarking isn't so open and very fragmented.   submitted by   /u/jinnyjuice [link]   [comments] 38 r/LocalLLaMA community 10d ago Qwen3.8-Max matches Kimi K3 and DeepSeek V4 Flash Qwen3.8-Max (2.4T) is another massive contribution to the open weight community. On benchmarks, it performs closely to Kimi K3 and DeepSeek V4 flash across all categories and is better at coding and software tasks. Qwen3.8-27B will also be open weight soon too. Weights are being… 11 NVIDIA Developer Blog official-blog 10d ago NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage Storage is an active part of every agentic AI workflow. As agents retrieve enterprise knowledge, access persistent memory, reuse key-value (KV) cache data,... 17 r/LocalLLaMA community 10d ago "Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out there and share knowledge if there is any interest. I also wasn't satisfied with… 32 r/LocalLLaMA community 10d ago V4-Flash-0731 - vibes after first weekend of use Spent way too much time with V4-Flash-0731 this weekend and wanted to share my vibes as briefly as possible. I sent it through a bit of real-work and some of my personal benchmarks. My quick thoughts are: Quantization hits this thing like a truck - I've tried a bunch of the Q2… 14 r/LocalLLaMA community 11d ago [RELEASE] SupraBrain-50M-v0.1 Hey there! So today we're releasing SupraBrain-50M, a hybrid language model that combines Gated DeltaNet linear recurrence with Sliding-Window Attention and Surprise-Gated update mechanisms to deliver very strong performance. Here are the benchmarks:… 19 Hugging Face Daily Papers research 11d ago Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants Abstract AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request… 15 Hugging Face Daily Papers research 11d ago ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Abstract Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a… 4 r/LocalLLaMA community 11d ago I benchmarked classic vector RAG vs Google's new OKF format vs both combined — same corpus, same 7 questions, all local (Ollama + ChromaDB) Google Cloud published OKF (Open Knowledge Format) on June 12th — a spec for storing curated knowledge as a directory of markdown files with YAML frontmatter. One concept per file, linked to each other, with an index.md for progressive disclosure. The only required field is… 7 Hugging Face Daily Papers research 11d ago SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift Abstract RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object detectors remain underexplored in this safety-critical domain. Limited cross-architecture benchmarking and insufficient… 32 Smol AI News news-outlet 11d ago Qwen 3.8 Max **Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with… 21 arXiv — Machine Learning research 11d ago Fast Rates for Swap-Agnostic Learning of Proper Losses arXiv:2607.28856v1 Announce Type: new Abstract: Swap-agnostic learning strengthens classical agnostic learning by allowing the comparator to select a different hypothesis on each level set of the learner's predictions. This benchmark captures prediction-dependent postprocessing,… 23 arXiv — Machine Learning research 11d ago FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents arXiv:2607.28945v1 Announce Type: new Abstract: Synthetic tabular data is increasingly used in privacy-preserving data sharing, data augmentation, and to mitigate downstream classifier bias. State-of-the-art tabular diffusion models such as TabDDPM and TabSyn achieve excellent… 21 arXiv — Machine Learning research 11d ago Learning Lookahead Lemmas for Neural Network Verification arXiv:2607.29051v1 Announce Type: new Abstract: State-of-the-art neural network verifiers use the branch-and-bound procedure as their core solving mechanism. We introduce an inprocessing framework for neural network verification driven by the lookahead procedure. Under this… 28 arXiv — Machine Learning research 11d ago Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives arXiv:2607.29064v1 Announce Type: new Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding. This… 21 arXiv — Machine Learning research 11d ago Simulation Code Generation for Fluid Systems using Large Language Models: Benchmarking Models and Prompting Strategies arXiv:2607.29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications. In this study, we explore how LLMs can be harnessed to automatically translate a neutral… 26 arXiv — Machine Learning research 11d ago GQ-FSL: Green Quantized Federated Split Learning arXiv:2607.29659v1 Announce Type: new Abstract: Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. While federated split learning (FSL) mitigates on-device… 28 arXiv — NLP / Computation & Language research 11d ago Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation arXiv:2607.28801v1 Announce Type: new Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric… 20 arXiv — NLP / Computation & Language research 11d ago TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal… 6 arXiv — NLP / Computation & Language research 11d ago Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monitoring, and human escalation. Yet evaluation often remains model-centric:… 26 arXiv — NLP / Computation & Language research 11d ago Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art arXiv:2607.29066v1 Announce Type: new Abstract: Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural Language Processing (NLP) offers a data-driven… 15 arXiv — NLP / Computation & Language research 11d ago M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models arXiv:2607.29125v1 Announce Type: new Abstract: Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations… 12 arXiv — NLP / Computation & Language research 11d ago Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding arXiv:2607.29196v1 Announce Type: new Abstract: Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identifying intended objects or referents, and withholding… 19 arXiv — NLP / Computation & Language research 11d ago ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation arXiv:2607.29539v1 Announce Type: new Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it… 32 arXiv — NLP / Computation & Language research 11d ago FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models arXiv:2607.29602v1 Announce Type: new Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic… 5 arXiv — NLP / Computation & Language research 11d ago Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review arXiv:2607.28631v1 Announce Type: cross Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose… 27 arXiv — NLP / Computation & Language research 11d ago Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery arXiv:2603.03322v2 Announce Type: replace Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However, rigorously evaluating an AI's capacity for knowledge discovery remains a critical… 19 Page 5 of 10 · 500 articles ← Newer Older →