News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow r/LocalLLaMA community 6d ago [MASSIVE RELEASE] Supra2-IMG - a tiny 100M text-to-image model - SOTA quality and open release! Hey everyone! It has been quite a while since the last SupraLabs model - but today we've something special for y'all: Supra2-IMG It's a 100M parameter DiT text-to-image model trained entirely from scratch in under 10 hours on a single H100 on Runpod. It can generate… 32 TechCrunch — AI news-outlet 6d ago Where will the next breakout startup come from? Benchmark’s full partnership weighs in at TechCrunch Disrupt 2026 Where will the next breakout startup come from? Benchmark’s full partnership weighs in on the main stage at TechCrunch Disrupt 2026. Save up to $200 before September 25, 11:59 p.m. PT. Register now. 5 r/LocalLLaMA community 6d ago Putting the question before the context took my local Qwen from 89% to 100% on a decision benchmark, and from ~400 ms to ~80 ms Small, free finding. I use local Qwen models for typed decisions: a state plus a question with fixed allowed answers, and I read the probability of each answer from the logits of one forward pass instead of generating text. I used to build the prompt like you would for a human:… 12 r/LocalLLaMA community 6d ago [Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff" https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=webp&s=74765dbd409f4c221640f9f6000a685f6fdbb242 Spent weekend benchmarking the Splash engine (by Incoai) and extending its architecture to native 8-bit on Apple Silicon (M5 Pro, 64 GB unified memory). Splash is… 10 arXiv — Machine Learning research 7d ago SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs? arXiv:2609.21190v1 Announce Type: new Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and… 17 arXiv — Machine Learning research 7d ago Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization arXiv:2609.21197v1 Announce Type: new Abstract: Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker… 16 arXiv — Machine Learning research 7d ago OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems arXiv:2609.21527v1 Announce Type: new Abstract: Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score… 26 arXiv — Machine Learning research 7d ago Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning arXiv:2609.21909v1 Announce Type: new Abstract: In this work, we conduct a systematic comparison of two state-of-the-art motion-imitation reinforcement learning (MIRL) pipelines, one built on SCONE/HyFyDy and one built on MuJoCo/MyoSim. HyFyDy emphasizes physiological realism… 34 arXiv — Machine Learning research 7d ago Benchmarking World Models for Continual Learning on Compositional Tasks arXiv:2609.22055v1 Announce Type: new Abstract: A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge… 10 arXiv — Machine Learning research 7d ago BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings arXiv:2609.22064v1 Announce Type: new Abstract: Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural… 14 arXiv — NLP / Computation & Language research 7d ago PhysioBench: A Unified Benchmark for Physiological Signal Question Answering arXiv:2609.20836v1 Announce Type: new Abstract: Physiological signals support diverse clinical and monitoring tasks, yet existing physiological signal foundation models typically require task-specific adaptation for each task. Natural language provides a common interface for… 34 arXiv — NLP / Computation & Language research 7d ago VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering arXiv:2609.20843v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) enables models to answer natural-language questions through structured graph reasoning and has achieved substantial progress across many benchmarks and applications. Recently, multimodal… 28 arXiv — NLP / Computation & Language research 7d ago TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar arXiv:2609.20832v1 Announce Type: new Abstract: We introduce TatBLiMP, the first benchmark of linguistic minimal pairs for Tatar (tt, ISO 639-3 tat), a Qypchaq Turkic language written in Cyrillic. To our knowledge it is the first grammaticality evaluation for Tatar language… 28 arXiv — NLP / Computation & Language research 7d ago MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs arXiv:2609.20850v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related… 29 arXiv — NLP / Computation & Language research 7d ago $\mu^2$-Bench: A Multilingual Machine Unlearning Benchmark arXiv:2609.20945v1 Announce Type: new Abstract: Undesired information such as harmful content and private data propagates through Multilingual Large Language Models (LLMs) via direct training and indirect cross-linguistic spread. Multilingual Machine Unlearning (MMU) aims to… 5 arXiv — NLP / Computation & Language research 7d ago Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection arXiv:2609.21179v1 Announce Type: new Abstract: Neural morphological generation systems often achieve high aggregate accuracy on benchmark datasets, yet such performance can conceal systematic errors clustered in rare morphological subclasses. We present an orthography-aware… 10 arXiv — NLP / Computation & Language research 7d ago Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction arXiv:2609.21392v1 Announce Type: new Abstract: Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of… 28 arXiv — NLP / Computation & Language research 7d ago Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations arXiv:2609.21490v1 Announce Type: new Abstract: Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that… 19 arXiv — NLP / Computation & Language research 7d ago Chinese Competitive Debating Dataset and Benchmark arXiv:2609.21637v1 Announce Type: new Abstract: Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared… 17 arXiv — NLP / Computation & Language research 7d ago PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction arXiv:2609.21673v1 Announce Type: new Abstract: Probabilistic Graphical Models (PGMs), especially Bayesian Networks (BNs), expose directed structure and probabilistic parameters, making them natural symbolic targets for neurosymbolic AI. Yet training text-to-parameterized-BN… 30 arXiv — NLP / Computation & Language research 7d ago CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords arXiv:2609.21722v1 Announce Type: new Abstract: Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily… 24 arXiv — NLP / Computation & Language research 7d ago QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge arXiv:2609.22038v1 Announce Type: new Abstract: We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval,… 4 arXiv — NLP / Computation & Language research 7d ago Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance arXiv:2511.13254v2 Announce Type: replace Abstract: Large Language Models (LLMs) have displayed remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive computational resources and careful orchestration of… 33 arXiv — NLP / Computation & Language research 7d ago Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA arXiv:2602.00279v2 Announce Type: replace Abstract: Reliable uncertainty quantification (UQ) is essential for safe deployment of large language models (LLMs) in scientific question answering, where long-form outputs exceed practical human verification at scale. We introduce the… 17 arXiv — NLP / Computation & Language research 7d ago MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks arXiv:2602.16313v2 Announce Type: replace Abstract: Existing evaluations of agents with memory typically assess memorization and action in isolation. One class of benchmarks evaluates memorization by testing recall of past conversations or text but fails to capture how memory is… 18 arXiv — NLP / Computation & Language research 7d ago JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems arXiv:2604.23478v3 Announce Type: replace Abstract: Large language models are widely used to judge the output of other language models, yet whether a judge returns the same verdict when the same request is worded differently remains largely unexamined. We study that question… 23 r/MachineLearning community 7d ago Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D] In February OpenAI stopped reporting SWE-bench Verified and recommended other labs stop too. Every frontier model they tested could reproduce the human-written reference fix, or verbatim details of the problem statement, for some tasks. Progress had slowed to six points in six… 33 r/LocalLLaMA community 8d ago Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3) Hey everyone, After wrestling with hardware and engine issues for days, I finally got Qwen 3.8 Next running properly on my multi-GPU rig. Thought I’d share the setup journey, benchmarks, and thermal results for anyone trying something similar. Seeing all the ongoing memes on… 16 r/LocalLLaMA community 8d ago With Gemini 4, bench goes up. They claimed open-weight models are dangerous but the benchmarks say otherwise.   submitted by   /u/Intrepid_Travel_3274 [link]   [comments] 22 r/LocalLLaMA community 8d ago Ternary-Bonsai-2-27B-PQ2_0 is not completely lobotomized I decided to run prism-ml/Ternary-Bonsai-2-27B-PQ2_0 through my own set of UNSCIENTIFIC benchmarks. I needed something to compare it to, so I decided I would compare with another 27B model by filesize: unsloth/Qwen3.8-27B-UD-IQ2_XXS. Since anyone considering running a 27B model… 23 r/MachineLearning community 8d ago Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D] Since OpenAI retired SWE-bench Verified in February (every frontier model tested could reproduce reference fixes for some tasks; underspecified tests rewarded knowing the intended fix), I've been trying to write down precisely what a decontamination report can and can't… 27 TechCrunch — AI news-outlet 8d ago Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models. 32 r/LocalLLaMA community 9d ago Prefil of local models vs opus and astra Why does no one talk about what the prefil speeds of these API providers are vs running locally. People with sparks or strix halos only seem to focus on decode without considering how much slower it is because of slow pp. Are there any benchmarks / figures of how fast the APIs… 8 r/LocalLLaMA community 9d ago M5 Ultra and M6 Chip Benchmark Results Reveal Graphics Performance First benchmarks   submitted by   /u/DustNearby2848 [link]   [comments] 18 NVIDIA Developer Blog official-blog 9d ago Benchmarking LLM Inference at Scale with AIPerf You’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast? Your instincts might lead you to send... 24 r/LocalLLaMA community 9d ago We benchmarked 24 LLMs against human writers on 475 creative writing prompts We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts. Creative writing is subjective, so the rankings aren't meant to predict what any one person will prefer. Instead, they predict what a large… 26 r/LocalLLaMA community 9d ago Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison Hey r/LocalLLaMA , Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark . We wanted to see… 24 r/LocalLLaMA community 10d ago bonsai's document reveal how much cherry picked their headlines are bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks. that qwen3.5 is a… 19 r/LocalLLaMA community 10d ago Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo Thanks for the exceptional support ( https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i_literally_built_the_jev_architecture_one_year/ ) and for the dozens of requests to make a generic model, run benchmarks, and create an HF space so anyone can test it. So here you go,… 30 arXiv — Machine Learning research 10d ago Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data arXiv:2609.20218v1 Announce Type: new Abstract: Large language models can label a tabular row from a plain-English description with no training - a capability now shipping in mainstream spreadsheet tools such as Microsoft Copilot in Excel and Anthropic's Claude for Excel -… 27 arXiv — NLP / Computation & Language research 10d ago Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry arXiv:2609.19154v1 Announce Type: new Abstract: While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training… 4 arXiv — NLP / Computation & Language research 10d ago V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering arXiv:2609.19879v1 Announce Type: new Abstract: Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the… 11 arXiv — NLP / Computation & Language research 10d ago PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces arXiv:2609.19883v1 Announce Type: new Abstract: Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained,… 8 arXiv — NLP / Computation & Language research 10d ago KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms arXiv:2609.19916v1 Announce Type: new Abstract: Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary… 32 arXiv — NLP / Computation & Language research 10d ago Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence arXiv:2609.19965v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical… 28 arXiv — NLP / Computation & Language research 10d ago Benchmarking LLM Compliance with China AI Generated Content Regulations arXiv:2609.19989v1 Announce Type: new Abstract: The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper… 36 arXiv — NLP / Computation & Language research 10d ago SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment arXiv:2609.20584v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive… 22 arXiv — NLP / Computation & Language research 10d ago What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks arXiv:2609.19182v1 Announce Type: cross Abstract: Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of… 38 arXiv — NLP / Computation & Language research 10d ago From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization arXiv:2609.19630v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the… 25 Vercel — AI dev-tools 10d ago Run Terminal-Bench and other Harbor evals on Vercel Sandbox You can now run Harbor evals on Vercel Sandbox. Harbor is the open-source harness behind Terminal-Bench , whose registry includes many other benchmarks such as SWE-bench, tau3-bench and OSWorld. Pass --env vercel to harbor run and each trial executes in its own isolated… 18 Page 3 of 10 · 500 articles ← Newer Older →