News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow r/LocalLLaMA community 8h ago Coding benchmarks that are quickly showcasing deep capability While we see for frontier models similar scores among famous coding benchmarks, across: DeepSWE, Terminal-Bench, LiveCodeBench, Code-Arena ELO. Here are in my opinion some next level benchmarks that really define deep intelligence , and complete capability in Software… 18 r/MachineLearning community 13h ago Is designing a memory graph around known data structure “overfitting” if I never touch the questions? [D] building a missing data infrastructure and started benchmarking long multi-session conversations (LoCoMo). I know the data looks like: people, facts, claims, events, timestamps, relations. So I extract those into a graph. I did not look at the QA pairs while building extractors… 26 r/LocalLLaMA community 18h ago Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant So after all my work, yeah, Raymond did it better, so I ported his work over, extended it turboX, extended it multiple other models (he had only Qwen models), and benchmarked the crap out of it to make sure it was worth it still. So really the credit goes to Raymond (… 9 r/MachineLearning community 20h ago Search agent beats GPT-6 Astra on benchmarks, just days after release [N]   submitted by   /u/Neither_You_5673 [link]   [comments] 4 r/MachineLearning community 2d ago Gpt 5,6,7: Does it even matter? The (ghost) productivity question. [D] an observation : GPT-5-class models are genuinely capable(They are) of doing a substantial fraction of knowledge work, why haven’t we seen a noticeable productivity shock in the real economy yet? Is AI actually less economically useful than the benchmarks suggest—or are… 32 r/LocalLLaMA community 2d ago I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code ( C code), the results were not completely unexpected but some quants were definitely underwhelming. TLDR : Best overall: bartowski/Qwen3.8-27B-IQ4_XS . Best… 32 r/LocalLLaMA community 2d ago Updated my benchmark with a new vLLM based recipe for Qwen 3.8 Flash Next : now up to 98/100 (instead of 91 previously) I was using: weights https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4 with the optimized SGLANG (patched) from https://old.reddit.com/r/BlackwellPerformance/comments/1w04xb7/qwen38_flashnext_on_1x_rtx_pro_6000_171_ts_c1_428/ Now I'm using: weights (AWQ W4A16) from:… 11 Hugging Face Daily Papers research 2d ago Last Translation Benchmark Abstract The Last Translation Benchmark introduces peer-reviewed, multimodal examples that break leading translation models alongside handcrafted verification rules for reliable, actionable evaluation. Generated by thinkingmachines/Inkling-Small For scientific progress, we need… 22 r/LocalLLaMA community 2d ago Any good alternatives to Artificial Analysis? So I used to use Artificial Analysis to compare models. The recent 61 score for Astra had me look into the results more granularly and I was very disappointed with what I found. Different pages reporting different scores for the same model on the same benchmark. Scores on a… 11 Hugging Face Daily Papers research 2d ago Environment Evolution for Terminal Agents Abstract Environment evolution incrementally raises task difficulty off-policy to sustain continuous learning signals for terminal agents, improving benchmark performance through multi-agent harnesses. Generated by thinkingmachines/Inkling-Small Scaling interactive and… 9 Latent.Space news-outlet 2d ago [AINews] GPT-6 Astra: OpenAI’s biggest LLM launch of all time new SOTA computer use and coding, 2.5x pricier per token, but WAY cheaper per task, less monitorable. overall, a very successful launch of OpenAI’s new frontier model class. 6 r/MachineLearning community 2d ago GPT-6 is released [N] Benchmark scores (GPT-6 uses a harness for ARC-AGI-3, and is at about 60% without one): https://preview.redd.it/v7nik4nbtfnh1.png?width=1378&format=png&auto=webp&s=a6ec04b5b87e7f2dce748b275d878ab0243f751d https://openai.com/index/gpt-6-astra/   submitted by  … 36 arXiv — Machine Learning research 2d ago Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails arXiv:2609.03504v1 Announce Type: new Abstract: Restricted eigenvalue (RE) bounds govern stable recovery by norm-regularized estimators. For isotropic sub-Gaussian measurements, the benchmark sample size is $1+w(A)^2$, where $w(A)$ is the Gaussian width of the normalized descent… 12 arXiv — Machine Learning research 2d ago LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues arXiv:2609.03507v1 Announce Type: new Abstract: Tracking depression from multi-session counseling dialogues requires estimating both current symptom severity and how it changes across sessions. Yet progress on this task is constrained by the scarcity of longitudinal counseling… 17 arXiv — Machine Learning research 2d ago WeatherNext 3: Increasing resolution and performance of global weather models with raw observations arXiv:2609.03582v1 Announce Type: new Abstract: State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best… 36 arXiv — Machine Learning research 2d ago RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models arXiv:2609.04007v1 Announce Type: new Abstract: Despite strong performance on held-out electroencephalography (EEG) data, seizure detectors may fail under real-world acquisition variability, artifacts, and adversarial inputs. We introduce RobustSeiz, an open-source,… 37 arXiv — NLP / Computation & Language research 2d ago BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events arXiv:2609.02895v1 Announce Type: new Abstract: Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social… 28 arXiv — NLP / Computation & Language research 2d ago Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards arXiv:2609.02899v1 Announce Type: new Abstract: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions:… 33 arXiv — NLP / Computation & Language research 2d ago LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation arXiv:2609.02954v1 Announce Type: new Abstract: Identifying the issues disputed between litigating parties is a crucial component of real-world litigation. However, legal issues remain comparatively underexplored in legal AI research. In this work, we study the computational… 13 arXiv — NLP / Computation & Language research 2d ago SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking arXiv:2609.03047v1 Announce Type: new Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what… 16 arXiv — NLP / Computation & Language research 2d ago FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models arXiv:2609.03331v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually… 19 arXiv — NLP / Computation & Language research 2d ago FrameBench:A Language Understanding Benchmark Based on Frame Semantics arXiv:2609.03370v1 Announce Type: new Abstract: In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent… 29 arXiv — NLP / Computation & Language research 2d ago Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory arXiv:2609.03394v1 Announce Type: new Abstract: Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in… 7 arXiv — NLP / Computation & Language research 2d ago When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents arXiv:2609.03467v1 Announce Type: new Abstract: Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather… 37 arXiv — NLP / Computation & Language research 2d ago KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records arXiv:2609.03597v1 Announce Type: new Abstract: Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten… 24 arXiv — NLP / Computation & Language research 2d ago Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements arXiv:2609.03654v1 Announce Type: new Abstract: The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and… 25 arXiv — NLP / Computation & Language research 2d ago Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks arXiv:2609.03734v1 Announce Type: new Abstract: BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit… 30 arXiv — NLP / Computation & Language research 2d ago Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation arXiv:2609.03814v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can… 10 arXiv — NLP / Computation & Language research 2d ago Last Translation Benchmark arXiv:2609.04173v1 Announce Type: new Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are… 20 arXiv — NLP / Computation & Language research 2d ago MedQA-MM: Shortcuts Behind Medical Visual Reasoning arXiv:2609.03261v1 Announce Type: cross Abstract: A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the… 11 arXiv — NLP / Computation & Language research 2d ago HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews arXiv:2609.03580v1 Announce Type: cross Abstract: The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination… 7 arXiv — NLP / Computation & Language research 2d ago RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents arXiv:2609.03773v1 Announce Type: cross Abstract: Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone.… 27 r/LocalLLaMA community 2d ago Has anyone already tried IFM's new K2-Horizon-MoVA-36B-A4B? How good/bad is it against comparable MoEs the same size? How does it compare against Qwen 3.6 35BA3B? Since we don't have 3.8 35B this seems like an upgrade if we look at some benchmarks like terminal bench, but they don't have SWE bench pro on the benchmarks table, and i don't… 38 Hugging Face Daily Papers research 2d ago LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Abstract LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art… 26 Hugging Face Daily Papers research 2d ago Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Abstract Fully quantizing hybrid LLMs—including recurrent Gated DeltaNet layers—to 4-bit NVFP4 preserves accuracy across long-context and reasoning benchmarks by localizing outliers and exploiting robust delta-rule dynamics. Generated by thinkingmachines/Inkling-Small Hybrid… 30 r/LocalLLaMA community 2d ago Qwen 3.8 27B Vs. Qwen 3.6 27B on oMLX Quality: 81.1 → 87.7 (+8%) Speed: 35 → 29 tok/s (−16%) Runtime: 8m51s → 44m39s (5x longer) Output tokens: 18K → 78K (🤯) Noticeably better quality, but you're paying for it with tokens and time. Full benchmark results (all hardware, all quants): llm-bench.io Qwen3.6-27B Vs.… 28 r/LocalLLaMA community 2d ago The benchmarks the big labs don't want you to see   submitted by   /u/jd_3d [link]   [comments] 32 r/LocalLLaMA community 3d ago How does meta spark 1.3 match claude fable 5 on benchmarks? Did anyone try it in agentic coding? How was your experience with it and is it just benchmaxed or is it really that good?   submitted by   /u/Personal-Try2776 [link]   [comments] 29 OpenAI official-blog 3d ago GPT-6 Astra: A new generation of intelligence Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science. 6 Hugging Face Daily Papers research 3d ago Exploring Collaboration between a language and a non-language agent Abstract A benchmark of collaborative chess tasks shows that integrating continuous subagent representations directly into language models via learned state tokens outperforms text-based verbalization and scales effectively. Generated by thinkingmachines/Inkling-Small LLMs are… 12 Hugging Face Daily Papers research 3d ago Post-Training Language Models for Gold-Medal Performance in Coding Competitions Abstract A specialization pipeline combining curated problems, synthetic reasoning, supervised fine-tuning, and reinforcement learning trains competitive programming models that exceed top human scores on IOI benchmarks using iterative test-time refinement. Generated by… 16 Hugging Face Daily Papers research 3d ago SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions Abstract SnapBench introduces paired corruption benchmarks for mobile snap-and-ask retrieval, revealing that image noise severely degrades multimodal retrieval and proposing an adaptive fusion method to calibrate modality reliability. Generated by thinkingmachines/Inkling-Small… 17 arXiv — Machine Learning research 3d ago Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal Control arXiv:2609.01676v1 Announce Type: new Abstract: Reinforcement learning achieves strong traffic signal control performance in simulation, yet policies trained in simulators often fail once deployed in the real world, a failure known as the Sim-to-Real gap. When RL is applied to… 25 arXiv — Machine Learning research 3d ago Reinforcement Learning and Rule-Based Peer-to-Peer Pricing in Residential PV-BES Communities arXiv:2609.01680v1 Announce Type: new Abstract: This paper compares rule-based and learning-based pricing mechanisms for peer-to-peer (P2P) electricity trading in residential photovoltaic communities. The rule-based benchmarks comprise bill-sharing as an ex post allocation… 29 arXiv — Machine Learning research 3d ago CAT-Flow: Curvature-Adaptive sTeps for Flow Matching arXiv:2609.01746v1 Announce Type: new Abstract: Flow Matching has emerged as a leading framework for generative modeling, powering state-of-the-art systems such as FLUX and Stable Diffusion 3.5. However, the iterative nature of its ODE-based sampling process creates a… 33 arXiv — Machine Learning research 3d ago hLLM: Single Pass Decoding for Generative Reranking arXiv:2609.01807v1 Announce Type: new Abstract: Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the… 9 arXiv — NLP / Computation & Language research 3d ago Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation arXiv:2609.02006v1 Announce Type: cross Abstract: A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inheritance distiller, Low-Rank Clone (LRC), deploys… 6 arXiv — Machine Learning research 3d ago AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers arXiv:2609.02339v1 Announce Type: new Abstract: World modeling requires a predictive model to maintain and update an internal state adequate for reasoning about the consequences of actions. We introduce the AGI Maze Prediction Datasets and Benchmark, a lightweight controlled… 9 arXiv — Machine Learning research 3d ago FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making arXiv:2609.01691v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to make decisions from visual inputs. We introduce FAIRLENS, a benchmark and evaluation framework for measuring both the fairness and the validity of VLM responses in three… 21 arXiv — NLP / Computation & Language research 3d ago MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models arXiv:2609.01772v1 Announce Type: new Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual… 29 Page 1 of 10 · 500 articles Older →