News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow Vercel — AI dev-tools 10d ago Run Terminal-Bench and other Harbor evals on Vercel Sandbox You can now run Harbor evals on Vercel Sandbox. Harbor is the open-source harness behind Terminal-Bench , whose registry includes many other benchmarks such as SWE-bench, tau3-bench and OSWorld. Pass --env vercel to harbor run and each trial executes in its own isolated… 18 r/LocalLLaMA community 10d ago First M5 Ultra benchmarks just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!   submitted by   /u/Ashefromapex… 8 r/LocalLLaMA community 11d ago AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR I let AI run my phone 60 real tasks, back to back, on the OnePlus I use every day Best text model still failed 43% of them Peak chip temp 98.2 C 69% of the battery gone The benchmark is AndroidLife, and this is the first of 11 models, qwen3.8-27b from Alibaba, running in text… 24 arXiv — Machine Learning research 11d ago EdgeReMIND: A Scalable, Top-Ranked Memorization Baseline for Temporal Multi-Relational Link Prediction arXiv:2609.17916v1 Announce Type: new Abstract: Temporal link prediction on the Temporal Graph Benchmark 2.0 (TGB 2.0) faces a scalability ceiling: on the benchmark's three largest datasets, every existing embedding method runs out of memory or exceeds the time budget. These… 13 arXiv — Machine Learning research 11d ago iMINDBench: iEEG Multi-Institution Neural Decoding Benchmark arXiv:2609.18104v1 Announce Type: new Abstract: Intracranial electroencephalography (iEEG) is widely used to record electrical activity directly from electrodes inside the human brain, making it an attractive modality for neural decoding. However, progress in iEEG decoding,… 6 arXiv — Machine Learning research 11d ago Reliable Virtual Sensing: A Multi-Domain Benchmark for Robustness Under Sensor Failures arXiv:2609.18396v1 Announce Type: new Abstract: Virtual sensing, the estimation of hard-to-measure quantities from available sensor measurements, is a critical enabler for control and monitoring in cyber-physical systems. However, when sensors fail, learning-based predictors can… 15 arXiv — Machine Learning research 11d ago StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction arXiv:2609.18949v1 Announce Type: new Abstract: We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed agentic systems on diagnosing peg stress and forecasting… 14 arXiv — NLP / Computation & Language research 11d ago From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings arXiv:2609.17538v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for structured information extraction from documents, yet their behavior under realistic OCR noise remains poorly understood. We present a systematic benchmark of open-source… 7 arXiv — NLP / Computation & Language research 11d ago English Word Sense Disambiguation in 2026: When the Labels Become the Bottleneck arXiv:2609.17554v1 Announce Type: new Abstract: In English all-words word sense disambiguation (WSD), the labels, not the models, have become the bottleneck: frontier LLMs are accurate enough that the errors surviving in the gold standard decide benchmark rankings -- in the test… 31 arXiv — NLP / Computation & Language research 11d ago DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning arXiv:2609.18135v1 Announce Type: new Abstract: State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema linking and SQL generation. However, existing work trains separate models for each task, failing to leverage the… 24 arXiv — NLP / Computation & Language research 11d ago TeochewBench: A Human-Reviewed Benchmark for Teochew Hanzi Translation arXiv:2609.18156v1 Announce Type: new Abstract: Teochew has a substantial speaker community and exhibits distinctive lexical, syntactic, and pragmatic features, yet textual resources for evaluating large language models remain limited. We present TeochewBench, a human-reviewed… 28 arXiv — NLP / Computation & Language research 11d ago Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors arXiv:2609.18203v1 Announce Type: new Abstract: Human values are deep motivational orientations that shape human behaviors. In e-commerce, they reveal the stable drivers behind users' purchase decisions. Compared with short-term interests, consumer values better explain how… 26 arXiv — NLP / Computation & Language research 11d ago I code or AI code: A comparative evaluation of AI-rated scores in classroom observations arXiv:2609.18274v1 Announce Type: new Abstract: Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study… 26 arXiv — NLP / Computation & Language research 11d ago Fallacy Benchmarks Measure Scheme Recognition, Not Fallacy Detection arXiv:2609.18644v1 Announce Type: new Abstract: Fallacy-detection benchmarks pair fallacy classes with a single "valid" or "none" class that takes everything data collection did not label as a fallacy. This construction is misleading: a classifier can learn cues that do well on… 24 arXiv — NLP / Computation & Language research 11d ago HearInContext: A Benchmark for Implicit Context in Speech Recognition arXiv:2609.18680v1 Announce Type: new Abstract: Contextual ASR can benefit from semantic cues or from target words explicitly provided in the context. We introduce HearInContext, a Mandarin--English benchmark that pairs shared synthetic speech with assistant replies supporting… 13 arXiv — NLP / Computation & Language research 11d ago A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages arXiv:2609.18739v1 Announce Type: new Abstract: Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance. In this paper, we present a multi-step framework to enhance the annotation… 29 arXiv — NLP / Computation & Language research 11d ago ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts arXiv:2609.18844v1 Announce Type: new Abstract: Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often… 21 arXiv — NLP / Computation & Language research 11d ago How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards arXiv:2609.18908v1 Announce Type: new Abstract: Existing legal benchmarks cover diverse tasks, while continuous monetary remedies remain comparatively underexplored. We introduce ECtHR-NPD, to the best of our knowledge, the first benchmark for predicting non-pecuniary damage… 37 arXiv — NLP / Computation & Language research 11d ago Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking arXiv:2609.18909v1 Announce Type: new Abstract: Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score… 32 arXiv — NLP / Computation & Language research 11d ago A Benchmark Suite and Ground-Truth Methodology for Formal Verification of IEC 61131-3 Ladder Diagram Programs arXiv:2609.18994v1 Announce Type: new Abstract: We present the first benchmark suite for formal verification of Programmable Logic Controller (PLC) programs that combines controlled ground truth with coverage of both textual (Structured Text, ST) and graphical (Ladder Diagram,… 30 arXiv — NLP / Computation & Language research 11d ago WordPolo: Evaluating Language Models Through Iterative Semantic Feedback arXiv:2609.19006v1 Announce Type: new Abstract: Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into the quality or faithfulness of their reasoning processes. We… 23 arXiv — NLP / Computation & Language research 11d ago Benchmarking Large Language Models for Biomedical Relation Extraction arXiv:2609.19071v1 Announce Type: new Abstract: Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid architectures, and state-of-the-art LLMs (Gemini 2.0, OpenAI O-series, Qwen,… 33 arXiv — NLP / Computation & Language research 11d ago Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators arXiv:2609.19072v1 Announce Type: new Abstract: Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks… 31 arXiv — NLP / Computation & Language research 11d ago GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents arXiv:2609.17695v1 Announce Type: cross Abstract: A large language model (LLM) agent can follow more graph paths without acquiring more independent evidence. GraphEcho tests whether agents mistake these repeated encounters for additional corroboration. The benchmark varies path… 13 Hugging Face Daily Papers research 11d ago VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Abstract Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization… 4 NVIDIA Developer Blog official-blog 11d ago TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through... 14 r/MachineLearning community 11d ago GoBench: Evaluating LLMs on the game of Go [R] GoBench evaluates LLMs on 9x9 Go games against a ladder of KataGo opponents, from random to superhuman. It measures general reasoning ability, strongly correlates with ARC-AGI 2 (r=0.83 correlation), and remains highly unsaturated. GPT-6 Astra max achieves 2500 Elo, much lower… 7 r/LocalLLaMA community 11d ago Can current MiniCPM5-2B/any best SOTA under 10B + modern harness beat pre-March 2025 frontier models like Grok 3/GPT-4o? Hey experts! I genuinely have this question and would love a general consensus from people actually using these models. LLMs have advanced a lot in benchmarks and in practical usage. Even GPT-4o and Grok 3 were already enough for general chatting,search lookup, RP, etc. So if… 38 r/LocalLLaMA community 11d ago China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claims — models still lag in some benchmarks but are drastically cheaper to use   submitted by   /u/DustNearby2848 [link]   [comments] 38 r/LocalLLaMA community 11d ago Qwen3.8 Max (0902) scores 45 on the Artificial Analysis Intelligence Index, up 5 points in a month and back on top of China's leaderboard, nosing out GLM-5.3 (44.9) and Kimi K3 (43.8) https://preview.redd.it/17s4r20sgwph1.png?width=1265&format=png&auto=webp&s=1166261eb76a5675c7ff01b673edfe473cda08cc A 30-day upgrade on the 2.4T MoE takes the crown back from GLM-5.3, can't wait for Qwen 4.0   submitted by   /u/UmpireBorn3719 [link]   [comments] 23 r/LocalLLaMA community 12d ago [Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance https://preview.redd.it/e5wn8eyh7vph1.png?width=1080&format=png&auto=webp&s=5faeec866acc9e8eca8dc9e84a6661bee7bcf954 https://preview.redd.it/8ov5gl8j7vph1.png?width=1080&format=png&auto=webp&s=440320ca4646ff7dceaf2aa5ac2ef76faeeba525 New Qwen3.8-Flash-Next quantization using… 23 r/LocalLLaMA community 12d ago qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp time to re-benchmark Qwen Flash Next again   submitted by   /u/jacek2023 [link]   [comments] 24 arXiv — NLP / Computation & Language research 12d ago Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks arXiv:2609.16023v1 Announce Type: new Abstract: Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively, however, is expensive, and rubric-based evaluation has become the… 25 arXiv — NLP / Computation & Language research 12d ago ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian arXiv:2609.16393v1 Announce Type: new Abstract: We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and… 19 arXiv — NLP / Computation & Language research 12d ago RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue arXiv:2609.16614v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain… 15 arXiv — NLP / Computation & Language research 12d ago Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models arXiv:2609.16739v1 Announce Type: new Abstract: Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain… 18 arXiv — NLP / Computation & Language research 12d ago Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion arXiv:2609.16777v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a… 13 arXiv — NLP / Computation & Language research 12d ago RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation arXiv:2609.16900v1 Announce Type: new Abstract: Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling,… 23 arXiv — NLP / Computation & Language research 12d ago Towards Detecting AI-Assisted Responses in Online Surveys arXiv:2609.17317v1 Announce Type: new Abstract: The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey… 17 arXiv — NLP / Computation & Language research 12d ago ECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex Dialogue arXiv:2609.17360v1 Announce Type: new Abstract: Full-duplex spoken dialogue systems must distinguish interruptions that require yielding the floor from backchannels that permit continued speaking. Existing benchmarks typically evaluate events independently and may therefore… 28 arXiv — NLP / Computation & Language research 12d ago Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM arXiv:2609.17435v1 Announce Type: new Abstract: We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the… 25 arXiv — NLP / Computation & Language research 12d ago BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents arXiv:2609.16305v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback. In such settings, safety failures may emerge… 16 arXiv — NLP / Computation & Language research 12d ago Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios arXiv:2609.17056v1 Announce Type: cross Abstract: Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background… 28 r/LocalLLaMA community 12d ago I benchmarked IFM/K2-Horizon-7B on 16GB VRAM After two days of fighting with benchmarking infrastructure, I finally benchmarked IFM/K2-Horizon-7B on 16GB VRAM. TL/DR: it works, but far behind Qwen-3.8-27B. Setup I compared 3 models - Qwen3.8-27B, Ornith-1.5-9B and new IFM/K2-Horizon-7B - on the following setup: All models… 5 r/LocalLLaMA community 12d ago Cut Qwen3.8-27B Reasoning Tokens by 40% -- 3.8 'ThinkingCap' benchmarked! I doubt I'm in the minority here when I say I love Qwen models, but the overthinking is a major timekiller. It was bad in 3.6-27B, and it's worse in 3.8. I know there are some who say, "well that's how it achieves such a good performance/size ratio"... But now there's some… 36 r/MachineLearning community 12d ago TabPFN-3.5 is released as the next SOTA tabular foundation model [N] Prior Labs released their latest tabular foundation model, TabPFN-3.5 today. The model is top of both TabArena and BeyondArena and SOTA for 1M rows and up to 20k features It comes with: - TabPFN-3.5-Fast (in alpha): This one goes 6x faster than the base model -… 11 r/LocalLLaMA community 12d ago ByteShape Qwen 3.8 27B: To KL Diverge or Not to KL Diverge, Part 2: Metric Boogaloo Hey r/LocalLLaMA , We’ve released our full ShapeLearn GGUFs for Qwen 3.8 27B. Blog / Download models TL;DR 3.84 bpw (GPU-5) reaches 99.63% of BF16’s aggregate score of 8 benchmarks, being the most accurate quant we’ve evaluated; 3.23 bpw (GPU-4) reaches 98.72%. These average… 6 r/LocalLLaMA community 13d ago Voodoo Dynamic Quant - Now MIT Licensed Two months ago I announced I had found a new dynamic quant method called Voodoo Quant which was SOTA for the most aggressive quant levels on some smaller Qwen3.5 GGUF models. I kept the methodology private at the time, but I've seen too many requests for dyn quants for various… 6 arXiv — Machine Learning research 13d ago A derivative-fidelity failure mode in physics-informed neural networks: strengthened benchmark evidence from function-value training arXiv:2609.13171v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) use automatic differentiation to impose differential-equation residuals, but good agreement in function values does not necessarily imply accurate derivatives. This paper formulates… 16 arXiv — Machine Learning research 13d ago LLMs or Naive Bayes? Old Gems or New Ways arXiv:2609.13185v1 Announce Type: new Abstract: Large language models (LLMs) prompt a recurring question in research computing: should classical methods like Naive Bayes (NB) be retired? We benchmark Complement Naive Bayes against zero-shot and few-shot LLMs spanning four model… 23 Page 4 of 10 · 500 articles ← Newer Older →