News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID arXiv:2609.30287v1 Announce Type: new Abstract: AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support… 27 arXiv — NLP / Computation & Language research 7h ago A Benchmark Framework for Screening Automation in Systematic Reviews arXiv:2609.30298v1 Announce Type: new Abstract: Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by… 20 arXiv — NLP / Computation & Language research 7h ago The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge arXiv:2609.30604v1 Announce Type: new Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations,… 31 arXiv — NLP / Computation & Language research 7h ago From annotation to reasoning: Culture in language models arXiv:2609.30897v1 Announce Type: new Abstract: How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave… 30 arXiv — NLP / Computation & Language research 7h ago AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth arXiv:2609.30483v1 Announce Type: cross Abstract: Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against… 4 arXiv — NLP / Computation & Language research 7h ago JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models arXiv:2609.31142v1 Announce Type: cross Abstract: Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a… 33 arXiv — NLP / Computation & Language research 7h ago PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents arXiv:2609.31468v1 Announce Type: cross Abstract: LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean… 19 arXiv — NLP / Computation & Language research 7h ago Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer arXiv:2609.31587v1 Announce Type: cross Abstract: We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether… 14 r/LocalLLaMA community 20h ago We released VeriLoop E2 (27B, Apache-2.0). The design question behind it: should an LLM be allowed to commit its own state? Disclosure: I’m one of the authors of VeriLoop E2, a 27B model post-trained from Qwen3.8-27B. The weights are Apache-2.0; the harness we evaluate it in is not open source (details at the bottom). This is a release post, but rather than a benchmark dump I want to talk about the… 28 r/LocalLLaMA community 1d ago Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_nonogram_logic_puzzle/ . That thread shaped v1.2: Reasoning effort is explicit per run Every prompt and output is public. All current top ranking private and open weight… 7 r/MachineLearning community 1d ago Publication potential [D] Hi all, I may be being silly but I have just finished my masters thesis which benchmarked three deep learning architectures (one of which is novel), across varying preprocessing pipelines for the purpose of EEG motor imagery task classification SPECIFICALLY on a consumer grade… 17 r/LocalLLaMA community 1d ago For the longest time I’ve felt this sub should have a pinned section where a detailed post about each model should get featured. For instance whenever a model comes out, what’s the best engine to run it, the best harness and absolute minimum you need to get same or near same re results that the benchmark of that model claims. And whenever a quant from Unsloth guys comes out the guide can either be updated… 30 r/LocalLLaMA community 1d ago Qwen3.8-27B IQ3_XXS vs Qwen3.6-35B-A3B Q4_K_M Which one is better for difficult tasks like web scrapping, coding, using tools? Looking for any benchmarks because i couldn't actually find one after quite some digging   submitted by   /u/Loose_Doubt367 [link]   [comments] 30 r/LocalLLaMA community 1d ago My Reading Library: Evaluating LLMs on Android Tasks Can LLM agents actually get through a day in the life of a normal user? That question got me reading papers on Android agents and mobile benchmarks over the past few months. A few patterns kept showing up: Most benchmarks run on emulators, making real-device metrics difficult to… 9 r/LocalLLaMA community 3d ago Has anyone benchmarked AI agents against the SOLIDWORKS CSWA exam? Would be interesting right? Models are starting to score higher and higher on benchmarks like Parametric CAD Bench , but can they pass an actual exam? The Certified SOLIDWORKS Associate (CSWA) exam might be an interesting place to start. They have an sample exam on their… 6 r/LocalLLaMA community 3d ago Did anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some I let Codex do some eval on Qwen 3.8 Flash Next IQ4_XS (served via vllm and r9v) and Qwen 3.8 FP8 (served via vllm and radiance). Here are the results: Benchmark Flash-Next IQ4_XS Qwen3.8 27B FP8 Result MMLU-Pro 83.8% 75.0% Flash-Next GPQA Diamond 42.5% 27.5% Flash-Next GSM8K… 8 arXiv — Machine Learning research 3d ago fable.intermittent: benchmarking probabilistic forecasting methods for intermittent time series arXiv:2609.28607v1 Announce Type: new Abstract: Intermittent time series are common in spare-parts demand and retail sales. Since the cost of forecast errors is typically asymmetric, decisions such as inventory control require the full predictive distribution rather than a point… 17 arXiv — Machine Learning research 3d ago Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER arXiv:2609.29101v1 Announce Type: new Abstract: Medical Named Entity Recognition (NER) for low-resource languages remains a challenging task due to high linguistic variability and a scarcity of domain-specific annotated corpora. This work presents a comprehensive empirical… 31 arXiv — Machine Learning research 3d ago When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection arXiv:2609.29580v1 Announce Type: new Abstract: A released table is often treated as an i.i.d. sample, although its repeated rows may encode business frequency, repeated entities, joins, resampling, or extraction errors. We show that this ambiguity creates a hidden measurement… 35 arXiv — Machine Learning research 3d ago Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets arXiv:2609.29625v1 Announce Type: new Abstract: Across seven public educational prediction datasets, three passed all four pre-modeling reliability checks; the remaining four either failed group-aware generalization tests or lacked the provenance metadata needed to run them. One… 28 arXiv — Machine Learning research 3d ago TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening arXiv:2609.29740v1 Announce Type: new Abstract: Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and… 25 arXiv — NLP / Computation & Language research 3d ago Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks arXiv:2609.28673v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on… 23 arXiv — NLP / Computation & Language research 3d ago COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages arXiv:2609.28826v1 Announce Type: new Abstract: Machine translation (MT) for Indian languages remains constrained by the limited availability of high-quality, Indic-centric parallel corpora and evaluation benchmarks. Existing multilingual resources are largely constructed from… 24 arXiv — NLP / Computation & Language research 3d ago Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms arXiv:2609.29001v1 Announce Type: new Abstract: Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two… 8 arXiv — NLP / Computation & Language research 3d ago BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech arXiv:2609.29371v1 Announce Type: new Abstract: This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining… 8 arXiv — NLP / Computation & Language research 3d ago Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure arXiv:2609.29445v1 Announce Type: new Abstract: We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be… 9 arXiv — NLP / Computation & Language research 3d ago Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark arXiv:2609.29479v1 Announce Type: new Abstract: Prospective clinical actions, the follow-ups, orders, referrals, and instructions that deter-mine what happens to a patient next, are annotated today in thin fragments across incom-patible corpora: each records a text span and one… 11 arXiv — NLP / Computation & Language research 3d ago PROOF: Profiling Reliability of Object-Level Facts in Large Language Models arXiv:2609.29504v1 Announce Type: new Abstract: Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for… 14 arXiv — NLP / Computation & Language research 3d ago EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation arXiv:2609.29511v1 Announce Type: new Abstract: Machine Translation (MT) for low-resource languages remains far behind that of high-resource languages, and the gap is widest in specialised domains, where parallel data is scarce or entirely absent. We present EnSiTa, a trilingual… 27 arXiv — NLP / Computation & Language research 3d ago Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap arXiv:2609.29559v1 Announce Type: new Abstract: Arabic-Russian machine translation (MT) remains under-explored due to the rich morphology of Arabic and low lexical overlap between the two languages. We benchmark seven fine-tuned neural machine translation (NMT) models against… 11 arXiv — NLP / Computation & Language research 3d ago ModularSQL: A Runtime Guardrail for the Multiplicity Blind Spot in Text-to-SQL arXiv:2609.29573v1 Announce Type: new Abstract: Text-to-SQL systems are increasingly deployed on production databases, where queries that pass benchmark evaluation can still produce results that distort downstream workflows. Standard set-based execution accuracy (Set-EX)… 36 arXiv — NLP / Computation & Language research 3d ago TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) arXiv:2609.29733v1 Announce Type: new Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles. While these systems achieve state-of-the-art performance, their applicability and… 32 arXiv — NLP / Computation & Language research 3d ago Benchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages arXiv:2609.29798v1 Announce Type: new Abstract: This paper presents an end-to-end study of automatic speech recognition (ASR) for adolescent health communication in three Ghanaian languages (Twi, Dagbani, and Ewe). The work proceeds in three connected stages; First, we benchmark… 33 arXiv — NLP / Computation & Language research 3d ago Multi-Task Learning by using Contextualized Word Representations for Syntactic Parsing of a Morphologically Rich Language arXiv:2609.29855v1 Announce Type: new Abstract: We address the challenge of syntactic parsing for Urdu, a morphologically rich language, and present state-of-the-art results for both constituency and dependency parsing. This paper offers four major contributions: 1) the… 9 arXiv — NLP / Computation & Language research 3d ago Artificial Societies Benchmark: A Validation Framework for Synthetic Research arXiv:2609.30030v1 Announce Type: new Abstract: A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to… 20 arXiv — NLP / Computation & Language research 3d ago Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS arXiv:2609.28988v1 Announce Type: cross Abstract: We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer… 27 arXiv — NLP / Computation & Language research 3d ago Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots arXiv:2609.29043v1 Announce Type: cross Abstract: General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to interpret diverse natural language commands and generate multi-step action sequences in real home environments. Conventional… 6 r/LocalLLaMA community 3d ago normalize benchmarks from different time period LiveBench has benchmark snapshots from different points in time. Could someone run an agent to normalize the values across these snapshots so we can compare model strength consistently from 2024 through 2026? Right now, it’s difficult to make meaningful comparisons across the… 18 r/LocalLLaMA community 3d ago 7900 XTX — two "low-thinking" Qwen 3.8 27B quants (Swift + ThinkingCap) vs the regular quant First, do they actually produce less tokens? Yes. Total tokens per benchmark run (4 scenarios): base quant ~66k, ThinkingCap ~49k (−26%), Swift ~45k (−33%). So the "less thinking" is real — and Swift cuts the most. Then the cost: and this is where it got interesting. The two… 29 r/LocalLLaMA community 3d ago ThinkingCap 3.8-27B vs. Swift 3.8-27B vs. Qwen 3.8-27B Benchmarks With the release of ThinkingCap-Qwen3.8-27B , I thought it would be worthwhile to do a comparison between the original Qwen3.8-27B, the new ThinkingCap, and Swift-Qwen3.8-27B . Both Swift which I already reviewed , and ThinkingCap do exactly the same thing: they reduce the… 8 r/LocalLLaMA community 4d ago Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark. I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all… 24 r/LocalLLaMA community 4d ago Folks, have you purchased the Mac M5 Ultra with 256GB yet? We need serious benchmarks, because we only get YouTube clowns influencers results Ok, so the Mac M5 Ultra (256GB) hit the market, but the only publicly available benchmarks material are flashy YouTube "clown influencers" videos. We need serious numbers to evaluate whether Apple’s silicon can actually compete with Nvidia’s current GPU‑centric workflows or not.… 11 arXiv — Machine Learning research 4d ago What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus arXiv:2609.26826v1 Announce Type: new Abstract: Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing… 13 arXiv — Machine Learning research 4d ago COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation arXiv:2609.26853v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing… 5 arXiv — Machine Learning research 4d ago QUARTET: Quad-branch cross-Attention and Random-walk Traces for Enhancing Transformers on Relational Graphs arXiv:2609.26855v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-table databases as heterogeneous temporal graphs, and graph transformers currently achieve state-of-the-art performance on benchmarks like RelBench. However, the current leading model,… 20 arXiv — Machine Learning research 4d ago An open benchmark for machine learning-based polymer property prediction arXiv:2609.27036v1 Announce Type: new Abstract: Polymer property prediction lacks open, standardized benchmarks that enable rigorous comparison of machine-learning methods, with existing resources covering only a narrow fraction of polymer architectures, such as homopolymers. We… 16 arXiv — Machine Learning research 4d ago A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems arXiv:2609.27201v1 Announce Type: new Abstract: Sequential RecSys are central to modern personalization, exploiting user's historical interaction sequences to drive next-step decisions. Deep learning models, particularly CNN and Transformer-based architectures, have proven… 32 arXiv — Machine Learning research 4d ago Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel arXiv:2609.27867v1 Announce Type: new Abstract: A forecasting benchmark reports which method won. We show that the answer is set by the evaluator's choices before any model is fitted. We benchmark 24 forecasting methods and one textbook reference, including six 2025-era time… 26 arXiv — NLP / Computation & Language research 4d ago Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms arXiv:2609.26942v1 Announce Type: new Abstract: Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weight LLMs to generate kinship… 7 arXiv — NLP / Computation & Language research 4d ago Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court arXiv:2609.26945v1 Announce Type: new Abstract: Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in… 19 Page 1 of 10 · 500 articles Older →