News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 4d ago Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure arXiv:2609.27176v1 Announce Type: new Abstract: Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a… 28 arXiv — NLP / Computation & Language research 4d ago Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach arXiv:2609.27205v1 Announce Type: new Abstract: Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the first grapheme-to-phoneme (G2P)… 17 arXiv — NLP / Computation & Language research 4d ago Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models arXiv:2609.27372v1 Announce Type: new Abstract: Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset,… 28 arXiv — NLP / Computation & Language research 4d ago PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models arXiv:2609.27395v1 Announce Type: new Abstract: Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number,… 15 arXiv — NLP / Computation & Language research 4d ago Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models arXiv:2609.27510v1 Announce Type: new Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is… 30 arXiv — NLP / Computation & Language research 4d ago Evaluating Open-Weight LLMs for Turkish Domain Documents Under Retrieval and Hardware Constraints arXiv:2609.28007v1 Announce Type: new Abstract: Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document… 33 arXiv — NLP / Computation & Language research 4d ago Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing arXiv:2609.28026v1 Announce Type: new Abstract: We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined… 18 arXiv — NLP / Computation & Language research 4d ago Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark arXiv:2609.28090v1 Announce Type: new Abstract: Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds… 4 arXiv — NLP / Computation & Language research 4d ago Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following arXiv:2609.28395v1 Announce Type: new Abstract: Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings… 34 arXiv — NLP / Computation & Language research 4d ago What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit arXiv:2609.27408v1 Announce Type: cross Abstract: Benchmarks for vision-language models offer their answer choices in some convention: a letter, a color name, a pixel coordinate. That convention is treated as neutral. We find it is not, and that the limits a benchmark reports… 26 Don't Worry About the Vase community 4d ago Claude Opus 5.5: The System Card Introducing the world’s most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5. 7 r/LocalLLaMA community 4d ago M5U base 96GB inference numbers for Q3.8FN after 112M tokens TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted So the good news is that I got the base model on launch day with only 64 core GPU. All benchmarks are for current maxed out model,… 34 r/LocalLLaMA community 4d ago MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam MiMo-V2.6-Pro has an insanely high score of 46 on AA, putting it at the head of the opensource models available. It also costs pennies. Flash is not out on AA yet, but it costs less than half on datacenter and is slightly below on Xiaomi's own benchmarks. It also fits in 192GB,… 12 OpenAI official-blog 5d ago Introducing MentalHealthBench MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations. 7 The Information — AI news-outlet 5d ago Shares of Chinese AI Model Firms Fall After Report of Regulatory Probe Hong Kong-listed shares of Z.ai and MiniMax plunged on Wednesday, after The Information reported that China’s internet regulator is probing potential data leaks to Anthropic. Z.ai, also known as Zhipu, dropped 12.4%, MiniMax fell 4% and Alibaba declined 4.4%. The benchmark Hang… 37 TechCrunch — AI news-outlet 5d ago “We’re already fighting yesterday’s battle”: Greece’s prime minister gets candid about AI Most leaders on a trade mission stick to the pitch, but when I interviewed Greek Prime Minister Kyriakos Mitsotakis this week, he also admitted that no government is ready for what AI is about to do. 7 arXiv — NLP / Computation & Language research 5d ago Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers arXiv:2609.25237v1 Announce Type: cross Abstract: Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and… 32 arXiv — NLP / Computation & Language research 5d ago Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices arXiv:2609.25645v1 Announce Type: cross Abstract: Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval,… 32 arXiv — Machine Learning research 5d ago Can You Delete a Year of Market Data? Machine Unlearning Against Exact Retraining Oracles arXiv:2609.26242v1 Announce Type: new Abstract: When a data license expires, deleting stored records does not remove influence encoded in a trained forecaster. Machine unlearning seeks to remove this influence without retraining. We benchmark temporal unlearning with 3,200… 18 arXiv — NLP / Computation & Language research 5d ago Peerify: Benchmarking Peer-Review Claim Verification arXiv:2609.25046v1 Announce Type: new Abstract: Peer review plays a central role in scholarly publishing, yet verifying whether reviewer claims are supported by manuscript evidence remains a largely manual and time-consuming process. We present Peerify, a pipeline for… 36 arXiv — NLP / Computation & Language research 5d ago FrontierMath Erd\H{o}s arXiv:2609.25050v1 Announce Type: new Abstract: We introduce FrontierMath Erd\H{o}s (FME), a benchmark of 68 Erd\H{o}s problems that are open as of August 2026. To solve a task in FME, AI systems must resolve (prove or disprove) one of the 68 conjectures in the proof assistant… 34 arXiv — NLP / Computation & Language research 5d ago FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability arXiv:2609.25192v1 Announce Type: new Abstract: Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition… 35 arXiv — NLP / Computation & Language research 5d ago FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing arXiv:2609.25298v1 Announce Type: new Abstract: Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally… 32 arXiv — NLP / Computation & Language research 5d ago Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development arXiv:2609.25396v1 Announce Type: new Abstract: Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for… 11 arXiv — NLP / Computation & Language research 5d ago SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children arXiv:2609.26090v1 Announce Type: new Abstract: Language is the target of most early intervention for autistic children. Because the goal and the method change from child to child, the work falls to a teacher who takes one child at a time and judges each scene as it unfolds.… 18 arXiv — NLP / Computation & Language research 5d ago Differentiable Fuzzy Inference Layer: A Monotone, Compositional Ordinal Reasoning Head for Large Language Models arXiv:2609.26113v1 Announce Type: new Abstract: A state-of-the-art language model asked to interpret "most of most students passed" typically answers "most," though composing two instances of "most" yields a proportion closer to "some." We trace this failure to an architectural… 36 arXiv — NLP / Computation & Language research 5d ago Modality-Gated Deep Adapters: Adding a Modality to a Frozen Embedding Model with Exact Preservation arXiv:2609.26182v1 Announce Type: new Abstract: Multimodal embedding models are deployed at scale: retrieval indices, benchmark results, and behavioral audits all depend on the base model's exact outputs. Extending such a model to a new modality with existing parameter-efficient… 29 arXiv — NLP / Computation & Language research 5d ago Beyond Short Segments : Expanding Speaker Embeddings with Vector Archives arXiv:2609.25007v1 Announce Type: cross Abstract: The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information. To address this critical challenge, we propose the Vector Archive… 18 r/LocalLLaMA community 5d ago New 6B image model coming, AntLing just open sourced the Ming-Image-0.1-Design family • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard.… 12 r/MachineLearning community 5d ago LinearSolveBench: new benchmark for linear solvers [P] LinearSolverBench measures the ability of a model or harness to write fast, accurate, and general numerical solvers for large sparse linear systems in C. The goal is to encourage algorithmic advances in numerical methods for solving linear systems of equations.… 27 r/MachineLearning community 5d ago QontoFAQ: A better Information Retrieval Benchmark [R] Retrieval benchmarks sometimes feel benchmaxxed by models, so we wanted to find a way to tie it as close as possible to my objective: finding the article that answers a product question right . We worked on a new metric which seems more proportional to document relevance, and… 8 r/LocalLLaMA community 6d ago Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed with half the context on the same hardware. I benchmarked a couple dozen quants, vllm, sglang, llama.cp and benchmarked settings and… 26 arXiv — Machine Learning research 6d ago PAGE: Partition-Aware Gated KV-Cache Eviction arXiv:2609.22157v1 Announce Type: new Abstract: KV-cache eviction methods decide which tokens to keep but not whether to evict at all, so a benchmark mean can hide a class of inputs on which compression drives accuracy from 99\% to 0\%. We reframe eviction as a per-input… 6 arXiv — Machine Learning research 6d ago CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds arXiv:2609.22183v1 Announce Type: new Abstract: Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown. We present CleanScore, a black-box audit using scored outputs only. Each benchmark question… 36 arXiv — Machine Learning research 6d ago Dissecting Hierarchical Reasoning Models: A Mechanistic Study arXiv:2609.22197v1 Announce Type: new Abstract: We study Hierarchical Reasoning Model (HRM), a representative hierarchical Transformer-based latent reasoning model with many variants, on Sudoku, Maze, and ARC-AGI-2. We mechanistically understand how HRM reasons and what… 21 arXiv — Machine Learning research 6d ago The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families arXiv:2609.22216v1 Announce Type: new Abstract: Quantization enables deployment of large language models on resource-constrained clinical edge devices, but its effect on clinical accuracy and safety remains understudied. We evaluate five 7-8B parameter models at FP16, GPTQ-INT8,… 37 arXiv — Machine Learning research 6d ago Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles arXiv:2609.22220v1 Announce Type: new Abstract: Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these… 38 arXiv — Machine Learning research 6d ago Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark arXiv:2609.22222v1 Announce Type: new Abstract: Large language models can generate executable data-analysis code, but successful execution is not equivalent to a valid official-statistics result. This study asks whether authoritative metadata and execution feedback improve the… 24 arXiv — Machine Learning research 6d ago Benchmarking Hybrid Deep Learning Architectures for Predictive Maintenance in Industry 4.0 arXiv:2609.22583v1 Announce Type: new Abstract: Predictive maintenance in Industry 4.0 refers to using data from sensors, machines, and production systems to estimate when equipment is likely to fail, so maintenance can be planned before a breakdown occurs [1]. However, a model… 27 arXiv — Machine Learning research 6d ago Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark arXiv:2609.22584v1 Announce Type: new Abstract: As industrial processes grow in complexity, traditional Proportional-Integral-Derivative (PID) controllers are often insufficient for handling their non-linear, multi-input dynamics. We propose using advanced Deep Reinforcement… 26 arXiv — NLP / Computation & Language research 6d ago Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents arXiv:2609.22090v1 Announce Type: new Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM… 17 arXiv — NLP / Computation & Language research 6d ago An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents arXiv:2609.22114v1 Announce Type: new Abstract: Context compression is widely proposed as a way to cut the token bill of LLM coding agents, and public benchmarks report that aggressive compression preserves task-solving quality. These two facts do not imply the third one… 16 arXiv — NLP / Computation & Language research 6d ago Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models arXiv:2609.22152v1 Announce Type: new Abstract: Imagination performs as a high-level function of large language models (LLMs) which determines the potential of how an LLM creates unseen or creative content. While existing works have built a rich family of creativity benchmarks… 14 arXiv — NLP / Computation & Language research 6d ago PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations arXiv:2609.22200v1 Announce Type: new Abstract: LLM assistants and agentic systems log long multi-turn conversations. AI providers often scan these conversations for Personally Identifiable Information (PII) and mask the PII before storing or processing conversation data. Yet… 4 arXiv — NLP / Computation & Language research 6d ago Which Part of the Context Layer Does the Work? Separating Semantic Content from Retrieval Scaffolding in Text-to-SQL Agents arXiv:2609.22259v1 Announce Type: new Abstract: Context layers, curated documentation that an analytics agent fetches at query time, produce large accuracy gains on text-to-SQL benchmarks. A with/without comparison cannot say which part of the layer does the work: the semantic… 6 arXiv — NLP / Computation & Language research 6d ago Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models arXiv:2609.22452v1 Announce Type: new Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio,… 33 arXiv — NLP / Computation & Language research 6d ago Beyond Final-Token Classification: Heterogeneous Readouts for Evidence-Grounded Suicide Risk Detection arXiv:2609.22767v1 Announce Type: new Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label psychosocial factor detection, and extraction of supporting phrases. We introduce… 4 arXiv — NLP / Computation & Language research 6d ago MIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills Assessment arXiv:2609.22778v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require expert judgment remains unclear. We investigate this challenge in the context of… 15 Hugging Face official-blog 6d ago How UK AISI and EvalEval Are Making Benchmark Results Reproducible Back to Articles a]:hidden"> How UK AISI and EvalEval Are Making Benchmark Results Reproducible Published September 22, 2026 Update on GitHub Upvote 1 Avijit Ghosh evijit evaleval Jenny Chim j-chim evaleval Deep Joshi deeplumiere evaleval Srishti srishtiy evaleval Matt Kennedy… 26 r/MachineLearning community 6d ago Jev's calibration was measured. The LLMs won [D] Source: Jev Benchmarks Its training method is literally called "Reinforcement Learning for Calibrated Decisions." Calibration gap vs human labels (lower = better): Yes/no: Jev 5.0, Gemini 3.8 Flash 2.0 Pick-one: Jev 9.8, DeepSeek V4.1 Flash 2.8 Rubric: Jev 19.7, GLM-5.3 12.9 It… 29 Page 2 of 10 · 500 articles ← Newer Older →