Benchmark radar is all you need when doing benchmark research!<br>We kept running into new benchmarks while doing benchmark research, so we built a crawler that continuously collects benchmark-related signals from across the web. It pulls evidence from 37 public sources every day, and keeps updating. We also have a CLI tool, which can help you create LaTeX version related benchmark work in minutes from 12k+ benchmark/eval/dataset records.</p>\n","updatedAt":"2026-09-14T03:09:49.021Z","author":{"_id":"674d191fc406e58320e21dc0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674d191fc406e58320e21dc0/3WUBoDbpGxOu5-OJq0bxS.png","fullname":"Koutian Wu","name":"ktwu01","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8639634251594543},"editors":["ktwu01"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/674d191fc406e58320e21dc0/3WUBoDbpGxOu5-OJq0bxS.png"],"reactions":[],"isReport":false}},{"id":"6aa769b380e76656e55d9c88","author":{"_id":"6a9ece96bb45c77270327fb2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a9ece96bb45c77270327fb2/jllrZ6ykXKke1RXBWQl03.png","fullname":"CharlesMarshall","name":"CharlesM1969","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-14T03:27:47.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Interesting and timely work","html":"<p>Interesting and timely work</p>\n","updatedAt":"2026-09-14T03:27:47.782Z","author":{"_id":"6a9ece96bb45c77270327fb2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a9ece96bb45c77270327fb2/jllrZ6ykXKke1RXBWQl03.png","fullname":"CharlesMarshall","name":"CharlesM1969","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9185094833374023},"editors":["CharlesM1969"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6a9ece96bb45c77270327fb2/jllrZ6ykXKke1RXBWQl03.png"],"reactions":[],"isReport":false},"replies":[{"id":"6aa7854c5bf664c961ebdada","author":{"_id":"674d1f45ad20c4bc3db28318","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3XK-5Ioq4dyM4HFZugZzo.png","fullname":"Koutian Wu","name":"ktwu-utexas","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-14T05:25:32.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Thank you Charles!!!","html":"<p>Thank you Charles!!!</p>\n","updatedAt":"2026-09-14T05:25:32.497Z","author":{"_id":"674d1f45ad20c4bc3db28318","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3XK-5Ioq4dyM4HFZugZzo.png","fullname":"Koutian Wu","name":"ktwu-utexas","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.360305517911911},"editors":["ktwu-utexas"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3XK-5Ioq4dyM4HFZugZzo.png"],"reactions":[],"isReport":false,"parentCommentId":"6aa769b380e76656e55d9c88"}}]}],"primaryEmailConfirmed":false,"paper":{"id":"2609.11115","authors":[{"_id":"6aa762ae7ba345d44ad14936","name":"Koutian Wu","hidden":false},{"_id":"6aa762ae7ba345d44ad14937","user":{"_id":"66a2554d345b3106f49b5b86","avatarUrl":"/avatars/c3776b837e0e862c304478dff8d955b0.svg","isPro":false,"fullname":"Junjie Zhou","user":"JunjieZhou","type":"user","name":"JunjieZhou"},"name":"Junjie Zhou","status":"claimed_verified","statusLastChangedAt":"2026-09-14T09:19:12.658Z","hidden":false},{"_id":"6aa762ae7ba345d44ad14938","name":"Ergan Shang","hidden":false},{"_id":"6aa762ae7ba345d44ad14939","name":"Jiayu Wang","hidden":false},{"_id":"6aa762ae7ba345d44ad1493a","name":"Pengqian Han","hidden":false},{"_id":"6aa762ae7ba345d44ad1493b","name":"Junkai Wang","hidden":false},{"_id":"6aa762ae7ba345d44ad1493c","name":"Wanghan Xu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/674d191fc406e58320e21dc0/hkmlO07HPTX_qDkKkenMA.png","https://cdn-uploads.huggingface.co/production/uploads/674d191fc406e58320e21dc0/u8kzvCVvkwL2KUQactJFl.png","https://cdn-uploads.huggingface.co/production/uploads/674d191fc406e58320e21dc0/3eM9RWPPkTZRgIsiOVq5-.png","https://cdn-uploads.huggingface.co/production/uploads/674d191fc406e58320e21dc0/cT_jPJngjgNBFfzSNKb9B.png"],"publishedAt":"2026-09-10T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation","submittedOnDailyBy":{"_id":"674d191fc406e58320e21dc0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674d191fc406e58320e21dc0/3WUBoDbpGxOu5-OJq0bxS.png","isPro":false,"fullname":"Koutian Wu","user":"ktwu01","type":"user","name":"ktwu01"},"summary":"Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.","upvotes":46,"discussionId":"6aa762ae7ba345d44ad1493d","projectPage":"https://benchmark-radar.org","githubRepo":"https://github.com/ktwu01/benchmark-radar","githubRepoAddedBy":"user","ai_summary":"Benchmark Radar is a searchable living database and discovery engine for AI evaluation benchmarks that aggregates sources, score histories, and evidence to support benchmark selection and comparison.","ai_keywords":["large language models (LLMs)","AI benchmarks","benchmark catalog","model cards","score histories","benchmark saturation","Pareto frontier","command-line interface (CLI)"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"674d191fc406e58320e21dc0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674d191fc406e58320e21dc0/3WUBoDbpGxOu5-OJq0bxS.png","isPro":false,"fullname":"Koutian Wu","user":"ktwu01","type":"user"},{"_id":"69414b41d151400e5ba09c9f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/2d7XvG3s1bLUf8_VT2Smj.jpeg","isPro":false,"fullname":"Guangwei Zhang","user":"Changhu1933","type":"user"},{"_id":"672aea8bc7ebfd3d6ddde029","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/gBuchdBJ10AEekAFRqlJh.png","isPro":false,"fullname":"Xu","user":"XuYifanXUXU","type":"user"},{"_id":"679162e2c7f527ef36187078","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/UM8EvYaVh804KOpfNWC5z.png","isPro":false,"fullname":"Zesen Huang","user":"huangzs","type":"user"},{"_id":"6a9ece96bb45c77270327fb2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a9ece96bb45c77270327fb2/jllrZ6ykXKke1RXBWQl03.png","isPro":false,"fullname":"CharlesMarshall","user":"CharlesM1969","type":"user"},{"_id":"640942c83461c51cf73896c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/640942c83461c51cf73896c2/7ex6bXG53cPqGE9gbz2EA.jpeg","isPro":false,"fullname":"HAN PENGQIAN","user":"pengqian","type":"user"},{"_id":"66a1b7635f58258df73318a0","avatarUrl":"/avatars/f62f68ecfb7ae2d373459bb2ca7eeacc.svg","isPro":false,"fullname":"Junhao Hu","user":"DerekHJH","type":"user"},{"_id":"654c3a4009dd7ef52491c080","avatarUrl":"/avatars/c3c14a5e732f7034eb5c50a4a8a47107.svg","isPro":false,"fullname":"Wenjun Feng","user":"USTCKevinF","type":"user"},{"_id":"66a2554d345b3106f49b5b86","avatarUrl":"/avatars/c3776b837e0e862c304478dff8d955b0.svg","isPro":false,"fullname":"Junjie Zhou","user":"JunjieZhou","type":"user"},{"_id":"697f384501d0d68355bcb8c0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/8dP1njmzAC2VlzWKerYLZ.jpeg","isPro":false,"fullname":"JunQiu Zhang","user":"AlanMusk","type":"user"},{"_id":"68380b3bc66376c25e92cb58","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/O8voKNOFGCqOTqTQ_AAhI.png","isPro":false,"fullname":"Jiazhou Xu","user":"StardustXu","type":"user"},{"_id":"6a674400ff63e01ecf8c0120","avatarUrl":"/avatars/b6b3cea42f257a5984039b020e4fb486.svg","isPro":false,"fullname":"Yuhan Lei","user":"AstroHan","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.11115.md","query":{}}">
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Abstract
Benchmark Radar is a searchable living database and discovery engine for AI evaluation benchmarks that aggregates sources, score histories, and evidence to support benchmark selection and comparison.
Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.
Community
Benchmark radar is all you need when doing benchmark research!
We kept running into new benchmarks while doing benchmark research, so we built a crawler that continuously collects benchmark-related signals from across the web. It pulls evidence from 37 public sources every day, and keeps updating. We also have a CLI tool, which can help you create LaTeX version related benchmark work in minutes from 12k+ benchmark/eval/dataset records.
Interesting and timely work
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.