Hugging Face Daily Papers · · 4 min read

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Benchmark LLM Data Preparation ability</p>\n","updatedAt":"2026-07-27T06:59:53.585Z","author":{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","fullname":"Hao Liang","name":"lhpku20010120","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6693871021270752},"editors":["lhpku20010120"],"editorAvatarUrls":["/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.20465","authors":[{"_id":"6a6308f62ee212ed0e2a15fb","user":{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","isPro":false,"fullname":"Hao Liang","user":"lhpku20010120","type":"user","name":"lhpku20010120"},"name":"Hao Liang","status":"claimed_verified","statusLastChangedAt":"2026-07-24T08:45:04.337Z","hidden":false},{"_id":"6a6308f62ee212ed0e2a15fc","name":"Qifeng Cai","hidden":false},{"_id":"6a6308f62ee212ed0e2a15fd","name":"Yibo Lin","hidden":false},{"_id":"6a6308f62ee212ed0e2a15fe","name":"Jianzhuo Du","hidden":false},{"_id":"6a6308f62ee212ed0e2a15ff","name":"Qifeng Xia","hidden":false},{"_id":"6a6308f62ee212ed0e2a1600","name":"Sizhe Qiu","hidden":false},{"_id":"6a6308f62ee212ed0e2a1601","name":"Linzhuang Sun","hidden":false},{"_id":"6a6308f62ee212ed0e2a1602","name":"Meiyi Qiang","hidden":false},{"_id":"6a6308f62ee212ed0e2a1603","name":"Zhaoyang Han","hidden":false},{"_id":"6a6308f62ee212ed0e2a1604","name":"Xiaochen Ma","hidden":false},{"_id":"6a6308f62ee212ed0e2a1605","name":"Bohan Zeng","hidden":false},{"_id":"6a6308f62ee212ed0e2a1606","name":"Ruichuan An","hidden":false},{"_id":"6a6308f62ee212ed0e2a1607","name":"Conghui He","hidden":false},{"_id":"6a6308f62ee212ed0e2a1608","name":"Wentao Zhang","hidden":false}],"publishedAt":"2026-05-19T00:00:00.000Z","submittedOnDailyAt":"2026-07-27T00:00:00.000Z","title":"DataPrep-Bench: Benchmarking LLMs as Training Data Preparators","submittedOnDailyBy":{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","isPro":false,"fullname":"Hao Liang","user":"lhpku20010120","type":"user","name":"lhpku20010120"},"summary":"The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM-driven data preparation as comprising two complementary capabilities: data construction, which transforms raw sources into supervised training data, and data quality evaluation, which predicts the training value of candidate datasets before downstream training; throughout, \"quality\" refers to downstream training utility rather than surface-level textual properties. We introduce DataPrep-Bench, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains and multiple base models. For data construction, methods consume identical raw sources and are scored by fine-tuning a base model on their outputs jointly with Dolly-15k; alongside this track we release Data-Construction-Skill, a skill-guided agent that lifts the Dolly-only baseline by nearly 20 points absolute on Llama-3.1-8B Finance and is competitive with the strongest agent- and DataFlow-based methods in knowledge-extraction-dense domains. For data quality evaluation, scoring functions are scored by Pearson correlation with downstream performance on a shared candidate pool; we release the Distributional Alignment Score (DAS), a distribution-based evaluator that uses MMD between a candidate dataset and a domain proxy. DAS attains the strongest cross-model correlation in four of six domains and is the only metric clearing r > 0.70 simultaneously in Math, Science, and Medical, outperforming existing quality-, diversity-, and heuristic-based evaluators. DataPrep-Bench provides a unified, downstream-grounded framework for measuring progress on both capabilities as co-equal targets of LLM-driven data preparation.","upvotes":14,"discussionId":"6a6308f62ee212ed0e2a1609","projectPage":"https://datapreparationbench.github.io/","githubRepo":"https://github.com/OpenDCAI/Data-Preparation-Bench","githubRepoAddedBy":"user","githubStars":218},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","isPro":false,"fullname":"Hao Liang","user":"lhpku20010120","type":"user"},{"_id":"6217599529500f41901123f8","avatarUrl":"/avatars/8a0fe54e53fe6527c70a78598a0cd941.svg","isPro":false,"fullname":"Hao Liang","user":"lhbit20010120","type":"user"},{"_id":"66ac9567c97d2f0c88c3ac72","avatarUrl":"/avatars/14df8b5eed4ea756c93f61999c75e44f.svg","isPro":false,"fullname":"PKU_Baichuan","user":"PKU-Baichuan","type":"user"},{"_id":"670cd1d3d526bc93f9c6137b","avatarUrl":"/avatars/0c41a6a31102101cfe2a75f81bcd7ecd.svg","isPro":false,"fullname":"xu chang","user":"xccr","type":"user"},{"_id":"6a3a27a7eacedbe8ba163c1a","avatarUrl":"/avatars/8064c54ad12f00bf0bd47499e42f0370.svg","isPro":false,"fullname":"lin","user":"kalelee","type":"user"},{"_id":"65b7098af327f1f4e315294d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b7098af327f1f4e315294d/38JtHebJ_XYWciMA15vJ3.jpeg","isPro":false,"fullname":"Runming He","user":"blackBOX25I47","type":"user"},{"_id":"6618a60721d5003025004c96","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6618a60721d5003025004c96/_9DZG3lKbIt5KOj4WxsIJ.jpeg","isPro":false,"fullname":"Meiyi Qiang","user":"MeiyiQiang","type":"user"},{"_id":"65099d08f37afbab0d3fb268","avatarUrl":"/avatars/cef45b7c6b7c90bbef341a39a9bb51be.svg","isPro":false,"fullname":"Xiaochen Ma","user":"Sunnyhaze","type":"user"},{"_id":"68400c7b50cb0ac62e5fd9f2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68400c7b50cb0ac62e5fd9f2/UqFfQFbFsCxjLIcwIwdFx.png","isPro":false,"fullname":"Qihan Lin","user":"tunaaa126","type":"user"},{"_id":"67ca931e163cf5cf898e49b3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ca931e163cf5cf898e49b3/Fh7RsNtPSM2iBSY4dJZOW.jpeg","isPro":false,"fullname":"scuuy","user":"scuuy666","type":"user"},{"_id":"6880e7df575954f53124c68c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/NG-VcF3FvLUxmVkAUTu9w.png","isPro":false,"fullname":"Qifeng Xia","user":"PiarPP","type":"user"},{"_id":"65536513b052cff48e60dfd9","avatarUrl":"/avatars/5fb4f634eca7764900e1cae972c8dc44.svg","isPro":false,"fullname":"mike stone","user":"debugger123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.20465.md","query":{}}">
Papers
arxiv:2607.20465

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

Published on May 19
· Submitted by
Hao Liang
on Jul 27
#3 Paper of the day
Authors:

Abstract

The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM-driven data preparation as comprising two complementary capabilities: data construction, which transforms raw sources into supervised training data, and data quality evaluation, which predicts the training value of candidate datasets before downstream training; throughout, "quality" refers to downstream training utility rather than surface-level textual properties. We introduce DataPrep-Bench, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains and multiple base models. For data construction, methods consume identical raw sources and are scored by fine-tuning a base model on their outputs jointly with Dolly-15k; alongside this track we release Data-Construction-Skill, a skill-guided agent that lifts the Dolly-only baseline by nearly 20 points absolute on Llama-3.1-8B Finance and is competitive with the strongest agent- and DataFlow-based methods in knowledge-extraction-dense domains. For data quality evaluation, scoring functions are scored by Pearson correlation with downstream performance on a shared candidate pool; we release the Distributional Alignment Score (DAS), a distribution-based evaluator that uses MMD between a candidate dataset and a domain proxy. DAS attains the strongest cross-model correlation in four of six domains and is the only metric clearing r > 0.70 simultaneously in Math, Science, and Medical, outperforming existing quality-, diversity-, and heuristic-based evaluators. DataPrep-Bench provides a unified, downstream-grounded framework for measuring progress on both capabilities as co-equal targets of LLM-driven data preparation.

Community

Paper author Paper submitter about 6 hours ago

Benchmark LLM Data Preparation ability

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.20465
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.20465 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.20465 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.20465 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers