Hugging Face Daily Papers · · 5 min read

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools.</p>\n","updatedAt":"2026-09-10T03:40:18.771Z","author":{"_id":"65fd45473ccf43503350d837","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65fd45473ccf43503350d837/i8r1uUpD0dP1vrrWucO6E.jpeg","fullname":"Haote Yang","name":"Hoter","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8959338665008545},"editors":["Hoter"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65fd45473ccf43503350d837/i8r1uUpD0dP1vrrWucO6E.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06703","authors":[{"_id":"6aa21af0a2aeb74440b1de0e","name":"Yubin Wang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de0f","name":"Xingjian Wei","hidden":false},{"_id":"6aa21af0a2aeb74440b1de10","name":"Jiang Wu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de11","name":"Yinfan Wang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de12","name":"Boyu Zhu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de13","name":"Lin Zhang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de14","name":"Jianing Yu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de15","name":"Huazheng Zeng","hidden":false},{"_id":"6aa21af0a2aeb74440b1de16","name":"Ruiyi Ding","hidden":false},{"_id":"6aa21af0a2aeb74440b1de17","name":"Junyuan Gao","hidden":false},{"_id":"6aa21af0a2aeb74440b1de18","name":"Jiaxing Sun","hidden":false},{"_id":"6aa21af0a2aeb74440b1de19","name":"Lingli Ge","hidden":false},{"_id":"6aa21af0a2aeb74440b1de1a","name":"Haote Yang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de1b","name":"Jingchao Wang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de1c","name":"Aijia Guo","hidden":false},{"_id":"6aa21af0a2aeb74440b1de1d","name":"Qian Jiang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de1e","name":"Yurui Zhao","hidden":false},{"_id":"6aa21af0a2aeb74440b1de1f","name":"Wenjian Zhang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de20","name":"Chen Zhu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de21","name":"Lijun Wu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de22","name":"Xiaolei Yang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de23","name":"Haodong Chen","hidden":false},{"_id":"6aa21af0a2aeb74440b1de24","name":"Junjie Yuan","hidden":false},{"_id":"6aa21af0a2aeb74440b1de25","name":"Zichao Ye","hidden":false},{"_id":"6aa21af0a2aeb74440b1de26","name":"Shaowei Hou","hidden":false},{"_id":"6aa21af0a2aeb74440b1de27","name":"Jing Ye","hidden":false},{"_id":"6aa21af0a2aeb74440b1de28","name":"Jia Yu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de29","name":"Shan Wang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de2a","name":"Lijun Wu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de2b","name":"Jiantao Qiu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de2c","name":"Chao Xu","hidden":false},{"_id":"6aa21af0a2aeb74440b1de2d","name":"Yuqiang Li","hidden":false},{"_id":"6aa21af0a2aeb74440b1de2e","name":"Guangyu Wang","hidden":false},{"_id":"6aa21af0a2aeb74440b1de2f","name":"Bowen Zhou","hidden":false},{"_id":"6aa21af0a2aeb74440b1de30","name":"Dahua Lin","hidden":false},{"_id":"6aa21af0a2aeb74440b1de31","name":"Conghui He","hidden":false}],"publishedAt":"2026-09-06T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents","submittedOnDailyBy":{"_id":"65fd45473ccf43503350d837","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65fd45473ccf43503350d837/i8r1uUpD0dP1vrrWucO6E.jpeg","isPro":false,"fullname":"Haote Yang","user":"Hoter","type":"user","name":"Hoter"},"summary":"High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools. DianShi-RxnDB is available at https://dianshi.opendatalab.org.cn/ .","upvotes":7,"discussionId":"6aa21af0a2aeb74440b1de32","projectPage":"https://dianshi.opendatalab.org.cn/","ai_summary":"DianShi-RxnDB is an automated platform that extracts and normalizes millions of structured organic reactions from patents to support AI-driven chemistry research.","ai_keywords":["DianShi-RxnDB","AI4Chem","organic reaction extraction","patent reaction schemes","normalization pipeline","Model Context Protocol"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6aa225f26ddda571e841b5b4","name":"DianShiGroup","fullname":"DianShi Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65fd45473ccf43503350d837/Zp5cNBRwNo-zrkcSWZpft.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66018b3dbcf0790b8d87168c","avatarUrl":"/avatars/d81f24b72abbb9a9276d595f06516aea.svg","isPro":false,"fullname":"Yubin Wang","user":"yubinwang628","type":"user"},{"_id":"65fd45473ccf43503350d837","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65fd45473ccf43503350d837/i8r1uUpD0dP1vrrWucO6E.jpeg","isPro":false,"fullname":"Haote Yang","user":"Hoter","type":"user"},{"_id":"64203cb61ccd411979d8cb15","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64203cb61ccd411979d8cb15/1eorbN3hH16TwBeOBlbG9.jpeg","isPro":false,"fullname":"R0na1D99","user":"R0na1D99","type":"user"},{"_id":"66728659dc0cf94709b98827","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66728659dc0cf94709b98827/t5CNht2uGSm1AsPgmz2i-.jpeg","isPro":false,"fullname":"Wayne","user":"waynesapphire","type":"user"},{"_id":"6a65b7ac95b6ee04bfdc01e8","avatarUrl":"/avatars/449592b05794e45d9b08484cf197e6a4.svg","isPro":false,"fullname":"RxnOptBench Reviewer Artifact","user":"rxnoptbench-review","type":"user"},{"_id":"6541f9a9eccc4f48dce3d2fb","avatarUrl":"/avatars/04aa1b4268e5537377d75a2914c533bd.svg","isPro":false,"fullname":"zhiyuan zhao","user":"juliozhao","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6aa225f26ddda571e841b5b4","name":"DianShiGroup","fullname":"DianShi Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65fd45473ccf43503350d837/Zp5cNBRwNo-zrkcSWZpft.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06703.md","query":{}}">
Papers
arxiv:2609.06703

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

Published on Sep 6
· Submitted by
Haote Yang
on Sep 10
Authors:
,

Abstract

DianShi-RxnDB is an automated platform that extracts and normalizes millions of structured organic reactions from patents to support AI-driven chemistry research.

High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools. DianShi-RxnDB is available at https://dianshi.opendatalab.org.cn/ .

Community

Paper submitter about 4 hours ago

High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.06703
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.06703 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.06703 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.06703 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers