Hugging Face Daily Papers · · 4 min read

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Discovery happens through repeated rounds of proposing, testing and learning. LDM uses every result to focus the next round, finding promising programs, proteins and molecules with fewer trials.</p>\n","updatedAt":"2026-08-18T08:18:30.467Z","author":{"_id":"66996ea912210698d6fb453b","avatarUrl":"/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg","fullname":"Yihang Chen","name":"scyyc9","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9497265815734863},"editors":["scyyc9"],"editorAvatarUrls":["/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg"],"reactions":[],"isReport":false}},{"id":"6a841abb2802876ed5db4130","author":{"_id":"66f63948793e1463e4097b44","avatarUrl":"/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg","fullname":"Zhongwei Yu","name":"Ease-Onway","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-18T08:41:31.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This paper provides LDM as a unified architecture to combine LLM with Bayesian experimental design, and LLM can further learn experimental design through post training. The initial release provides case studies as proof of concept, showing the effectiveness of LDM in diverse open-ended design problems.","html":"<p>This paper provides LDM as a unified architecture to combine LLM with Bayesian experimental design, and LLM can further learn experimental design through post training. The initial release provides case studies as proof of concept, showing the effectiveness of LDM in diverse open-ended design problems.</p>\n","updatedAt":"2026-08-18T08:41:31.992Z","author":{"_id":"66f63948793e1463e4097b44","avatarUrl":"/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg","fullname":"Zhongwei Yu","name":"Ease-Onway","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.954204797744751},"editors":["Ease-Onway"],"editorAvatarUrls":["/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15669","authors":[{"_id":"6a8413b5b153becad167705c","name":"Zhongwei Yu","hidden":false},{"_id":"6a8413b5b153becad167705d","name":"Yan Song","hidden":false},{"_id":"6a8413b5b153becad167705e","name":"Xue Yan","hidden":false},{"_id":"6a8413b5b153becad167705f","name":"Anjie Liu","hidden":false},{"_id":"6a8413b5b153becad1677060","name":"Xingyu Lu","hidden":false},{"_id":"6a8413b5b153becad1677061","name":"Yihang Chen","hidden":false},{"_id":"6a8413b5b153becad1677062","name":"Huichi Zhou","hidden":false},{"_id":"6a8413b5b153becad1677063","name":"Siyuan Guo","hidden":false},{"_id":"6a8413b5b153becad1677064","name":"Luoyang Sun","hidden":false},{"_id":"6a8413b5b153becad1677065","name":"Sihan Chen","hidden":false},{"_id":"6a8413b5b153becad1677066","name":"Xiangning Yu","hidden":false},{"_id":"6a8413b5b153becad1677067","name":"Jun Wang","hidden":false}],"publishedAt":"2026-08-16T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search","submittedOnDailyBy":{"_id":"66996ea912210698d6fb453b","avatarUrl":"/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg","isPro":false,"fullname":"Yihang Chen","user":"scyyc9","type":"user","name":"scyyc9"},"summary":"Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.","upvotes":22,"discussionId":"6a8413b5b153becad1677068","projectPage":"https://largediscovery.net/","githubRepo":"https://github.com/yzailab/Large-Discovery-Models","githubRepoAddedBy":"user","ai_summary":"A recurrent Large Discovery Model couples generative proposal with a Bayesian non-parametric reward surrogate to guide uncertainty-aware search across molecules, proteins, and programs.","ai_keywords":["Large Discovery Model","recurrent architecture","Bayesian non-parametric reward surrogate","generative model","uncertainty-aware value","discovery memory","antibody design","molecular optimisation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"6a7949a95036d389aeb01b6c","name":"Yangtze-ailab","fullname":"yangtze-ailab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a79450d30eca859f1895e7a/ycrBoZgxQwk2y7rUJwc2K.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66996ea912210698d6fb453b","avatarUrl":"/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg","isPro":false,"fullname":"Yihang Chen","user":"scyyc9","type":"user"},{"_id":"628c5da32f09ccf530204dbe","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1653366416287-628c5da32f09ccf530204dbe.jpeg","isPro":false,"fullname":"Zhangyue Yin","user":"yinzhangyue","type":"user"},{"_id":"6363a769287b5ce02ed156a4","avatarUrl":"/avatars/72f1ed35ce0c20f566399c7f06261452.svg","isPro":false,"fullname":"wrara","user":"wrawar","type":"user"},{"_id":"645d7f107c7258d904e82749","avatarUrl":"/avatars/a4e9d47b281f18616c522c1a8b8ee7e5.svg","isPro":false,"fullname":"HuichiZhou","user":"Zhouhc","type":"user"},{"_id":"68320962388706bf96e68803","avatarUrl":"/avatars/f174196ad605f49307377d69e2baa550.svg","isPro":false,"fullname":"Leo Chen","user":"darkmoonlight1","type":"user"},{"_id":"66840c818f6c78ebd41f86ca","avatarUrl":"/avatars/626ffa3d262c02f964f15fa420af414c.svg","isPro":false,"fullname":"Kun Shao","user":"ShaoKun-Agent","type":"user"},{"_id":"69f881fe04afb047d6d54dfa","avatarUrl":"/avatars/7d2f96301d04d3bc52e79039666ac8f4.svg","isPro":false,"fullname":"Leo C","user":"xg12138","type":"user"},{"_id":"68abf3aaf626dce772ab8e4b","avatarUrl":"/avatars/1281623f0242f4c2be328849f7195e3c.svg","isPro":false,"fullname":"Memento","user":"AgentFly","type":"user"},{"_id":"65041307324053b21adcf01a","avatarUrl":"/avatars/0377287cdf55c698c660b8a3c79f43c7.svg","isPro":true,"fullname":"Ka Yiu Lee","user":"ycps051031","type":"user"},{"_id":"69f880fdb1e057b04faad6fb","avatarUrl":"/avatars/49a210b6d97e373bde02aa9da2d8a678.svg","isPro":false,"fullname":"nobodynose","user":"bigguy323","type":"user"},{"_id":"66f63948793e1463e4097b44","avatarUrl":"/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg","isPro":false,"fullname":"Zhongwei Yu","user":"Ease-Onway","type":"user"},{"_id":"69254ad9ea206fa9c6fb66f6","avatarUrl":"/avatars/4ff0f222cd69a62c25d02922124fabb4.svg","isPro":false,"fullname":"zl","user":"wangzl98","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a7949a95036d389aeb01b6c","name":"Yangtze-ailab","fullname":"yangtze-ailab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a79450d30eca859f1895e7a/ycrBoZgxQwk2y7rUJwc2K.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.15669.md","query":{}}">
Papers
arxiv:2608.15669

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

Published on Aug 16
· Submitted by
Yihang Chen
on Aug 18
Authors:
,

Abstract

A recurrent Large Discovery Model couples generative proposal with a Bayesian non-parametric reward surrogate to guide uncertainty-aware search across molecules, proteins, and programs.

Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.

Community

Paper submitter about 2 hours ago

Discovery happens through repeated rounds of proposing, testing and learning. LDM uses every result to focus the next round, finding promising programs, proteins and molecules with fewer trials.

This paper provides LDM as a unified architecture to combine LLM with Bayesian experimental design, and LLM can further learn experimental design through post training. The initial release provides case studies as proof of concept, showing the effectiveness of LDM in diverse open-ended design problems.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.15669
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.15669 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.15669 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.15669 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers