Discovery happens through repeated rounds of proposing, testing and learning. LDM uses every result to focus the next round, finding promising programs, proteins and molecules with fewer trials.</p>\n","updatedAt":"2026-08-18T08:18:30.467Z","author":{"_id":"66996ea912210698d6fb453b","avatarUrl":"/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg","fullname":"Yihang Chen","name":"scyyc9","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9497265815734863},"editors":["scyyc9"],"editorAvatarUrls":["/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg"],"reactions":[],"isReport":false}},{"id":"6a841abb2802876ed5db4130","author":{"_id":"66f63948793e1463e4097b44","avatarUrl":"/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg","fullname":"Zhongwei Yu","name":"Ease-Onway","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-18T08:41:31.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This paper provides LDM as a unified architecture to combine LLM with Bayesian experimental design, and LLM can further learn experimental design through post training. The initial release provides case studies as proof of concept, showing the effectiveness of LDM in diverse open-ended design problems.","html":"<p>This paper provides LDM as a unified architecture to combine LLM with Bayesian experimental design, and LLM can further learn experimental design through post training. The initial release provides case studies as proof of concept, showing the effectiveness of LDM in diverse open-ended design problems.</p>\n","updatedAt":"2026-08-18T08:41:31.992Z","author":{"_id":"66f63948793e1463e4097b44","avatarUrl":"/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg","fullname":"Zhongwei Yu","name":"Ease-Onway","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.954204797744751},"editors":["Ease-Onway"],"editorAvatarUrls":["/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15669","authors":[{"_id":"6a8413b5b153becad167705c","name":"Zhongwei Yu","hidden":false},{"_id":"6a8413b5b153becad167705d","name":"Yan Song","hidden":false},{"_id":"6a8413b5b153becad167705e","name":"Xue Yan","hidden":false},{"_id":"6a8413b5b153becad167705f","name":"Anjie Liu","hidden":false},{"_id":"6a8413b5b153becad1677060","name":"Xingyu Lu","hidden":false},{"_id":"6a8413b5b153becad1677061","name":"Yihang Chen","hidden":false},{"_id":"6a8413b5b153becad1677062","name":"Huichi Zhou","hidden":false},{"_id":"6a8413b5b153becad1677063","name":"Siyuan Guo","hidden":false},{"_id":"6a8413b5b153becad1677064","name":"Luoyang Sun","hidden":false},{"_id":"6a8413b5b153becad1677065","name":"Sihan Chen","hidden":false},{"_id":"6a8413b5b153becad1677066","name":"Xiangning Yu","hidden":false},{"_id":"6a8413b5b153becad1677067","name":"Jun Wang","hidden":false}],"publishedAt":"2026-08-16T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search","submittedOnDailyBy":{"_id":"66996ea912210698d6fb453b","avatarUrl":"/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg","isPro":false,"fullname":"Yihang Chen","user":"scyyc9","type":"user","name":"scyyc9"},"summary":"Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.","upvotes":22,"discussionId":"6a8413b5b153becad1677068","projectPage":"https://largediscovery.net/","githubRepo":"https://github.com/yzailab/Large-Discovery-Models","githubRepoAddedBy":"user","ai_summary":"A recurrent Large Discovery Model couples generative proposal with a Bayesian non-parametric reward surrogate to guide uncertainty-aware search across molecules, proteins, and programs.","ai_keywords":["Large Discovery Model","recurrent architecture","Bayesian non-parametric reward surrogate","generative model","uncertainty-aware value","discovery memory","antibody design","molecular optimisation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"6a7949a95036d389aeb01b6c","name":"Yangtze-ailab","fullname":"yangtze-ailab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a79450d30eca859f1895e7a/ycrBoZgxQwk2y7rUJwc2K.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66996ea912210698d6fb453b","avatarUrl":"/avatars/d898f7967d4d0785e0c7a1e94b7a237c.svg","isPro":false,"fullname":"Yihang Chen","user":"scyyc9","type":"user"},{"_id":"628c5da32f09ccf530204dbe","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1653366416287-628c5da32f09ccf530204dbe.jpeg","isPro":false,"fullname":"Zhangyue Yin","user":"yinzhangyue","type":"user"},{"_id":"6363a769287b5ce02ed156a4","avatarUrl":"/avatars/72f1ed35ce0c20f566399c7f06261452.svg","isPro":false,"fullname":"wrara","user":"wrawar","type":"user"},{"_id":"645d7f107c7258d904e82749","avatarUrl":"/avatars/a4e9d47b281f18616c522c1a8b8ee7e5.svg","isPro":false,"fullname":"HuichiZhou","user":"Zhouhc","type":"user"},{"_id":"68320962388706bf96e68803","avatarUrl":"/avatars/f174196ad605f49307377d69e2baa550.svg","isPro":false,"fullname":"Leo Chen","user":"darkmoonlight1","type":"user"},{"_id":"66840c818f6c78ebd41f86ca","avatarUrl":"/avatars/626ffa3d262c02f964f15fa420af414c.svg","isPro":false,"fullname":"Kun Shao","user":"ShaoKun-Agent","type":"user"},{"_id":"69f881fe04afb047d6d54dfa","avatarUrl":"/avatars/7d2f96301d04d3bc52e79039666ac8f4.svg","isPro":false,"fullname":"Leo C","user":"xg12138","type":"user"},{"_id":"68abf3aaf626dce772ab8e4b","avatarUrl":"/avatars/1281623f0242f4c2be328849f7195e3c.svg","isPro":false,"fullname":"Memento","user":"AgentFly","type":"user"},{"_id":"65041307324053b21adcf01a","avatarUrl":"/avatars/0377287cdf55c698c660b8a3c79f43c7.svg","isPro":true,"fullname":"Ka Yiu Lee","user":"ycps051031","type":"user"},{"_id":"69f880fdb1e057b04faad6fb","avatarUrl":"/avatars/49a210b6d97e373bde02aa9da2d8a678.svg","isPro":false,"fullname":"nobodynose","user":"bigguy323","type":"user"},{"_id":"66f63948793e1463e4097b44","avatarUrl":"/avatars/f96a4e0f8e6c251c6457ca0f8c1fd156.svg","isPro":false,"fullname":"Zhongwei Yu","user":"Ease-Onway","type":"user"},{"_id":"69254ad9ea206fa9c6fb66f6","avatarUrl":"/avatars/4ff0f222cd69a62c25d02922124fabb4.svg","isPro":false,"fullname":"zl","user":"wangzl98","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a7949a95036d389aeb01b6c","name":"Yangtze-ailab","fullname":"yangtze-ailab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a79450d30eca859f1895e7a/ycrBoZgxQwk2y7rUJwc2K.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.15669.md","query":{}}">
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
Abstract
A recurrent Large Discovery Model couples generative proposal with a Bayesian non-parametric reward surrogate to guide uncertainty-aware search across molecules, proteins, and programs.
Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.
Community
Discovery happens through repeated rounds of proposing, testing and learning. LDM uses every result to focus the next round, finding promising programs, proteins and molecules with fewer trials.
This paper provides LDM as a unified architecture to combine LLM with Bayesian experimental design, and LLM can further learn experimental design through post training. The initial release provides case studies as proof of concept, showing the effectiveness of LDM in diverse open-ended design problems.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.15669 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.15669 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.15669 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.