We introduce a large-scale framework for measuring how LLM-generated research ideas differ from human research ideas. We study ideation as a distributional alignment problem: given the same local literature context, do LLMs identify the same kinds of opportunities and construct the same kinds of contributions as researchers? We build from 11.7K papers across ML conferences and Nature Communications, reverse-engineering proximal prior works for each paper and extracting the human idea as a motivation–method pair. We then prompt nine LLMs to generate ideas from the same context and annotate both human and model ideas with a two-axis research-taste taxonomy covering opportunity patterns and method paradigms. Across models and domains, we find a stable gap that LLM ideas concentrate heavily on bridge-like motivations and synthesis methods, while human papers span a broader range research topic. Reasoning and richer full-paper context do not close this gap. Reasoning even sharpens the template. These results suggest that future AI ideation systems should optimize not only individual idea quality, but also diversity of research taste.</p>\n","updatedAt":"2026-07-06T17:17:56.500Z","author":{"_id":"666710655b45380937a5d472","avatarUrl":"/avatars/6e41eb68519c6bb99567af4ef54df2cf.svg","fullname":"Ziyu","name":"ziyuuc","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.875653862953186},"editors":["ziyuuc"],"editorAvatarUrls":["/avatars/6e41eb68519c6bb99567af4ef54df2cf.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.01233","authors":[{"_id":"6a4717796ee372f6920de29e","user":{"_id":"666710655b45380937a5d472","avatarUrl":"/avatars/6e41eb68519c6bb99567af4ef54df2cf.svg","isPro":false,"fullname":"Ziyu","user":"ziyuuc","type":"user","name":"ziyuuc"},"name":"Ziyu Chen","status":"claimed_verified","statusLastChangedAt":"2026-07-05T21:10:39.965Z","hidden":false},{"_id":"6a4717796ee372f6920de29f","name":"Yilun Zhao","hidden":false},{"_id":"6a4717796ee372f6920de2a0","name":"Arman Cohan","hidden":false}],"publishedAt":"2026-07-01T00:00:00.000Z","submittedOnDailyAt":"2026-07-06T00:00:00.000Z","title":"Measuring the Gap Between Human and LLM Research Ideas","submittedOnDailyBy":{"_id":"666710655b45380937a5d472","avatarUrl":"/avatars/6e41eb68519c6bb99567af4ef54df2cf.svg","isPro":false,"fullname":"Ziyu","user":"ziyuuc","type":"user","name":"ziyuuc"},"summary":"LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.","upvotes":6,"discussionId":"6a4717796ee372f6920de2a1","githubRepo":"https://github.com/ziyuuc/TasteGap","githubRepoAddedBy":"user","ai_summary":"Large language models generate research ideas that cluster around specific opportunity patterns and paradigms, diverging systematically from the broader and more diverse distributions found in human research papers.","ai_keywords":["research ideation","large language models","research taste taxonomy","opportunity patterns","research paradigms","human research papers","idea generation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":4,"organization":{"_id":"6532df27d690f3012efde84c","name":"yale-nlp","fullname":"Yale NLP Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65204db5b0e0d57453cb1809/9OAeiZ-BrN2g1h1yd6-1W.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"666710655b45380937a5d472","avatarUrl":"/avatars/6e41eb68519c6bb99567af4ef54df2cf.svg","isPro":false,"fullname":"Ziyu","user":"ziyuuc","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"62f662bcc58915315c4eccea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62f662bcc58915315c4eccea/zOAQLONfMP88zr70sxHK-.jpeg","isPro":true,"fullname":"Yilun Zhao","user":"yilunzhao","type":"user"},{"_id":"676ce7767fff9075b5d526fa","avatarUrl":"/avatars/bf6697163b91564a8d4b773d3f6420bf.svg","isPro":false,"fullname":"Kaiyan Zhang","user":"maxzky","type":"user"},{"_id":"67492b9e347c3876f22b3684","avatarUrl":"/avatars/0d80d23f7b10ce8bac689f6e8317a014.svg","isPro":false,"fullname":"Tiansheng Hu","user":"HughieHu","type":"user"},{"_id":"683c642b02c1a474a867964e","avatarUrl":"/avatars/63e44a9cf788ee7b3ad236407700ceca.svg","isPro":false,"fullname":"Jinbiao Wei","user":"mikeweii","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6532df27d690f3012efde84c","name":"yale-nlp","fullname":"Yale NLP Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65204db5b0e0d57453cb1809/9OAeiZ-BrN2g1h1yd6-1W.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.01233.md","query":{}}">
Measuring the Gap Between Human and LLM Research Ideas
Published on Jul 1
· Submitted by Ziyu on Jul 6 Abstract
Large language models generate research ideas that cluster around specific opportunity patterns and paradigms, diverging systematically from the broader and more diverse distributions found in human research papers.
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.
Community
We introduce a large-scale framework for measuring how LLM-generated research ideas differ from human research ideas. We study ideation as a distributional alignment problem: given the same local literature context, do LLMs identify the same kinds of opportunities and construct the same kinds of contributions as researchers? We build from 11.7K papers across ML conferences and Nature Communications, reverse-engineering proximal prior works for each paper and extracting the human idea as a motivation–method pair. We then prompt nine LLMs to generate ideas from the same context and annotate both human and model ideas with a two-axis research-taste taxonomy covering opportunity patterns and method paradigms. Across models and domains, we find a stable gap that LLM ideas concentrate heavily on bridge-like motivations and synthesis methods, while human papers span a broader range research topic. Reasoning and richer full-paper context do not close this gap. Reasoning even sharpens the template. These results suggest that future AI ideation systems should optimize not only individual idea quality, but also diversity of research taste.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.01233 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.01233 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.01233 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.