Hugging Face Daily Papers · · 4 min read

DrugGen 2: A disease-aware language model for enhancing drug discovery

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences.</p>\n","updatedAt":"2026-07-10T08:37:42.983Z","author":{"_id":"61990d48d7f09e0d8b7714de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61990d48d7f09e0d8b7714de/sL3Tc36iuBZr0_Dy0wWCc.jpeg","fullname":"Ali Motahharynia","name":"alimotahharynia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9050942063331604},"editors":["alimotahharynia"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/61990d48d7f09e0d8b7714de/sL3Tc36iuBZr0_Dy0wWCc.jpeg"],"reactions":[{"reaction":"😎","users":["alifardinia00"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.08404","authors":[{"_id":"6a50a88375fd3d966bd45f0f","name":"Ali Motahharynia","hidden":false},{"_id":"6a50a88375fd3d966bd45f10","name":"Mohammadreza Ghaffarzadeh-Esfahani","hidden":false},{"_id":"6a50a88375fd3d966bd45f11","name":"Mahsa Sheikholeslami","hidden":false},{"_id":"6a50a88375fd3d966bd45f12","name":"Navid Mazrouei","hidden":false},{"_id":"6a50a88375fd3d966bd45f13","name":"Matin Irajpour","hidden":false},{"_id":"6a50a88375fd3d966bd45f14","name":"Yousof Gheisari","hidden":false},{"_id":"6a50a88375fd3d966bd45f15","name":"Hajar Sirous","hidden":false}],"publishedAt":"2026-07-09T00:00:00.000Z","submittedOnDailyAt":"2026-07-10T00:00:00.000Z","title":"DrugGen 2: A disease-aware language model for enhancing drug discovery","submittedOnDailyBy":{"_id":"61990d48d7f09e0d8b7714de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61990d48d7f09e0d8b7714de/sL3Tc36iuBZr0_Dy0wWCc.jpeg","isPro":false,"fullname":"Ali Motahharynia","user":"alimotahharynia","type":"user","name":"alimotahharynia"},"summary":"Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences. DrugGen-2 was developed by fine-tuning a pre-trained GPT-2 model on a curated dataset of approved drugs linked to their diseases and targets, using a two-step strategy of supervised fine-tuning followed by reinforcement learning via group relative policy optimization (GRPO). This process was guided by reward functions optimizing for chemical validity, novelty, diversity, and high predicted binding affinity. When evaluated on five protein targets relevant to diabetic nephropathy, DrugGen-2 significantly outperformed baseline models (DrugGPT and DrugGen). It demonstrated a superior capacity to generate unique molecules, exhibited greater structural similarity to approved drugs, and achieved improved predicted binding affinities across all targets. Molecular docking analyses further supported these findings, identifying candidate ligands with strong binding potential, including compounds with predicted affinities (-9.917, -9.485, and -9.367) exceeding those of reference drugs such as enalapril for angiotensin-converting enzyme (-8.283). By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.","upvotes":12,"discussionId":"6a50a88475fd3d966bd45f16","projectPage":"https://huggingface.co/spaces/alimotahharynia/DrugGen-2","githubRepo":"https://github.com/alimotahharynia/DrugGen-2","githubRepoAddedBy":"user","ai_summary":"DrugGen-2 generates small molecules conditioned on disease ontology and target protein sequences through fine-tuning GPT-2 with supervised learning and reinforcement learning using GRPO, achieving superior molecular diversity and binding affinity compared to baseline models.","ai_keywords":["GPT-2","supervised fine-tuning","reinforcement learning","group relative policy optimization","chemical validity","molecular generation","binding affinity","molecular docking"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":3},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"61990d48d7f09e0d8b7714de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61990d48d7f09e0d8b7714de/sL3Tc36iuBZr0_Dy0wWCc.jpeg","isPro":false,"fullname":"Ali Motahharynia","user":"alimotahharynia","type":"user"},{"_id":"645d63c0ce72244df7b36be8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/645d63c0ce72244df7b36be8/09vhYAzgv1svwvQM4eIE9.jpeg","isPro":false,"fullname":"MoRezaGH","user":"Moreza009","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6434b6619bd5a84b5dcfa4de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6434b6619bd5a84b5dcfa4de/h8Q6kPNjFNc03wmdboHzq.jpeg","isPro":true,"fullname":"Young-Jun Lee","user":"passing2961","type":"user"},{"_id":"6789b4147bd5f0f1891d0525","avatarUrl":"/avatars/74c1fedf42bf3718eb88dcaf1a762f2a.svg","isPro":false,"fullname":"Nahid Yousefian","user":"nahidysf","type":"user"},{"_id":"67c40fa2c0a504dc033f6c8b","avatarUrl":"/avatars/a2a92b7bbc1d875687198bd18d2f9303.svg","isPro":false,"fullname":"hossein bekas","user":"hossein12365555555","type":"user"},{"_id":"67c412b1ecc2b3bb52331880","avatarUrl":"/avatars/7f6ebaec60f6f42753bb351efa74902e.svg","isPro":false,"fullname":"ali kobraiian","user":"kobraiian898989","type":"user"},{"_id":"651fce4ffc791e23f7208490","avatarUrl":"/avatars/d0967380dc3dbc7012f16b64a4a1a6f4.svg","isPro":false,"fullname":"Michel Kafold","user":"Michel-kafold009","type":"user"},{"_id":"671695e68c364caa43429de8","avatarUrl":"/avatars/0397a28b80ae0996c96677a7956e301f.svg","isPro":false,"fullname":"maxinose","user":"ahmadmax77","type":"user"},{"_id":"6716addaf51dd4be6650962c","avatarUrl":"/avatars/0c10d10407c45983cb470b289b33de80.svg","isPro":false,"fullname":"Sorour ragheb","user":"Sorour122","type":"user"},{"_id":"67181b533a5fdf2aeefdba6f","avatarUrl":"/avatars/52641a256152c64293c4e64e77e21ad6.svg","isPro":false,"fullname":"sual","user":"bettercall96","type":"user"},{"_id":"651e865843e7d4ac931b8519","avatarUrl":"/avatars/72833bf2f77842b6ca880453dfe67784.svg","isPro":false,"fullname":"ali fardini","user":"alifardinia00","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.08404.md","query":{}}">
Papers
arxiv:2607.08404

DrugGen 2: A disease-aware language model for enhancing drug discovery

Published on Jul 9
· Submitted by
Ali Motahharynia
on Jul 10
Authors:
,

Abstract

DrugGen-2 generates small molecules conditioned on disease ontology and target protein sequences through fine-tuning GPT-2 with supervised learning and reinforcement learning using GRPO, achieving superior molecular diversity and binding affinity compared to baseline models.

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences. DrugGen-2 was developed by fine-tuning a pre-trained GPT-2 model on a curated dataset of approved drugs linked to their diseases and targets, using a two-step strategy of supervised fine-tuning followed by reinforcement learning via group relative policy optimization (GRPO). This process was guided by reward functions optimizing for chemical validity, novelty, diversity, and high predicted binding affinity. When evaluated on five protein targets relevant to diabetic nephropathy, DrugGen-2 significantly outperformed baseline models (DrugGPT and DrugGen). It demonstrated a superior capacity to generate unique molecules, exhibited greater structural similarity to approved drugs, and achieved improved predicted binding affinities across all targets. Molecular docking analyses further supported these findings, identifying candidate ligands with strong binding potential, including compounds with predicted affinities (-9.917, -9.485, and -9.367) exceeding those of reference drugs such as enalapril for angiotensin-converting enzyme (-8.283). By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.

Community

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.08404
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers