Hugging Face Daily Papers · · 4 min read

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🧬 Introducing 𝐄𝐌𝐁𝐋 𝐀𝐈 𝐋𝐈𝐁𝐑𝐀𝐑𝐈𝐀𝐍.<br>A knowledge layer for life-science AI agents.<br>Natural-language question in. Evidence snippet out.<br>No vector database. No separate literature index. Always up to date.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/6500880f051fae19fc515eda/HeRwiCk-OsRY4HNvbg-SS.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6500880f051fae19fc515eda/HeRwiCk-OsRY4HNvbg-SS.png\" alt=\"x-slide-1\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/6500880f051fae19fc515eda/62NPdohDljlfyxrG9l0rC.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6500880f051fae19fc515eda/62NPdohDljlfyxrG9l0rC.png\" alt=\"x-slide-2\"></a></p>\n","updatedAt":"2026-08-03T14:30:14.042Z","author":{"_id":"6500880f051fae19fc515eda","avatarUrl":"/avatars/c69df184dfe44f7feca6e6f45edfc9b1.svg","fullname":"Luigi Sigillo","name":"luigi-s","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5757988691329956},"editors":["luigi-s"],"editorAvatarUrls":["/avatars/c69df184dfe44f7feca6e6f45edfc9b1.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28229","authors":[{"_id":"6a70355dbbe824e6bcc466fb","name":"Luigi Sigillo","hidden":false},{"_id":"6a70355dbbe824e6bcc466fc","name":"Matteo Silvestri","hidden":false},{"_id":"6a70355dbbe824e6bcc466fd","name":"Francesco Tabaro","hidden":false},{"_id":"6a70355dbbe824e6bcc466fe","name":"Rajat Bhatnagar","hidden":false},{"_id":"6a70355dbbe824e6bcc466ff","name":"Syed Irtaza Mubashar","hidden":false},{"_id":"6a70355dbbe824e6bcc46700","name":"Matt Jeffryes","hidden":false},{"_id":"6a70355dbbe824e6bcc46701","name":"Daljit Nijjer","hidden":false},{"_id":"6a70355dbbe824e6bcc46702","name":"Vittorio Perera","hidden":false},{"_id":"6a70355dbbe824e6bcc46703","name":"Ola Spjuth","hidden":false},{"_id":"6a70355dbbe824e6bcc46704","name":"Julio Saez-Rodriguez","hidden":false},{"_id":"6a70355dbbe824e6bcc46705","name":"Melissa Harrison","hidden":false},{"_id":"6a70355dbbe824e6bcc46706","name":"Fabio Petroni","hidden":false}],"publishedAt":"2026-07-30T14:00:27.000Z","submittedOnDailyAt":"2026-08-03T00:00:00.000Z","title":"EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents","submittedOnDailyBy":{"_id":"6500880f051fae19fc515eda","avatarUrl":"/avatars/c69df184dfe44f7feca6e6f45edfc9b1.svg","isPro":false,"fullname":"Luigi Sigillo","user":"luigi-s","type":"user","name":"luigi-s"},"summary":"The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than 16 points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about 8 points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at https://github.com/petroni-lab/librarian","upvotes":4,"discussionId":"6a70355dbbe824e6bcc46707","githubRepo":"https://github.com/petroni-lab/librarian","githubRepoAddedBy":"user","githubStars":8,"organization":{"_id":"6a70a76394d22bc61014432d","name":"EMBL-Rome","fullname":"European Molecular Biology Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6500880f051fae19fc515eda/UPaEOvAF76FN9DjbH_K0H.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"672a3417aa84596b8aa3ccba","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/eMyP3q3rAjRfSzmqpvJeC.png","isPro":false,"fullname":"Matteo Silvestri","user":"phdsilver22","type":"user"},{"_id":"6500880f051fae19fc515eda","avatarUrl":"/avatars/c69df184dfe44f7feca6e6f45edfc9b1.svg","isPro":false,"fullname":"Luigi Sigillo","user":"luigi-s","type":"user"},{"_id":"6270324ebecab9e2dcf245de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6270324ebecab9e2dcf245de/cMbtWSasyNlYc9hvsEEzt.jpeg","isPro":false,"fullname":"Kye Gomez","user":"kye","type":"user"},{"_id":"6a1d8f24bd8b36b69d8ae55a","avatarUrl":"/avatars/37b9f25c7e0f276b458aa2c6d4519b1e.svg","isPro":false,"fullname":"Francesco Tabaro","user":"ftabaro","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a70a76394d22bc61014432d","name":"EMBL-Rome","fullname":"European Molecular Biology Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6500880f051fae19fc515eda/UPaEOvAF76FN9DjbH_K0H.png"},"query":{}}">
Papers
arxiv:2607.28229

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

Published on Jul 30
· Submitted by
Luigi Sigillo
on Aug 3
Authors:
,

Abstract

The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than 16 points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about 8 points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at https://github.com/petroni-lab/librarian

Community

Paper submitter about 2 hours ago

🧬 Introducing 𝐄𝐌𝐁𝐋 𝐀𝐈 𝐋𝐈𝐁𝐑𝐀𝐑𝐈𝐀𝐍.
A knowledge layer for life-science AI agents.
Natural-language question in. Evidence snippet out.
No vector database. No separate literature index. Always up to date.

x-slide-1

x-slide-2

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.28229 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.28229 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.28229 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers