We built a point-in-time financial deep research benchmark, featuring questions and rubrics generated through a rigorous, quality-controlled data pipeline. Additionally, we contracted financial experts to validate our data, with each spending an average of 1.2 hours on this meticulous review process. Leading LLMs such Opus-5 with our Finance Harness only score 44.9% on our leaderboard, showcasing the significant challenge our benchmark presents. We invite everyone to contribute to our leaderboard!</p>\n","updatedAt":"2026-08-06T21:52:14.385Z","author":{"_id":"64c04feeb746fe51543c1b7a","avatarUrl":"/avatars/77b87fca59a64c2463d652d454d78b13.svg","fullname":"Han","name":"Rujun","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8933539390563965},"editors":["Rujun"],"editorAvatarUrls":["/avatars/77b87fca59a64c2463d652d454d78b13.svg"],"reactions":[],"isReport":false}},{"id":"6a753a7244433412ccf79321","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-07T01:52:50.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality](https://huggingface.co/papers/2607.12252) (2026)\n* [FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance](https://huggingface.co/papers/2607.19409) (2026)\n* [ICBCBench: An Industry Consortium Benchmark for Financial Deep Research](https://huggingface.co/papers/2606.17458) (2026)\n* [Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark](https://huggingface.co/papers/2606.18648) (2026)\n* [DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction](https://huggingface.co/papers/2606.18191) (2026)\n* [IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO](https://huggingface.co/papers/2606.23032) (2026)\n* [How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?](https://huggingface.co/papers/2606.26346) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.12252\">FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.19409\">FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.17458\">ICBCBench: An Industry Consortium Benchmark for Financial Deep Research</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.18648\">Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.18191\">DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.23032\">IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.26346\">How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-07T01:52:50.411Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7345870137214661},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27853","authors":[{"_id":"6a74ee80e1228e04b323809d","name":"Yijia Xiao","hidden":false},{"_id":"6a74ee80e1228e04b323809e","name":"Rujun Han","hidden":false},{"_id":"6a74ee80e1228e04b323809f","name":"Yanfei Chen","hidden":false},{"_id":"6a74ee80e1228e04b32380a0","name":"Zifeng Wang","hidden":false},{"_id":"6a74ee80e1228e04b32380a1","name":"Ke Jiang","hidden":false},{"_id":"6a74ee80e1228e04b32380a2","name":"Zhongying CuiZhu","hidden":false},{"_id":"6a74ee80e1228e04b32380a3","name":"Vishy Tirumalashetty","hidden":false},{"_id":"6a74ee80e1228e04b32380a4","name":"Wei Wang","hidden":false},{"_id":"6a74ee80e1228e04b32380a5","name":"Burak Gokturk","hidden":false},{"_id":"6a74ee80e1228e04b32380a6","name":"Tomas Pfister","hidden":false},{"_id":"6a74ee80e1228e04b32380a7","name":"Chen-Yu Lee","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"FinanceHarness: Autonomous Financial Deep Research Framework","submittedOnDailyBy":{"_id":"64c04feeb746fe51543c1b7a","avatarUrl":"/avatars/77b87fca59a64c2463d652d454d78b13.svg","isPro":false,"fullname":"Han","user":"Rujun","type":"user","name":"Rujun"},"summary":"Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. Even leading LLMs and agents score below 40% on the rubrics, showing that FinanceGym is challenging and leaves substantial headroom. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%. FinanceHarness is available at https://github.com/Yijia-Xiao/FinanceHarness.","upvotes":3,"discussionId":"6a74ee81e1228e04b32380a8","projectPage":"https://financegym.github.io/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64c04feeb746fe51543c1b7a","avatarUrl":"/avatars/77b87fca59a64c2463d652d454d78b13.svg","isPro":false,"fullname":"Han","user":"Rujun","type":"user"},{"_id":"64d1fe784dfd5df7076bfd1a","avatarUrl":"/avatars/e2bb108d1f6b1383f0e3f263c4747b3d.svg","isPro":false,"fullname":"Chen-Yu Lee","user":"chenyulee","type":"user"},{"_id":"67c772e077464ecefff08038","avatarUrl":"/avatars/23ad9716715bbaf934a646e45c8f5ae0.svg","isPro":false,"fullname":"Zoey CuiZhu","user":"cuizz11","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
FinanceHarness: Autonomous Financial Deep Research Framework
Published on Jul 30
· Submitted by Han on Aug 6 Abstract
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. Even leading LLMs and agents score below 40% on the rubrics, showing that FinanceGym is challenging and leaves substantial headroom. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%. FinanceHarness is available at https://github.com/Yijia-Xiao/FinanceHarness.
Community
We built a point-in-time financial deep research benchmark, featuring questions and rubrics generated through a rigorous, quality-controlled data pipeline. Additionally, we contracted financial experts to validate our data, with each spending an average of 1.2 hours on this meticulous review process. Leading LLMs such Opus-5 with our Finance Harness only score 44.9% on our leaderboard, showcasing the significant challenge our benchmark presents. We invite everyone to contribute to our leaderboard!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.27853 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.27853 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.