Hugging Face Daily Papers · · 5 min read

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.</p>\n","updatedAt":"2026-07-24T11:30:27.623Z","author":{"_id":"66c596fccf0439733e38913f","avatarUrl":"/avatars/f4805286596305996613689868c57fa2.svg","fullname":"Xianfu Cheng","name":"BuaaCXF","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8848273754119873},"editors":["BuaaCXF"],"editorAvatarUrls":["/avatars/f4805286596305996613689868c57fa2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.19238","authors":[{"_id":"6a63434f501b0a811675372f","name":"Xianfu Cheng","hidden":false},{"_id":"6a63434f501b0a8116753730","name":"Shiwei Zhang","hidden":false},{"_id":"6a63434f501b0a8116753731","name":"Jiyu Zhao","hidden":false},{"_id":"6a63434f501b0a8116753732","name":"Jian Yang","hidden":false},{"_id":"6a63434f501b0a8116753733","name":"Xinyuan Wang","hidden":false},{"_id":"6a63434f501b0a8116753734","name":"Ming Zhou","hidden":false},{"_id":"6a63434f501b0a8116753735","name":"Weixiao Zhou","hidden":false},{"_id":"6a63434f501b0a8116753736","name":"Xiangyuan Guan","hidden":false},{"_id":"6a63434f501b0a8116753737","name":"Xiang Li","hidden":false},{"_id":"6a63434f501b0a8116753738","name":"Zhenhe Wu","hidden":false},{"_id":"6a63434f501b0a8116753739","name":"Ziyi Ni","hidden":false},{"_id":"6a63434f501b0a811675373a","name":"Zhoujun Li","hidden":false},{"_id":"6a63434f501b0a811675373b","name":"Bingjing Xu","hidden":false}],"publishedAt":"2026-07-21T00:00:00.000Z","submittedOnDailyAt":"2026-07-24T00:00:00.000Z","title":"FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents","submittedOnDailyBy":{"_id":"66c596fccf0439733e38913f","avatarUrl":"/avatars/f4805286596305996613689868c57fa2.svg","isPro":false,"fullname":"Xianfu Cheng","user":"BuaaCXF","type":"user","name":"BuaaCXF"},"summary":"Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.","upvotes":4,"discussionId":"6a63434f501b0a811675373c","projectPage":"https://huggingface.co/datasets/Multilingual-Multimodal-NLP/FinanceComplexQA","githubRepo":"https://github.com/buaacxf/FinanceComplexQA","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"63ba7720fc454697637969f1","name":"Beihang","fullname":"Beihang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ba7666c138c8f2b7844b58/n98lZU9VWxYgWIkzE_6o4.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66c596fccf0439733e38913f","avatarUrl":"/avatars/f4805286596305996613689868c57fa2.svg","isPro":false,"fullname":"Xianfu Cheng","user":"BuaaCXF","type":"user"},{"_id":"69bf5ef49ac7b84490b5096c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/A0b-T9AqmIPHHRaSrnB1r.png","isPro":false,"fullname":"Ruoxi Tang","user":"james-gonzalezh","type":"user"},{"_id":"63f7767fbd28622c9b9915e9","avatarUrl":"/avatars/d2c624d2815cb0effb0c26a7687b4388.svg","isPro":false,"fullname":"Ziyi Ni","user":"Nicole-Yi","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63ba7720fc454697637969f1","name":"Beihang","fullname":"Beihang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ba7666c138c8f2b7844b58/n98lZU9VWxYgWIkzE_6o4.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.19238.md","query":{}}">
Papers
arxiv:2607.19238

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

Published on Jul 21
· Submitted by
Xianfu Cheng
on Jul 24
Authors:
,

Abstract

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.

Community

Paper submitter about 9 hours ago

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.19238
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.19238 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.19238 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers