Hugging Face Daily Papers · · 3 min read

HPD-Parsing: Hierarchical Parallel Document Parsing

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A New Paradigm for Document Parsing from the PaddleOCR Team</p>\n","updatedAt":"2026-07-22T04:20:45.479Z","author":{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","fullname":"cuicheng","name":"ChengCui","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":40,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7702298164367676},"editors":["ChengCui"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18839","authors":[{"_id":"6a6036b77e7f152167e47132","user":{"_id":"630ec94805b5c5280e3bf194","avatarUrl":"/avatars/3bd98d7e9bfc402f198e2ada704cefe7.svg","isPro":false,"fullname":"WEI","user":"WEISHU","type":"user","name":"WEISHU"},"name":"Shu Wei","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:40:47.143Z","hidden":false},{"_id":"6a6036b77e7f152167e47133","name":"Jingjing Wu","hidden":false},{"_id":"6a6036b77e7f152167e47134","name":"Lingshu Zhang","hidden":false},{"_id":"6a6036b77e7f152167e47135","name":"Qunyi Xie","hidden":false},{"_id":"6a6036b77e7f152167e47136","name":"Hao Zou","hidden":false},{"_id":"6a6036b77e7f152167e47137","name":"Le Xiang","hidden":false},{"_id":"6a6036b77e7f152167e47138","name":"Xu Fan","hidden":false},{"_id":"6a6036b77e7f152167e47139","name":"Yangliu Xu","hidden":false},{"_id":"6a6036b77e7f152167e4713a","user":{"_id":"67d96e68939c3823ab2e06a5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FJx_Nf_5aD1w1s0RqbyGn.png","isPro":false,"fullname":"Manhui Lin","user":"gggdddfff","type":"user","name":"gggdddfff"},"name":"Manhui Lin","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.335Z","hidden":false},{"_id":"6a6036b77e7f152167e4713b","name":"Xiaolong Ma","hidden":false},{"_id":"6a6036b77e7f152167e4713c","user":{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","isPro":false,"fullname":"cuicheng","user":"ChengCui","type":"user","name":"ChengCui"},"name":"Cheng Cui","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.328Z","hidden":false},{"_id":"6a6036b77e7f152167e4713d","name":"Tengyu Du","hidden":false},{"_id":"6a6036b77e7f152167e4713e","name":"YY","hidden":false}],"publishedAt":"2026-07-21T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"HPD-Parsing: Hierarchical Parallel Document Parsing","submittedOnDailyBy":{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","isPro":false,"fullname":"cuicheng","user":"ChengCui","type":"user","name":"ChengCui"},"summary":"Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autoregressive trajectory, creating a sequential bottleneck that grows with document length. Such full-page sequential generation overlooks a key property of document parsing: layout must be analyzed globally, whereas block content can be parsed in parallel. Based on this observation, we introduce HPD-Parsing, which replaces full-page autoregressive generation with a Hierarchical Parallel Decoding paradigm. A main layout branch organizes the overall document structure and dynamically assigns block-level content decoding to concurrent branches, while progressive multi-token prediction (P-MTP) further reduces the decoding steps within each branch. Experiments on public benchmarks show that HPD-Parsing achieves 4,752 tokens per second, delivering 2.62times the throughput of the fastest existing document parsing model and 3.06times that of the vanilla autoregressive baseline, while maintaining competitive parsing accuracy. These results establish hierarchical parallel decoding as an effective alternative to full-page autoregressive generation, opening a new direction for efficient unified document parsing.","upvotes":5,"discussionId":"6a6036b77e7f152167e4713f","organization":{"_id":"62067d5d3906f102bc9658bd","name":"PaddlePaddle","fullname":"PaddlePaddle","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1654942635336-5f3ff69679c1ba4c353d0c5a.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","isPro":false,"fullname":"cuicheng","user":"ChengCui","type":"user"},{"_id":"630ec94805b5c5280e3bf194","avatarUrl":"/avatars/3bd98d7e9bfc402f198e2ada704cefe7.svg","isPro":false,"fullname":"WEI","user":"WEISHU","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6780890b816943dee855d100","avatarUrl":"/avatars/cd73a57d8078b6ede0a7f0aee04b28d8.svg","isPro":false,"fullname":"Wujingjing","user":"wjj1999","type":"user"},{"_id":"65e48259a4e46e644ebf4f6a","avatarUrl":"/avatars/4950c9e0b1bedd357fd6fa606665d4a5.svg","isPro":false,"fullname":"Johny_XIE","user":"johnyxie","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"62067d5d3906f102bc9658bd","name":"PaddlePaddle","fullname":"PaddlePaddle","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1654942635336-5f3ff69679c1ba4c353d0c5a.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18839.md","query":{}}">
Papers
arxiv:2607.18839

HPD-Parsing: Hierarchical Parallel Document Parsing

Published on Jul 21
· Submitted by
cuicheng
on Jul 22
Authors:

Abstract

Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autoregressive trajectory, creating a sequential bottleneck that grows with document length. Such full-page sequential generation overlooks a key property of document parsing: layout must be analyzed globally, whereas block content can be parsed in parallel. Based on this observation, we introduce HPD-Parsing, which replaces full-page autoregressive generation with a Hierarchical Parallel Decoding paradigm. A main layout branch organizes the overall document structure and dynamically assigns block-level content decoding to concurrent branches, while progressive multi-token prediction (P-MTP) further reduces the decoding steps within each branch. Experiments on public benchmarks show that HPD-Parsing achieves 4,752 tokens per second, delivering 2.62times the throughput of the fastest existing document parsing model and 3.06times that of the vanilla autoregressive baseline, while maintaining competitive parsing accuracy. These results establish hierarchical parallel decoding as an effective alternative to full-page autoregressive generation, opening a new direction for efficient unified document parsing.

Community

Paper author Paper submitter about 5 hours ago

A New Paradigm for Document Parsing from the PaddleOCR Team

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.18839
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.18839 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers