End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline.</p>\n","updatedAt":"2026-08-07T04:56:03.653Z","author":{"_id":"669205f1ccca14aa8f13f770","avatarUrl":"/avatars/11ce274e93345fe3790ac9fa687e2bcb.svg","fullname":"Hao Yu","name":"Longin-Yu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8380445241928101},"editors":["Longin-Yu"],"editorAvatarUrls":["/avatars/11ce274e93345fe3790ac9fa687e2bcb.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06146","authors":[{"_id":"6a7564fde1228e04b32382db","name":"Hao Yu","hidden":false},{"_id":"6a7564fde1228e04b32382dc","name":"Jiabo Zhan","hidden":false},{"_id":"6a7564fde1228e04b32382dd","name":"Kang Liu","hidden":false},{"_id":"6a7564fde1228e04b32382de","name":"Linnan Zhao","hidden":false},{"_id":"6a7564fde1228e04b32382df","name":"Dongxu Yue","hidden":false},{"_id":"6a7564fde1228e04b32382e0","name":"Rui Chen","hidden":false},{"_id":"6a7564fde1228e04b32382e1","name":"Jinglin Wang","hidden":false},{"_id":"6a7564fde1228e04b32382e2","name":"Chong Sun","hidden":false},{"_id":"6a7564fde1228e04b32382e3","name":"Chen Li","hidden":false},{"_id":"6a7564fde1228e04b32382e4","name":"Jing Lyu","hidden":false},{"_id":"6a7564fde1228e04b32382e5","name":"Chun Yuan","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"PaDoc: Layout-Grounded Parallel Decoding for Document Parsing","submittedOnDailyBy":{"_id":"669205f1ccca14aa8f13f770","avatarUrl":"/avatars/11ce274e93345fe3790ac9fa687e2bcb.svg","isPro":false,"fullname":"Hao Yu","user":"Longin-Yu","type":"user","name":"Longin-Yu"},"summary":"End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline. Code is available at https://github.com/Longin-Yu/Padoc","upvotes":10,"discussionId":"6a7564fee1228e04b32382e6","githubRepo":"https://github.com/Longin-Yu/Padoc","githubRepoAddedBy":"user","githubStars":5,"organization":{"_id":"665c80b1d2b102a5bb77a4de","name":"tsinghua-uni","fullname":"Tsinghua University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6199e76004b5da0c05211e25/wS8ToqQakyAIW5DKuoATi.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"669205f1ccca14aa8f13f770","avatarUrl":"/avatars/11ce274e93345fe3790ac9fa687e2bcb.svg","isPro":false,"fullname":"Hao Yu","user":"Longin-Yu","type":"user"},{"_id":"66eb8ebae604909be7f0eea0","avatarUrl":"/avatars/b826e25f8b164a1f185bc76992bee1d8.svg","isPro":false,"fullname":"Jiabo Zhan","user":"TonyZhan","type":"user"},{"_id":"68ce3eb08db3012f60069099","avatarUrl":"/avatars/5e02fb7fa7e3dfa2a3d4deccc35a44da.svg","isPro":false,"fullname":"Jinglin Wang","user":"wang-0538","type":"user"},{"_id":"674d092c6421c58761fc83eb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674d092c6421c58761fc83eb/lQSRWX_YyTzpRq2aJtwYT.png","isPro":false,"fullname":"Xingsong Ye","user":"Yesianrohn","type":"user"},{"_id":"6a50ae91b83044b170ac97ea","avatarUrl":"/avatars/c17fd9f9fc61a3024f0c4323ceb962c7.svg","isPro":false,"fullname":"psp","user":"luckin201125","type":"user"},{"_id":"62ea2d48b493272c269e8f34","avatarUrl":"/avatars/89b215dafa503b51ab212a9b63c82aca.svg","isPro":false,"fullname":"Hanyu Lai","user":"hanyullai","type":"user"},{"_id":"65781534e390cfd40998d7af","avatarUrl":"/avatars/85d3be7f74b9d959ba9b3ccc04398536.svg","isPro":false,"fullname":"Hongyang Wei","user":"nonwhy","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"652a8226b355406e2dc9943f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/nK_ddRC-NVcpweDg3uFwj.jpeg","isPro":false,"fullname":"Rui Chen","user":"Conroy","type":"user"},{"_id":"65a5231a087d8a2e9cc2414b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a5231a087d8a2e9cc2414b/wj0l5R5LmBUG-E8XTMdBM.jpeg","isPro":false,"fullname":"cuicheng","user":"ChengCui","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"665c80b1d2b102a5bb77a4de","name":"tsinghua-uni","fullname":"Tsinghua University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6199e76004b5da0c05211e25/wS8ToqQakyAIW5DKuoATi.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06146.md","query":{}}">
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
Published on Aug 6
· Submitted by Hao Yu on Aug 7 Abstract
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline. Code is available at https://github.com/Longin-Yu/Padoc
Community
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.06146 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.