I compared even more parsers on 14 PDF-parsing capabilities using different types
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| In a previous post, I compared MinerU, Granite-Docling, and PaddleOCR-VL. Many commentors suggested I added their favorite parsers. So I did. And also added some new capabilities to differentiate the top models. Here is the full list of parser compared:
What I found:
Same disclosure as before: the three original VLM rows ran on hexread.com (my product), everything else ran locally or an L4. EDIT: Sources, raw outputs, test files and scripts are in this repo for reference: alaamroue/pdf-parser-bench [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.