r/LocalLLaMA · · 1 min read

CJK Manga/Manhwa/Manhua 150M OCR model (hayai-ocr-v2) outperforming PaddleOCR-VL-For-Manga

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

CJK Manga/Manhwa/Manhua 150M OCR model (hayai-ocr-v2) outperforming PaddleOCR-VL-For-Manga

https://preview.redd.it/o7794vsszqih1.png?width=687&format=png&auto=webp&s=3b217db46fb51c517e103d12e2e2ca9813b7774f

https://preview.redd.it/ci1as7nmzqih1.png?width=800&format=png&auto=webp&s=811e0b30de0f76ee548ef7701359674d8f8bf32a

I trained a custom model with a custom decoder and siglip2-naflex vision encoder that performs better than PaddleOCR-VL-For-Manga while being more than 10x faster and smaller. Please try it out at hayai-ocr-v2 and let me know if it's any good for your particular task. I will integrate this model soon in the hayai-ocr python library.

NOTE: Finetune and Pretrain refers to different eval datasets.

submitted by /u/KingDutchIsBad455
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA