logoalt Hacker News

throwaw12today at 11:17 AM5 repliesview on HN

that's difficult as well, how do you k ow where to split?


Replies

johndoughtoday at 11:51 AM

There are models specifically for splitting an image into text regions, e.g. PP-DocLayoutV3 https://huggingface.co/PaddlePaddle/PP-DocLayoutV3

I am using a stripped-down minimal version of it which I uploaded here, since I am not a fan of huge dependency trees: https://github.com/99991/simple-pp-doclayoutv3

Another recent model for this task is Unlimited-OCR: https://github.com/baidu/Unlimited-OCR

kgwgktoday at 12:23 PM

Text is often written as separate lines (and paragraphs) at least in some languages.

wongarsutoday at 11:47 AM

Let the model do the splitting. A 800x800px image should be enough to make those decisions

grog454today at 11:43 AM

Overlap the splits?

vrganjtoday at 11:40 AM

Presumably a small cheap model could do that part?