logoalt Hacker News

eviksyesterday at 7:59 PM2 repliesview on HN

> The original character (𡚴) was not added to JIS or Unicode until much later and doesn't display on most sites for me

Why didn't they simly replace the original bad one?

> nine hundred pages. Imagine tracking down a single character without a page reference

Not that hard to imagine, OCR existed back then?


Replies

gucci-on-fleekyesterday at 9:49 PM

> Not that hard to imagine, OCR existed back then?

How do you train OCR on a character that doesn't exist though? Even these days, OCR often confuses the 52 Latin characters, so I wouldn't expect as good of results with older technology and some 20k CJK characters.

show 1 reply
Kyeyesterday at 8:37 PM

OCR was slow and unreliable and was for a very long time.

show 1 reply