logoalt Hacker News

A spectre is haunting Unicode

154 pointsby sensanatytoday at 2:34 PM47 commentsview on HN

Comments

joshdavhamtoday at 4:52 PM

The author Paul McCann (polm) is one of my favourite programmers out there!

He’s done awesome work in the Japanese NLP space over the last decade which has really helped me in my language learning projects.

He maintains a mecab (Japanese tokenizer) wrapper for Python [0], has a book on Japanese NLP written for English speakers [1] and also worked on Spacy at one point [2].

[0] https://github.com/polm/fugashi [1] https://www.japanesenlp.com/ [2] https://spacy.io/

gweinbergtoday at 8:27 PM

It occurs to me that we can use 彊 to mean "a completely unknown concept that cannot be named". For example if you ask, "when Cthulhu rises from its slumber, what is thefirst thing it will do? Probably it will 彊.

show 2 replies
erjiangtoday at 5:59 PM

I think there’s evidence found for the origin of “彁” as the result of a poor scan of a newspaper article. Look up “彁 新聞” to find some japanese sources about this.

xelxebartoday at 10:38 PM

Xu Bing has a book that consists entirely of invented characters:

https://en.wikipedia.org/wiki/A_Book_from_the_Sky

hnfongtoday at 5:24 PM

Well, vast swaths of the Kangxi dictionary (which serves as "sources" for probably most of the CJK characters) are such "ghost" characters as described in the article...

The peculiar properties of CJK characters and the philosophy (apparently the Japanese did not like Unicode's tendencies towards Aristotelian essentialism) under which they were implemented in Unicode probably singlehandedly forced unicode to expand beyond the BMP....

show 3 replies
sedatktoday at 6:14 PM

Fascinating. But, I guess it's better to have superfluous invalid characters than missing real ones.

amaketoday at 11:24 PM

Should probably have "(2008)" in the title

Dwedittoday at 10:46 PM

Saw headline, expected branch prediction vulnerability involving Unicode, was surprised at a completely different topic.

evikstoday at 7:59 PM

> The original character (𡚴) was not added to JIS or Unicode until much later and doesn't display on most sites for me

Why didn't they simly replace the original bad one?

> nine hundred pages. Imagine tracking down a single character without a page reference

Not that hard to imagine, OCR existed back then?

show 2 replies
panzitoday at 7:36 PM

Is anyone using these characters now for anything? No youth language or online slang using it?

show 1 reply
philipovtoday at 5:31 PM

"- the spectre of communism. All the powers of old encoding have entered into a holy alliance to exorcise this spectre..."

show 1 reply