comments (5)

  • The author Paul McCann (polm) is one of my favourite programmers out there!

    He’s done awesome work in the Japanese NLP space over the last decade which has really helped me in my language learning projects.

    He maintains a mecab (Japanese tokenizer) wrapper for Python [0], has a book on Japanese NLP written for English speakers [1] and also worked on Spacy at one point [2].

    [0] https://github.com/polm/fugashi [1] https://www.japanesenlp.com/ [2] https://spacy.io/

    joshdavham

  • I think there’s evidence found for the origin of “彁” as the result of a poor scan of a newspaper article. Look up “彁 新聞” to find some japanese sources about this.

    erjiang

  • Well, vast swaths of the Kangxi dictionary (which serves as "sources" for probably most of the CJK characters) are such "ghost" characters as described in the article...

    The peculiar properties of CJK characters and the philosophy (apparently the Japanese did not like Unicode's tendencies towards Aristotelian essentialism) under which they were implemented in Unicode probably singlehandedly forced unicode to expand beyond the BMP....

    hnfong

  • Fascinating. But, I guess it's better to have superflous invalid characters than missing real ones.

    sedatk

  • "- the spectre of communism. All the powers of old encoding have entered into a holy alliance to exorcise this spectre..."

    philipov