I was hoping someone's done it before, but apparently not, so here's what I've ended up with. The module below (I call it unicodedata2
) extends unicodedata
and provides script_cat(chr)
which returns a tuple (Script name, Category) for a unicode char. Example:
# coding=utf8
import unicodedata2
print unicodedata2.script_cat(u'Ф') #('Cyrillic', 'L')
print unicodedata2.script_cat(u'の') #('Hiragana', 'Lo')
print unicodedata2.script_cat(u'★') #('Common', 'So')
The module: https://gist.github.com/2204527
与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…