【问题标题】:What are the upper and lower bound for Chinese char in UTF-8?UTF-8 中中文字符的上限和下限是多少?
【发布时间】:2012-02-28 06:47:51
【问题描述】:

我想在python中设置一个包含所有ord()的中文字符:

对于英语,相当于:

english = set(range(ord('a'),ord('z') + 1 ) +
              range(ord('A'),ord('Z') + 1 ))

【问题讨论】:

  • 您不想直接在 UTF-8 中执行此操作,您想生成 Unicode 代码点并将它们转换为 UTF-8。
  • 你或许可以在这里找到你需要的东西:unicode.org/charts
  • 汉字存在于整个 Unicode 中多个不相交的集合中。
  • 中文范围有很多,但有些平台(唉,不是 Python)允许您查询脚本的代码点范围。

标签: python cjk


【解决方案1】:

来自 Unicode 标准(v6.0,第 12.1 节),

在Unicode标准的七个主要块中可以找到汉字,如表12-2所示

Table 12-2. Blocks Containing Han Ideographs

Block                                   | Range       | Comment
----------------------------------------+-------------+-----------------------------------------------------
CJK Unified Ideographs                  | 4E00–9FFF   | Common
CJK Unified Ideographs Extension A      | 3400–4DBF   | Rare
CJK Unified Ideographs Extension B      | 20000–2A6DF | Rare, historic
CJK Unified Ideographs Extension C      | 2A700–2B73F | Rare, historic
CJK Unified Ideographs Extension D      | 2B740–2B81F | Uncommon, some in current use
CJK Compatibility Ideographs            | F900–FAFF   | Duplicates, unifiable variants, corporate characters
CJK Compatibility Ideographs Supplement | 2F800–2FA1F | Unifiable variants

在这些块之外还有一些附加功能:

Table 12-3. Small Extensions to the URO

Range     | Version | Comment
----------+---------+-------------------------------------------------
9FA6–9FB3 | 4.1     | Interoperability with HKSCS standard
9FB4–9FBB | 4.1     | Interoperability with GB 18030 standard
9FBC–9FC2 | 5.1     | Interoperability with commercial implementations
9FC3      | 5.1     | Correction of mistaken unification
9FC4–9FC6 | 5.2     | Interoperability with ARIB standard
9FC7–9FCB | 5.2     | Interoperability with HKSCS standard

要使用集合操作来构造它们的一组序数值,您可以这样做:

chinese = set(range(0x4E00, 0xA000) +
              range(0x3400, 0x4DC0) +
              range(0x20000, 0x2A6E0) +
              range(0x2A700, 0x2B740) +
              range(0x2B740, 0x2B820) +
              range(0xF900, 0xFB00) +
              range(0x2F800, 0x2FA20) +
              range(0x9FA6, 0x9FCC))

但请注意,该集合包含超过 75000 个字符,因此它可能不是最紧凑或最有效的数据结构。

另外,如果您坚持对文字字符使用 ord(),则需要使用 32 位 unicode 文字形式:

>>> ord(u'\U00002F800')
194560

【讨论】:

    猜你喜欢
    • 2016-09-13
    • 2010-12-14
    • 2012-03-20
    • 2011-05-24
    • 1970-01-01
    • 2010-11-06
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多