Chinese Character Code for Information Interchange
Early traditional Chinese encoding used in libraries, precursor to Unicode.
The Chinese Character Code for Information Interchange (CCCII) is a character set developed by the Chinese Character Analysis Group in Taiwan, first published in 1980 and expanded in 1982 and 1987. It is used mostly by library systems and is one of the earliest established and most sophisticated encodings for traditional Chinese, predating Big5 and CNS 11643. It is distinguished by its unique system for encoding simplified versions and other variants of its main set of hanzi characters.
- first_published
- 1980
- expanded
- 1982, 1987
- developed_by
- Chinese Character Analysis Group (Taiwan)
- primary_use
- Library systems
- character_set_size_1987
- approximately 53940 code points
- character_set_size_1989_draft
- 75684 code points
- variant_used_by_Library_of_Congress
- East Asian Character Code (EACC)
Quick Facts
- Standard
- MARC-8, ANSI/NISO Z39.64 (both EACC version)
- Classification
- TBCS for CJK based on the ISO 2022 structure, JACKPHY component of MARC
- Lang
- Chinese, Japanese, Korean
Facts from the source article.
Lore & Background
CCCII was designed as a 94n set per ISO/IEC 2022, with each Chinese character represented by a 3-byte code. Its maximum theoretical capacity is 830,584 characters, though practical encoding is less due to reserved spaces for variant characters. The 94 ISO 2022 planes are grouped into 16 layers, with layer 1 containing non-hanzi and hanzi, layer 2 simplified characters, and layers 3–12 further variants. Layers 13 and 14 add Japanese and Korean support, while layer 16 holds other characters. A variant of an earlier version, the East Asian Character Code (EACC), is used by the Library of Congress as part of MARC-8.
Reader's Guide
CCCII's significance lies in its early and sophisticated encoding of traditional Chinese characters, including a systematic approach to variant forms. It was a direct precursor to Unicode: work at Apple based on the Research Libraries Group's CJK Thesaurus, used to maintain EACC, was incorporated into Unicode's Unihan set. Despite its innovative design, adoption was limited to libraries in the United States, Hong Kong, and Taiwan due to specialized hardware requirements and difficulty in determining root versus variant characters. As of 2009, EACC remained in extensive use for bibliographic purposes. The structure of CCCII has been both praised for its thoughtfulness and criticized for its fixed, one-dimensional approach to variant relationships. Unicode hanzi characters are referenced to CCCII and EACC codes in the Unihan database, though mapping is not complete due to differing unification criteria.
Did You Know?
- CCCII predates Big5 (1984) and CNS 11643 (1986).
- EACC contains fewer characters than the most recent versions of CCCII.
- CCCII defines roughly 53940 code points as of its 1987 edition.
- The code 0x212320 is used by some implementations as an ideographic space.
Origins and Structural Architecture
Developed by the Chinese Character Analysis Group in Taiwan, CCCII first appeared in 1980 and saw major expansions in 1982 and 1987. It stands as one of the earliest and most technically refined encodings for traditional Chinese, arriving before both Big5 in 1984 and CNS 11643 in 1986. The system is built on the ISO/IEC 2022 94n framework, meaning every character is encoded as a three-byte sequence where each byte is a 7-bit value falling between 0x21 and 0x7E. This yields a theoretical ceiling of 94 cubed, or 830,584, code points, though the actual repertoire is smaller because variant glyphs consume related planes. The 94 planes are organized into sixteen layers of six planes apiece, with the final layer holding only four. Layer 1 carries non-hanzi symbols and the most frequently used hanzi; layers 2 through 12 accommodate simplified forms and additional variants; layers 13 and 14 serve Japanese kana and Korean hangul respectively; layer 15 is reserved; and layer 16 holds miscellaneous characters. The 1987 edition defined roughly 53,940 code points, and a 1989 draft pushed that figure to 75,684.
The Layered Variant Character System
CCCII's most distinctive innovation is its treatment of character variants through a homologous layering scheme. Layer 1 holds the primary traditional Chinese forms, while layer 2 places their simplified counterparts at identical row and cell coordinates. Layers 3 through 12 extend the same positional alignment to further variant shapes, so a single conceptual character can be addressed across multiple layers at the same grid position. This multi-dimensional mapping was praised by Ken Lunde, who called it one of the most well thought-out character set standards from Taiwan and said its structure was to be truly admired, though he noted that OpenType variant form substitution could deliver comparable functionality. Christian Wittern of Hanazono University offered a sharper critique, arguing that the relationships among character variants are too intricate to be captured in a fixed, one-dimensional, hard-wired codetable. The scale of the challenge is evident in the 1989 draft, which listed 44,167 unique characters alongside 31,517 variant entries, underscoring just how deeply the layered architecture had to penetrate the complexity of Chinese script.
Adoption, Limitations, and the Library Domain
Despite its technical depth, CCCII never achieved broad consumer adoption. Its primary domain was library and bibliographic systems. By 1995, the encoding was in use mainly in libraries across the United States, Hong Kong, and Taiwan. A trimmed variant, EACC (East Asian Character Code, designated ANSI/NISO Z39.64), was folded into the Library of Congress's MARC-8 framework, where it was assigned the private-use F-byte 0x31 under the ANSI X3.41 implementation of ISO 2022. EACC contained only 15,686 characters, a modest subset of CCCII's full repertoire. Several practical obstacles constrained wider uptake: reliance on specialized hardware, the difficulty of deciding when a root form versus a variant should be displayed, and the absence of firmly established reference glyphs to guide those decisions. Outside the library world, Big5 became the dominant encoding for Chinese in those territories, especially before Unicode gained traction. Even as late as 2009, EACC remained in extensive bibliographic use, and the Library of Congress published mapping tables linking EACC to Unicode for hanzi, hangul, kana, and punctuation.
Precursor to Unicode and Enduring Cross-Reference Legacy
CCCII's influence on modern character encoding extends far beyond its own operational life. Its most tangible legacy is embedded in Unicode's Unihan database. At Apple, engineers built a CJK character cross-reference database drawing on the Research Libraries Group's CJK Thesaurus, which had been used to maintain EACC. That work was directly incorporated into the development of Unicode's Unihan set. Within Unihan, individual hanzi carry reference keys kCCCII and kEACC that point back to their corresponding codes in those earlier systems, preserving a traceable link to the 1980s encoding. However, because Unicode's unification criteria—shaped by Japanese JIS X 0208 and by the Association for a Common Chinese Code in China—differ from CCCII's own variant-handling philosophy, not every CCCII variant maps to a distinct Unicode code point. CCCII thus occupies a foundational position in the history of CJK encoding, its layered variant logic and cross-reference data having seeded structures that continue to inform how Chinese characters are organized and referenced in today's universal character set.
Frequently Asked Questions
Who created the Chinese Character Code for Information Interchange?
The Chinese Character Analysis Group, a team based in Taiwan, developed CCCII and first published it in 1980. It was later expanded in 1982 and again in 1987 to grow its character coverage.
What sets CCCII apart from Big5 or CNS 11643?
CCCII predates both Big5 and CNS 11643, making it one of the earliest formal encodings built specifically for traditional Chinese. Its most distinctive trait is a built-in mechanism for encoding simplified and other variant forms of the hanzi in its main set.
Where is CCCII actually used in the real world?
Its primary home is in library systems, where it handles cataloging and information interchange for traditional Chinese texts. It is not a consumer-facing encoding but rather a specialized tool for institutional record-keeping.
How many characters does the CCCII set contain?
The 1987 revision brought the set to roughly 53,940 code points. A 1989 draft further expanded that figure to approximately 75,684 code points.
Why do fans of Chinese text encoding history consider CCCII important?
As one of the earliest and most sophisticated encodings designed for traditional Chinese, it established groundwork that later influenced broader standards such as Unicode. It is often cited as a key precursor in the long evolution of how Chinese characters are represented digitally.
More in Taiwanese inventions 1-24
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
