UTF-63
Β· By Paul BrinkmeierUTF-8 is almost everywhere these days. The web runs on it, most common file formats use it, macOS, Linux and increasingly even Windows systems use it as the default option. It makes sense: It's quite compact for mostly-ASCII text and therefore most keywords and symbols of text-based file formats and protocols.
What if I told you that in this age of memory bandwidth-constrained workloads you are using a wasteful encoding for text storage? Most text these days is 99% ASCII with a few emoji sprinkled in (source: Made it up). Even if we ignore the costly multibyte emoji encodings, using UTF-8, 1 out of every 8 bits of memory is wasted on the encoding, because 7-bit ASCII characters are encoded using a whole byte, i.e. we have an encoding overhead of roughly 15%.
For example, consider the UTF-8 encoding of the string βπππ encodingβ (see here for codepoints):
11110000100111111001100110001000 π 11110000100111111001100110001001 π 11110000100111111001100110001010 π 00100000011001010110111001100011 enc 01101111011001000110100101101110 odin 01100111 g
Colored spans show which code points are encoded by which bits. Neutral (white on black) bits are encoding overhead required by UTF-8. For emoji, the most important medium of modern communication, this encoding is obviously horrible. Every monkey requires a full 32 bits to encode its 21 bits. And even for good old reliable ASCII, useless bits are sprinkled here and there. There's gotta be a better way!
UTF-63
Enter: An entirely new way to encode Unicode!
10000111110110010010000000111110 ππ 11001001001000011111011001001010 π 00100000110010111011101100011110 enco 11111100100110100111011101100111 ding
For our entirely random example, this UTF-63 encoding uses only 16 bytes, a clear win over the whopping 21 bytes required by UTF-8 (that's 30% more!). So how does it work?
Decoding is dead simple: We consider each set of 8 bytes a Big Endian 64-bit code unit. If the most significant bit is 0, the other 63 bits contain nine Basic Latin characters (i.e. ASCII characters) at 7 bits each. If it's 1, they contain any three Unicode code points at 21 bits each (the maximum codepoint is U+10FFFF). When any of the encoded code points is U+0000 NULL, the string ends there (yes, we're bringing null-terminated strings back. So. many. applications. don't allow NULL anyways).
Usage
Admittedly, this page actually still uses UTF-8, simply because most modern browsers don't support UTF-63 yet. To my knowledge, most OSes, desktop applications and programming languages don't either. It wouldn't surprise me if VLC supported UTF-63 subtitles or something though. You can do your part by spreading the word among friends, family, flatmates, coworkers and anybody in between.