Skip to content
HN On Hacker News ↗

UTF-8000

▲ 134 points • 122 comments • by vismit2000 • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,702
PEAK AI % 0% · §1
Analyzed
Sep 20
backend: pangram/v3.3
Segments scanned
1 windows
avg 1702 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,702 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

Unlimited UTF-8! ASCII ⊆ UTF-8 ⊆ UTF-8000. No special cases introduced. All properties preserved. Try out the reference implementation with $ pipx install UTF-8000. UTF-8000 is in no way endorsed by or representative of the Unicode Consortium. This is a fun standalone project / proposal. TLDR / Examples ASCII 1 0xxxxxxx UTF-8 2 110xxxxx 10xxxxxx 3 1110xxxx 10xxxxxx 10xxxxxx 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx UTF-8000 5 111110xx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 6 1111110x 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 7 11111110 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 8 11111111 100xxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 9 11111111 1010xxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10 11111111 10110xxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx ... 22 11111111 10111111 10111111 10110xxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx ... 10xxxxxx ... There is nothing special-case-y about the example 22-byte code unit here. It is just a good prototypical example, demonstrating the power of UTF-8000 with multiple start bytes. There are only two special cases, both of which are inherited from UTF-8: ASCII as is, and 2-byte UTF-8 having 4 mandatory content bits to check against overlong encoding as opposed to 5 for all longer length code units. Anatomy Here is anatomical diagram of the example 22-byte code unit from the tldr. See the glossary for more information on the definitions of the terms. Byte number four is exciting! It is a continuation byte, a start byte, the final start byte, has content bits, and has only some of the mandatory content bits, which are straddled across the final start byte and first non-start byte. The main contribution of UTF-8000's specification is clarity on splitting the highest bits of the first byte of UTF-8 code units into self-synchronization bits and start bits, and then making it clear how to stripe the start bits across the continuation bytes if needed, to achieve arbitrarily large code units. Glossary These terms are ordered somewhat by chronology of first requirement, rather than alphabetically, for convenience. Terms used within definitions are underlined clickable hyperlinks. Term Definition Codepoint A non-negative integer, aka an unsigned integer. Code Unit A sequence of UTF-8000 bytes that encode a single codepoint. First Byte The first, one and only, byte that begins a UTF-8000 code unit. The self-synchronization prefix of a first byte is either 0 for ASCII or 11 for multi-byte code units. This term is not synonymous with start byte. A first byte is necessarily a start byte, but not the other way around. It is for this reason that first byte is sometimes also known as first start byte. Fun observation: because of the self-synchronization prefix 0 the upper hex nibble of ASCII bytes can only be one of 0, 1, 2, 3, 4, 5, 6, 7. This term is mutually exclusive with continuation byte due to self-synchronization. Continuation Byte A byte beyond the first byte of a multi-byte UTF-8000 code unit. The self-synchronization prefix of a continuation byte is 10, which is also known as the continuation prefix bits. Fun observation: because of the self-synchronization prefix 10 the upper hex nibble of continuation bytes can only be one of 8, 9, A, B. This term is mutually exclusive with first byte due to self-synchronization. Self-Synchronization Prefix The highest bits of every UTF-8000 byte that indicate whether it is a first byte or a continuation byte. The possible self-synchronization prefixes form a prefix-free tree: .----0 First byte for ASCII `----1---0 Continuation byte for multi-byte UTF-8000 `----1 First byte for multi-byte UTF-8000 This piece of the clever architecture of UTF-8, which UTF-8000 inherits, provides the property of self-synchronization at a byte level: we can instantaneously tell what kind of byte we are looking at, and where it should belong in a code unit, just by looking at these highest bits. This is most useful when decoding part of a file encoded in UTF-8000. If we randomly seek through the file to an arbitrary byte, we can unambiguously tell whether we are at a first byte whence we can begin decoding a new code unit immediately, or that we are at a continuation byte whence we need to seek a little further on in order to find the next first byte in order to begin decoding. Nor do we have to process any bytes prior to our seek position in order to discover some global state or the context of the byte we have seek-ed to; a first byte is always unambiguously a first byte wherever it appears, which we can deduce by its self-synchronization prefix being either 0 or 11. This is useful not only for random access, but also for error recovery. Suppose that we are decoding an error-prone stream of UTF-8000 bytes and that whenever when we encounter an error (e.g. a rogue 0xC0 byte) we wish to keep calm and carry on instead of immediately exiting. We can yield Unicode replacement characters U+FFFD and then await the next first byte, discarding anything in the interim. See the Wikipedia article for self-synchronizing code for more general info. These bits are highlighted in bright cyan. Start Byte A byte containing one or more start bits. The start bytes exist contiguously at the beginning of a UTF-8000 code unit. The power of UTF-8000 is that we can have multiple start bytes, to achieve arbitrary code unit lengths, to encode arbitrarily large codepoints. Sometimes it is sensible to colloquially also include ASCII as a start byte when we are talking about the bytes towards the start of a code unit, even though ASCII bytes have no start bits. Every non-ASCII code unit has at least one start byte. The first start byte is the first byte, and it is followed by zero or more continuation bytes that are also start bytes. Therefore because a UTF-8000 code unit can have multiple start bytes, this term is not synonymous with first byte. In restricting to only UTF-8 without UTF-8000, this term is synonymous with first byte. This is because UTF-8-length code units only require one start byte, whether using up to 4 bytes in the current UTF-8 standard (RFC 3629 (2003)), or using up to 6 bytes in former standards (RFC 2044 (1996) and RFC 2279 (1998)). Start Bits The unary-code sequence of bits contained in the start bytes of a multi-byte UTF-8000 code unit that tells us the length of the code unit in bytes. For a code unit made of n bytes the start bits are n-2 1 bits followed by a terminating 0 bit. To be clear, the start bits include this terminating zero bit. Thus the start bits sequence is of length n-1 and looks like 111...10. The possible start bits sequences form a prefix-free tree: .----0 Two byte UTF-8 `----1---0 Three byte UTF-8 `----1---0 Four byte UTF-8 `----1---0 Five byte UTF-8000 `----... n byte UTF-8000 For an n byte code unit where n < 8 the start bits all fit together snugly in the first byte. Otherwise they are striped across as many of the first few bytes as they need, filling the free bits that are not occupied by continuation prefix bits. This is another piece of the clever architecture of UTF-8, which UTF-8000 inherits, that provides the property of self-punctuation also known as a prefix code or a prefix-free code: when decoding a multi-byte code unit, once we have read to the end of the start bytes, that is we have encountered the terminating 0 bit, we know exactly how many bytes we expect in that code unit. Notwithstanding errors we can therefore succeed in decoding the code unit by reading exactly that many bytes, and no more. This avoids a problem of dumber variable-length encodings whose code units do not intrinsically indicate their length: one has to read beyond the last byte of a code unit, that is one reads the first byte of the next code unit, in order to know that the current code unit has finished. For very dumb encodings which have neither self-synchronization nor self-punctuation, to make random access possible one would have to put dedicated auxiliary bytes, punctuation like a comma byte, between code units to be able to tell where one ends and another begins. See the Wikipedia articles for prefix code and unary coding for more general info. This term is mutually exclusive with content bits. These bits are highlighted in bright magenta. Content Byte A byte containing one or more content bits. A byte being a content byte does not imply that it is a continuation byte. For example a 3-byte code unit begins with 1110xxxx, which contains 4 content bits and is not a continuation byte. A byte being a continuation byte does not imply that it is a content byte. For example a 22-byte code unit contains 10111111 as its second byte, which is a continuation byte and has no content bits. Content Bits The sequence of bits in a code unit beyond the start bits and to the end of the code unit, in which the codepoint's binary bits are stored. For example a 3-byte code unit, which has the form 1110xxxx 10xxxxxx 10xxxxxx, has 16 content bits. For ASCII there are 7 content bits. These seven bits xxxxxxx combined with a byte's highest bit being set to the self-synchronization prefix 0 means that ASCII is perfectly included into UTF-8 without being altered. Thus ASCII code units take the form 0xxxxxxx. Otherwise for an n byte code unit, where n > 1, there are 5n+1 content bits. This is how we arrive at that formula: We start with n blank bytes, each of which has 8 bits. For each byte 2 bits are taken by the self-synchronization prefix. Then an additional n-1 bits are taken by the start bits. Thus there are 8n - 2n - (n-1) = 5n+1 bits left for content bits. Another way to think about the 5 in this formula is by extending from n-1 bytes to n bytes by appending another continuation byte. By doing this we gain 6 free bits in the continuation byte, but we lose 1 bit to the longer start bits sequence, thus overall we gain 6-1 = 5 bits for content bits. This term is mutually exclusive