Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,029 words · 6 segments analyzed
Why would anyone do this? I saw someone on twitter arguing that saving data in JSON was apparently not what Real™ developers do. Obviously, I had to become a Real™ Developer too. Turns out, the answer was binary. Naturally, I had a brilliant idea: “How hard could it be to make my own binary format?” Surely it’s just a little wb . ( ˶ˆᗜˆ˵ ) It was, unfortunately, not just a wb. I ended up building an entire binary schema language that can shrink JSON payloads by 80%. jBin What is binary packing? Say we got some data "hello world" then it would be translated in ascii to 104 101 108 108 111 32 119 111 114 108 100 h e l l o [SPACE] w o r l d So h becomes 104 in ASCII. Since these ASCII values fit within 8 bits, each character takes up 1 byte. h e l 01101000 01100101 01101100 l o [SPACE] 01101100 01101111 00100000 w o r 01110111 01101111 01110010 l d 01101100 01100100 so we can just write it using a lil bitta c: FILE *f = fopen("file.bin", "wb"); unsigned char data[] = "hello world"; fwrite(data, 1, sizeof(data) - 1, f); fclose(f); Simple enough. Now let’s try writing 104.
Obviously, we could just write 104 as ASCII characters: '1' '0' '4' → 49 48 52 But that’s 3 bytes for a number that only needs 1 byte. So if we want to save those 2 bytes, we need some way of telling the decoder, “hey, this is an integer, not a string.” You could add a header, you would only be adding an additional byte (well depends on how many types you got..
hopefully you don’t have more than 128 types…if you do you got bigger issues. great! lets just use the first byte to represent our type and second to represent our data. [TYPE][DATA] say 0 is int and 1 is string so 104 would be 00000000 01101000 and string would be.. 00000001 01101000 oh wait….. that would only give is h we need a way to represent different lengths of data. welp lets just get another byte. that should represent the length of our string. so now our binary becomes [TYPE][LENGTH][DATA] great! now we can represent our string like this: [TYPE] [LENGTH] [DATA] (STRING) (11) 00000001 00001011 00... h e l 01101000 01100101 01101100 l o [SPACE] 01101100 01101111 00100000 w o r 01110111 01101111 01110010 l d 01101100 01100100 GREAT! now we could pack both strings and ints together! say we wanted to represent "userid": 123 now you could just package it all together [TYPE:STRING][LENGTH:6][WORD:userid][TYPE:INT][LENGTH:-][DATA:123] Great! we can represent 123 as a 1 byte number with 2 bytes of header. but notice, we are not really using LENGTH field for ints? why need it then? waste of bytes eh? WELL… if we get rid of it, how does our binary reader know where the header ends? It needs some way to say “okay, the header is done, start reading the actual data now.” huh. what can we use to represent that a byte is ending. A length byte for the header, perhaps? Ehh. That’s redundant. We’d be removing the length field just to add another length field. But hey, we could use a bit in the header itself. We could have one bit say: I am not the last byte in the header. There’s more. You might think: why not use the LSB? Well, then we’d only be able to represent even numbers. Which is… not ideal. So we’ll use the MSB instead. so now our tag looks something like this: [CONTINUATION BIT][7 BITS OF DATA] If the continuation bit is 1, there’s another header byte. If it’s 0, the header is done. using this, we can just have our 123 be [TYPE=0][DATA=123] and if its a string. [TYPE=1][LENGTH=11][DATA=104]... so its of type 1, length 11 but wait…what if the length is greater than 127? with 7 bits you can only represent up to 127! We use the same thing! but for ints! if the first bit is 1 then the int continues. 128 can be written as: 10000001 00000000 ^ MSB / continuation bit (in big endian) what the binary reader will do: Reads the first byte.
The MSB is 1, so there’s another byte. The remaining 7 bits are 1. Reads the second byte. Its MSB is 0, so this is the last byte. Its remaining 7 bits are 0. Combines the two 7-bit values to get 128.
This is a kind of varint (variable-length integer). The encoding we’re using here is little-endian: the least-significant 7 bits come first. 10000000 00000001 Hey this is great, innit? You can represent different types in the same binary and your binary parser will read them all correctly But notice, We are storing this data per field. [TYPE][DATA] [TYPE][DATA] [TYPE][DATA] And most data isn’t just a bunch of random values floating around. It’s usually structured. Take a C struct: struct { int i; char *s; int a[10]; } this would be say on a 32 bit system. [32-bit int] [32-bit pointer] [10 × 32-bit ints] and we didn’t have to add headers everytime. because we know the type of the data from the struct itself. Hmm. I wonder if we can do this for our binary data… And yes, we can. That’s what a schema is! so for our struct our schema can just be: i: int s: char * a: list(int) The schema lets us know the type without storing the type alongside every value. now our binary format doesn’t need to worry about the type! it only need to worry the size of the data! That’s what protobuf does So lets think about all the different sizes of data we can have. we got ints, we got floats, bools, strings. We can treat ints and bools as varints, while floats are fixed-width: f32 or f64. Strings are different.
We can’t just encode their bytes as a varint, because the bytes themselves are the actual data we need to preserve. So instead, we need to know how many bytes belong to the string before we start reading it.