Skip to content
HN On Hacker News ↗

Coding Theory: A Playful Introduction #SoME5

▲ 31 points • 1 comments • by vismit2000 • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,466
PEAK AI % 0% · §1
Analyzed
Sep 21
backend: pangram/v3.3
Segments scanned
1 windows
avg 1466 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,466 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

This post is designed for people with absolutely no idea about coding theory. Refer to the introduction in case you want to skip the basic stuff. Switch to light-mode and a bigger display for better experience. A Problem Imagine you are a smol child in 1900s; you want to talk with your best friend at late night who is also your opposite neighbour. You don’t want to disturb others (or perhaps wanna talk in secret), so you try speaking softly but the distance is too long to reach them. Thankfully, both of you can see each other from your bedroom windows. The problem is how will you two communicate? Each of you has a flashlight 🔦 that you can use. Think about it!1 Codes for Communication Attempt 1 (Drawing) Well of course, you turn ON your flashlight and start drawing letters, you create the shape of \(\textrm{I}\) with one single vertical stroke of your flashlight and then your friend can see the line and understand it. then you start drawing \(\textrm{L}\) with one vertical stroke and an horizontal stroke below it from where you ended and realise that it will look as $\textrm{L}$ to your friend, so you need to draw $\textrm{L}$ then your friend interprets it correctly as \(\textrm{L}\). then you draw an oval \(\textrm{O}\) but wonder what if its interpreted as \(\textrm{0}\), then you send \(\textrm{V}\) with two simple storkes, then when you send $\textrm{E}$ you soon realise a bigger problem, interpreting symbols with many strokes isn’t easy afterall the old strokes dont stay in the air until you finish the letter. So you wanted to do something more precise. Attempt 2 (Blinking) Till now, your flashlight was always ON, you realise you haven’t turned it OFF since you started and then it clicks, what if you tried blinking (turning your flashlight ON and then OFF). To keep it simple, you assign \(\textrm{A}\) as \(1\) blink, \(\textrm{B}\) as \(2\) blinks, and so on \(\textrm{Z}\) as \(26\) blinks. This works! To send \(\textrm{I LOVE}\), you first blink your flashlight \(9\) times then wait for some time and blink it \(12, 15, 22\) and \(5\) times. Your friend counts the number of blinks (\(\#\)blinks) and computes the letter sent. Of course, you need to make sure there is sufficient pauses between blinks, otherwise \(\textrm{I L}\) can be interpreted as \(\textrm{U}\) with \(9+12=21\) blinks. also, you will need different pauses between \(\textrm{I}\), \(\textrm{L}\) and \(\textrm{L}\), \(\textrm{O}\) as one separates words and one separates letters of same word. Now \(\textrm{I LOVE}\) is \(9+12+15+22+5=63\) blinks, and it gets the job done. We can do better, but first celebrate as you have just discovered Coding Theory 🎊 Introduction What you developed was a code, i.e., a system for transferring information (in this case among people). This code helped you communicate the message (something to be sent) by converting it into an encoded message (something that’s actually sent). A message is made up of symbols (letters in our case) and an encoded message is made up of codewords (\(\#\)blinks in our case). When you were converting a letter into \(\#\)blinks, you were encoding and your friend was decoding when they were converting \(\#\)blinks back into the letter. Coding Theory is the study of such codes. This shouldn’t be confused with the popular term with the same name ‘coding’ which means writing a computer program i.e., instructions for a computer. Now, let’s get back to the problem. Attempt 3 (Frequency Analysis) This discovery is great and as a result you want to send \(\textrm{I LOVE CODING THEORY}\) next, well guess what, it turns out to be \(206\) blinks, too long :( But then, here you find your first breakthrough, you realise there is no need to map \(\textrm{A$-$Z}\) to \(\textrm{1$-$26}\) sequentially. From your futuristic experience of Scrabble, you know that letter \(\textrm{E}\) comes up the most in english langauge, so why not assign it \(1\) blink instead of \(5\). and we can go ahead and assign second most frequent letter of the alphabet \(\textrm{T}\) as \(2\) blinks, to the third most frequent \(\textrm{A}\) as \(3\) blinks, and so on until the least frequent \(\textrm{Q}\) and \(\textrm{Z}\) as \(25\) and \(26\) blinks respectively, this ensures that we use fewer blinks for popular letters and hence can send our message faster. Below plot, gives an idea of frequency of each letter. Convince yourself that for any two pair of letters, it is better to represent the more frequent letter with lesser \(\#\)blinks to minimise the expected total number of blinks. Frequency Analysis ($\%$) of English letters (Image by Nandhp Public domain, via Wikimedia Commons) Letter \(\#\)blinks Letter \(\#\)blinks \(\textrm{A}\) \(3\) \(\textrm{N}\) \(6\) \(\textrm{B}\) \(20\) \(\textrm{O}\) \(4\) \(\textrm{C}\) \(12\) \(\textrm{P}\) \(19\) \(\textrm{D}\) \(10\) \(\textrm{Q}\) \(25\) \(\textrm{E}\) \(1\) \(\textrm{R}\) \(9\) \(\textrm{F}\) \(16\) \(\textrm{S}\) \(7\) \(\textrm{G}\) \(17\) \(\textrm{T}\) \(2\) \(\textrm{H}\) \(8\) \(\textrm{U}\) \(13\) \(\textrm{I}\) \(5\) \(\textrm{V}\) \(21\) \(\textrm{J}\) \(23\) \(\textrm{W}\) \(15\) \(\textrm{K}\) \(22\) \(\textrm{X}\) \(24\) \(\textrm{L}\) \(11\) \(\textrm{Y}\) \(18\) \(\textrm{M}\) \(14\) \(\textrm{Z}\) \(26\) Using this table, \(\textrm{I LOVE CODING THEORY}\) is shortened to \(138\) blinks, a whopping \(33\%\) reduction! Punctuation Before optimising our code further, let’s discuss the important topic of punctuation. When we are sending our blinks, there are actually three levels of pauses that we need to take between blinks for accurate decoding, these are pauses between blinks of same letter pauses between blinks of different letters pauses between blinks of different words The technical term for these pauses is punctuation and it has been an important part of our codes till now. Attempt 4 (Morse Code) Let’s try to formalise our previous system. Instead of writing blinks over and over again, we can use codewords to represent our encoded message, so the letter \(\textrm{A}\) means \(\bullet\), \(\textrm{B}\) means \(\bullet\bullet\), \(\textrm{C}\) means \(\bullet\bullet\bullet\), and so on. This is essentially Base \(1\) system of counting. Where, we literally have same number of \(\bullet\)’s (the length of the codeword) as number of blinks, akin to how ancient people used number of sticks to count their number of sheeps. Notice that in this way, both our message and the corresponding encoded message can be represented by different sets of symbols, for our message the symbols are the alphabets whereas for the encoded message, the symbols are the dots, each specific collection of such dots form a codeword. Here, we were using only one symbol \(\bullet\), but what if we use two symbols instead? That’s Base \(2!\) With two symbols (say \(\bullet\) and \(-\) (called bits)) the possibilities explode, initially we had one codeword for every length, now there are \(2^n\) codewords with length \(n\), for eg, for length \(2\), possible codewords are \(\bullet\bullet\), \(\bullet-\), \(-\bullet\) and \(--\). As, we now have many more codewords of short length, this will again shorten the average length of the codewords and in turn, hopefully reduce \(\#\)blinks. But, what does the symbols \(\bullet\) and \(-\) represent here? Well, in Morse Code, people denote \(\bullet\) by a short blink (dot) and \(-\) by long blink (dash), in particular, a blink in a dash is three times as long as a dot. Here, every letter of the alphabet can be written as a codeword comprised of dots and dashes as shown in below table Letter Symbol Letter Symbol \(\textrm{A}\) \(\bullet -\) \(\textrm{N}\) \(-\bullet\) \(\textrm{B}\) \(-\bullet\bullet\bullet\) \(\textrm{O}\) \(---\) \(\textrm{C}\) \(-\bullet-\bullet\) \(\textrm{P}\) \(\bullet--\bullet\) \(\textrm{D}\) \(-\bullet\bullet\) \(\textrm{Q}\) \(--\bullet-\) \(\textrm{E}\) \(\bullet\) \(\textrm{R}\) \(\bullet-\bullet\) \(\textrm{F}\) \(\bullet\bullet-\bullet\) \(\textrm{S}\) \(\bullet\bullet\bullet\) \(\textrm{G}\) \(--\bullet\) \(\textrm{T}\) \(-\) \(\textrm{H}\) \(\bullet\bullet\bullet\bullet\) \(\textrm{U}\) \(\bullet\bullet-\) \(\textrm{I}\) \(\bullet\bullet\) \(\textrm{V}\) \(\bullet\bullet\bullet-\) \(\textrm{J}\) \(\bullet---\) \(\textrm{W}\) \(\bullet--\) \(\textrm{K}\) \(-\bullet-\) \(\textrm{X}\) \(-\bullet\bullet-\) \(\textrm{L}\) \(\bullet-\bullet\bullet\) \(\textrm{Y}\) \(-\bullet--\) \(\textrm{M}\) \(--\) \(\textrm{Z}\) \(--\bullet\bullet\) Notice again that \(\textrm{E}\) is just a single \(\bullet\), the shortest possible codeword, similarly \(\textrm{T}\) and \(\textrm{A}\) are also given pretty short codewords which are \(-\) and \(\bullet-\) respectively meanwhile \(\textrm{Q}\) has \(3\) dashes and a dot making it the longest letter in terms of \(\#\)blinks, suggesting that certain kind of frequency analysis was considered while designing this code. We have a total of $2+2^2+2^3+2^4=30$ four-letter Morse Code combinations, but only 26 English alphabets. This leaves room for few accented characters like Ä, Ö, Ü and Ş to get shorter codewords compared to other accents. Unlike, previous codes decoding this code is slightly (but not too much :) challenging. The following tree helps in decoding received codewords back to messages, it is essentially the previous table but converted into a tree. once we receive a codeword, we start from the root of the tree at the extreme left and go to above branch for each dot and below branch for each dash; the letter we settle at after the codeword is done is the corresponding symbol of message. Morse Code Decoding for English letters Let’s try to decode the text below, I have added appropriate punctuation of length \(1\) dot, \(1\) dash and \(2\) dashes to distinguish between symbols, letters, and words respectively