Git 3.0's upcoming SHA-256 default will be a costly mistake
Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,703 words · 1 segments analyzed
Where to begin? I've been sitting on this for a few years, mostly because there are smarter people who have been concentrating on this and I don't love being a back seat driver. However, I think that the Git 3.0 release is about to cost everyone a lot of time and angst for little benefit, and virtually nobody knows what's coming. So grab some popcorn and let me tell you a tale of how one of the new upcoming Git 3.0 breaking changes is about to be a huge, costly, global train wreck of a change for almost no practical value. I'll keep this short as many of you probably know this at a basic level. Git is what's known as a content addressable database. This means that if you want to store and transmit data in it, Git will calculate a hash of the contents and use that in a key/value database as the key (the value being the content). The same content always gets the same hash, globally. Git's use of SHA-1 hashes as keys in it's key-value object database This is nice, because it means that the same file content is never stored twice. There is also a cool property where commits encode the hash of the commit that came before it, which means that this integrity essentially propagates - you can't change the hash of anything without changing the hash of everything that comes after it. This gives it "cryptographic integrity", meaning that hashing the latest commit essentially also hashes potentially millions of file contents, trees and commits that came before it. In Git, this hash function has always been SHA-1. This was what Linus picked in 2005 when Git was started and it's worked pretty well for 20 years - it's relatively fast and impossible in a practical sense for two different files to accidentally hash to the same value. In fact, as far as I'm aware, this has never happened in the history of every file, tree and commit ever made in Git in every repository ever created - billions and billions of them. Mathematically, for SHA-1’s 160-bit output, the birthday bound means that you would need about 1.4 septillion random files (1.4 quadrillion billion files - 1,400,000,000,000,000 billion - it's impossible to effectively describe) in a single project to have file hashes accidentally collide. There is a problem though, which is that mathematically, SHA-1 is now considered semi-"broken" because there have been published collision attacks (SHAttered in 2017, SHA-1 is a Shambles in 2020) - not really practical to exploit in any demonstrated way, but now theoretically possible. So in response to this, after a huge amount of work by very smart people, the upcoming Git 3.0 release is planning to change it's default hashing algorithm from the semi-"broken" SHA-1 to the stronger SHA-256 algorithm. But first, let's pause. Before we dig into this, what does "broken" mean? This is important to understand, because it's probably not what normal people would think "broken" means. From a cryptographic hashing standpoint, broken in this sense essentially means that finding collisions isn't impossible. In other words, if you throw enough money and GPUs at the problem, you can, for some content shapes, come up with two different things that hash to the exact same value, on purpose. What can you do with a broken SHA-lor? Earl-y in the morrrrn... This means that while it's still nearly impossible to accidentally have two reasonable files with different contents, it's not technically impossible to manufacture two different files that hash to the same thing. This means that there theoretically exist attack vectors where someone could replace one file's contents with a malicious version and Git can't tell the difference because the hash is mathematically identical. These papers showed that SHA-1 has a property that in theory can be exploited by modern GPU farms to produce purposeful collisions on the order of a few tens of thousands of dollars today that SHA-256 does not have (google "linear message schedule sha-1 vs sha-256" if you super-duper care...). Sounds really scary and concerning, right? Won't somebody please think of the children!?! Well, not really. But let's take one minute to step back and talk about these collisions. There are two main problems with hash functions when you're considering attacks on them (assuming I had to massively, stupidly, simplify things). One is a "collision attack" and the other is a "second-preimage attack". A "collision attack" is when you know you're going to be an attacker but pretend to be a good guy until you gain trust. You generate two files on purpose with the same hash - one is benign and the other is malicious. You give people the benign one until you're trusted and then switch it with the malicious one because the hash matches and you know Git can't tell the difference. You can even get signed tags or commits on trees that have the benign file and make it look like the bad file was signed. A "second-preimage attack" is when you see a file you want to replace and then create a second file that does something malicious that also happens to match that hash so you can get people to unknowingly pull it down instead. Importantly, this does not mean that the author of the original file needs to also be the attacker. Collision versus second-preimage So, the most important thing that I would like to emphasize before this rant begins is that while second-preimage attacks have a more concerning aspect to them, nearly no widely used hash function ever used is susceptible to it. Git could be using MD5 (considered an incredibly broken hash function) and still be effectively immune from a second-preimage attack. Again, MD5 is what is considered a completely "broken" hash function, which SHA-1 is not - it's much stronger. What do I mean by "effectively immune"? If every one of the roughly 3 billion GPUs on Earth were magically replaced with an RTX 5090, and every one spent 100% of its time doing nothing but MD5, the expected time to brute-force a particular preimage would still be about 16 billion years (~11 billion median), or roughly the age of the universe.3.4×1038checks3×109GPUs×2.2×1011hashes/sec≈ 5.2×1017 seconds ≈ 16 billion years So, any realistic interesting attack vector therefore relies on a collision attack, meaning the person who introduces the original file has also pre-computed the malicious version and intends to inject it after it's accepted. However, let's be massively, unrealistically conservative and even assume for argument's sake that a preimage is easy. Let's say that you found a way to create a second preimage matching the hash of any known file on your laptop in an hour. Now you can target any file you want to replace and come up with a different malicious file that matches the hash effortlessly. Congratulations! You can now pwn any codebase on earth! Egh. Hm, one tiny problem. How do you: get that file to be fetched by people who don't know you and have never pulled a previous version of this content get them to run it in a way that is useful to you Every argument and problem set after this depends on the answers to these questions, yet the answers to these questions are generally not part of the conversation around this problem. Even if we assume that SHA-1 was trivial to break, even if we assume that second-preimage production was possible (or even cheap), it does not make the attack vectors that are available actually very easy to exploit. The reason why is that hashing is not really the mechanism of trust in the SCM world. It certainly provides an element of cryptographic integrity, but it is not what trust is fundamentally based on. Linus literally argues this at the birth of Git. I really think people should not consider the sha1 the "security". The real security is in distribution.Linus Torvalds, 2005 Trust is based on "where do you pull from?" and it always has been. For one minute, let's consider what an actual attack looks like in the world of getting untrusted content into codebases. This is, of course, the worst case scenario of what we're looking at here. Some source code file (or more likely in a hash attack case, binary file) is inserted maliciously without you knowing. It turns out, this actually happens a fair amount. Not because people are spending hundreds of thousands of dollars on GPUs to generate random entropy to throw in after a null byte, so that some binary file can happen to match a SHA-1 checksum. No, it currently happens in the real world because someone socially engineers an exploit to get write access to some npm package used by millions of projects. Now it's not one binary that's difficult to detect, it's every single file in a dependency that is blindly pulled in with no previous checksumming match by every project that has this project in its package.json file. This is maybe a billion times simpler, cheaper and more likely to succeed than trying to brute force a hash collision and manipulate an untrusted fetch. If I wanted to get untrusted code into Android, it's so much simpler to bribe or convince the maintainer of a popular downstream project, take it over and inject difficult to detect code into an already trusted source URL than try to engineer some easy to detect hash collision and put it into a URL that nobody would ever pull from. Do you think there is no tired open source maintainer who wouldn't give over maintenance of a highly used, unloved project for a $40k lump sum payment? Voila, now you don't need to rent GPUs, it takes one day and you can replace any file you want with any content you wish. In other words, hash collision attacks are maybe the dumbest possible way to get untrusted code on a system when unpaid open source maintainers and low-trust package management forges exist. To go back to Linus's argument, I don't pull code from https://github.com/rust-lang/rust because I trust the GPG signature that signed the latest commit SHA and just assume that any random source is fine.