Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,541 words · 1 segments analyzed
In our previous blog post in this series, we introduced tine for building immutable images. Now that we have images, we need a means to securely distribute them. In this article we’re introducing Quarry, our “dumb as rocks” software delivery toolchain! Software Delivery When one stops to consider what makes Linux systems unique, one of the first things that springs to mind is their software delivery mechanisms; traditionally via package managers. A global system for managing software installation and updates for an entire system was a radical idea. Many distributions ended up building their own tools to manage and distribute software. Some folks argue that this fragmentation has hurt Linux, but this situation is arguably inevitable as the different tools distributions have built reflect the differing goals of those projects. At Amutable, our goal is to deliver high-integrity Linux systems. This is reflected in all aspects of what we do, and software delivery is no different. Key Requirements Let’s now take a look at our key requirements. Be dumb. The protocol must not rely on a smart server to generate responses to clients; the update payloads and metadata should be static blobs served via "dumb" static hosting and the "smarts" should all live in the client. This makes mirroring very easy to implement, caching quite flexible and cheap (static blobs are a CDN's bread-and-butter with egress costs ranging from "insanely cheap" to "free"), and gives us a lot of flexibility for future alternative transport mechanisms. Provide granular ownership. The signature and general distribution model must allow for different "owners" of signed data. The major pitfall we really want to avoid is the "third-party signing key" model where we are forced into acting like a certificate authority for all third-party data. Aside from being a needless administrative overhead, such models reduce the overall security of authorisation schemes because an object signed for a vendor often becomes authorised for all machines. Allow for flexible key escrow. As part of granular ownership it should be possible to move between key escrow by us (or some other party) and full key ownership without any impact on clients. This might seem like a given but it is important to keep in mind because hardware-backed signing keys are often not exportable or rely on bespoke protocols that may not be interoperable. Also, not all signing models allow for the entity that signs something to be completely replaced safely. Be autonomous. The update scheme must be robust and descriptive enough for machines to be able to perform completely autonomous updates, while still providing all of the mechanisms people expect from modern deployment systems: staged rollouts and blue-green deployments are probably the most obvious. Support granular addressability. In addition to being able to push generic operating system updates to millions of machines, it must be possible to push to more fine-grained groups of machines (even to individual machines) so that we can also use our software delivery mechanism as a more general control and provisioning mechanism. Operate on a need-to-know basis. The flip-side of granular addressability is that the structure of every organisation's fleet, as well as their general deployment scheme, is now represented directly in the structure of the repositories. A single node getting compromised should not permit an attacker to glean any more information about the wider organisation the node is owned by — it should only have access to provisioning data directly pertaining to it. Be secure. It goes without saying, but the scheme we use must be safe against all known attacks against similar software distribution schemes and must be built using sound security principles. Some Background To understand how Quarry works, we’ll first need to take a look at the projects it leverages for installing and distributing artefacts. systemd-sysupdate systemd-sysupdate is a low-level installer primarily focused on installing and updating Unified Kernel Images (UKIs) and Discoverable Disk Images (DDIs) and provides a generic scheme for immutable image-based operating system updates. It supports installing DDIs by flashing them to disk partitions, installing DDIs as system extensions (sysexts and confexts), installing UKIs to the EFI boot partition, as well as providing other rather generic facilities for installing files from remote sources to local targets. Note: In the rest of this article “sysupdate” will refer to the general scheme used by systemd-sysupdate while systemd-sysupdate will refer to the actual program. When compared to other update systems for Linux, sysupdate has a fairly unique mechanism for describing properties of update sources and targets. These properties include: what payloads exist; how payloads are versioned; where new payloads should be pulled from; how payloads should be installed; and which payloads need to be updated together in groups. All of these are represented by on-disk transfer files. The Update Framework (TUF) In the mid-2000s there were a few high-profile attacks against software repositories which led to a few researchers taking a higher-level look at what kinds of attacks should be protected against in theory and what actually are protected against. They found that many software delivery systems had fundamental security issues and lacked robust mechanisms to recover from key compromises. At the same time, the Tor Project was working along similar lines on an update scheme that needed to be protected against all sorts of man-in-the-middle attacks. A joint paper was written describing the design and The Update Framework (TUF) was born. Why TUF? There were quite a few aspects of TUF which made us conclude that it was a very solid base for Quarry. One of TUF’s primary areas of focus is the problem of how to deal with key compromises of software repositories (the seminal paper was even titled “Survivable Key Compromise in Software Update Systems”.) They recognised that most software distribution schemes have a flawed ownership model where single signing keys are used for several distinct purposes, reducing the security of the scheme due to the muddling of rare-but-high-risk operations and common-and-lower-risk operations. TUF instead separates the different stages of the artefact publishing flow into distinct roles with segregated signing keys. The core idea is that the entity signing an object should be the one with the most information about its correctness. TUF then provides very clear procedures for how the compromise of any role’s key can be safely resolved in-band. The targets role signs lists of artefact names, sizes, and hashes. In the traditional software repository model, the idea is for the targets role keys to be owned by the actual developers. It can also delegate to other custom roles (with different signing keys) to allow for many different parties to distribute software in a repository dynamically. (In TUF-speak artefacts are called “target files”, hence “targets role”.) The snapshot role signs a snapshot of the targets roles’ metadata (versions and hashes) to ensure that clients see a consistent snapshot of all of the software in the repository. This is usually owned by the publisher of a repository. The timestamp role just signs a version and hash of the snapshot role. All metadata in TUF has an explicit expiry to defend against freeze attacks. Having a very small description in the form of a timestamp role allows for very short TUF metadata expiries without clients needing to needlessly re-pull other role data. This is usually owned by a publisher of a repository. The root role signs a description of the set of public keys allowed to sign for the top-level roles (as well as the threshold of keys needed for each). This is owned by the publisher of a repository but the keys are almost always stored offline as they are the most security-sensitive keys in TUF. The following diagram may make this clearer.1 Rendering diagram… The best part is that TUF is entirely implemented using static files — a perfect match for our need for a “dumb” protocol that is very easily mirrored! The authors of TUF also spent a lot of time thinking about common pitfalls in software distribution: freeze attacks, mix-and-match attacks, endless stream attacks, fast-forward attacks, etc. This mirrored our own experiences with software distribution protocols that had made us keenly aware of the menagerie of vulnerabilities such protocols are liable to have. Introducing Quarry After evaluating other schemes and briefly considering building our own, we concluded TUF would be a great foundation to build upon. So behold, Quarry — our TUF-based client and publishing tool. Quarry fills in a lot of gaps in upstream tooling (such as providing a properly managed programmatic and transactional publishing flow) and also includes a few TUF extensions we found indispensable. Per-machine Repositories In order to facilitate the granularity we need, every machine needs to have a dedicated source of installable artefacts targeted to them. Machines that are grouped together with shared artefacts will also likely want to have them deduplicated by grouping them into dedicated sources per grouping. Unfortunately, the only really natural way of doing this with TUF is quite top-down — each organisation would have a single repository with each group or machine represented as a subdirectory, potentially using delegations. This is the polar opposite of being need-to-know; structurally, every machine would know about how every other machine in an organisation is configured! There are also some more in-the-weeds issues that make this approach unsavoury.