Pangram verdict · v3.3
We believe that this entire text is AI.
AI likelihood · overall
AIArticle text · 1,491 words · 1 segments analyzed
After 145 days of uninterrupted uptime, it was time to type one of those commands that feels slightly more consequential when the machine in question quietly runs half of your digital life: reboot By that point, the server had moved roughly 51 TiB inbound and 5 TiB outbound according to btop. It had been serving websites, storing files, collecting metrics, recording cameras, receiving television, running home automation, resolving DNS, hosting databases, building software, processing weather data, coordinating IoT devices and doing dozens of other small jobs that are easy to forget until the machine is unavailable. And I had just upgraded it from Debian 12 Bookworm to Debian 13 Trixie. A major operating-system upgrade is probably the most brittle state such a server can enter. For a while, you deliberately create a system consisting of an old running kernel, newly replaced userspace libraries, processes still holding old binaries in memory, stopped services, upgraded services, temporarily disabled third-party repositories and a new kernel waiting for its first boot. Yet a few hours later everything was back. Not merely back, either. The machine was running with dramatically lower load, roughly half the memory consumption and more than 10 GB of old packages and debris removed. The interesting part is not that Debian can be upgraded. The interesting part is what this says about running a serious home server on plain Linux, directly on the hardware, without putting Proxmox underneath everything and without reflexively wrapping every daemon in Docker. The server that quietly became infrastructure This is not a dedicated NAS or an old laptop running Pi-hole. It is my main local Linux server and effectively the backbone of my infrastructure. Among other things, it runs: local PHP applications and development environments MariaDB and Valkey Samba and general NAS duties like TimeMachine backups CoreDNS MQTT openHAB and home automation TVHeadend with a Digital Devices DVB card WeeWX and weather-station processing Prometheus monitoring for the wider infrastructure Grafana UniFi Network Server Frigate NVR with video analysis remote and local backups scheduled jobs, synchronization and automation scripts various development tools a collection of Docker containers where containers actually make sense internal software such as Mainframe and PentaPaper It is a complex workload. That does not mean it needs a complex architecture. That distinction matters. A surprising amount of homelab discussion starts from the assumption that every new service should become another Docker container, another VM, another LXC container or another layer under a hypervisor. I use all of those technologies where appropriate. I just don't consider them goals in themselves. Sometimes the cleanest architecture for a service really is: systemd ↓ service rather than: hypervisor ↓ virtual machine ↓ container runtime ↓ container ↓ service Modern Linux is extraordinarily good at running multiple workloads simultaneously. That is one of the things Unix-like systems have spent decades becoming good at. DNS consumes almost nothing. MQTT mostly waits. Web applications are bursty. MariaDB caches useful data and sleeps between queries. openHAB reacts to events. TVHeadend spends much of its life moving streams. Prometheus periodically scrapes metrics. WeeWX wakes up to process measurements. Even something comparatively substantial like Frigate does not mean the rest of the machine suddenly needs its own cluster. Most workloads do not peak simultaneously. That makes consolidation extremely efficient. Bare metal does not mean a snowflake server There is another assumption I increasingly disagree with: that installing software directly on Linux inevitably means endless manual configuration, forgotten files under /etc, mysterious package conflicts and a machine nobody dares to touch after three years. It certainly can mean that. It doesn't have to. This server is managed using Ansible. Its important configuration is infrastructure as code. Services are monitored. Data is backed up. Most software comes either directly from Debian or from a small number of deliberate upstream repositories. The operating model is closer to professional infrastructure than to a collection of shell commands copied from forum posts. That changes the risk model substantially. If a machine fails, the important question is not: Can I somehow repair this exact installation? It is: Do I understand the state well enough to reproduce it? Those are very different situations. It is the same philosophy behind my backup strategy. I have written separately about why I don't blindly consider RAID1 a substitute for backups on a machine like this. Mirroring protects against a particular hardware failure. It does nothing useful when I accidentally delete a file and the deletion is immediately mirrored to the second disk. For some datasets, periodic synchronization to additional drives—some of which can remain spun down most of the time—is a much better match for the actual failure modes I care about. The principle is the same throughout the system: design around the problem, not around the fashionable solution. A major upgrade is where this architecture gets tested Normal operation proves surprisingly little. A service that has been running for two years can continue running because nobody has disturbed it. A major distribution upgrade is different. This is where dependencies move underneath you. At one point during the Debian 12 → Debian 13 migration, the conceptual state of the machine looked approximately like this: old running kernel + partially upgraded userspace + old processes with old libraries mapped + new libraries on disk + some stopped services + some already upgraded services + third-party repositories temporarily disabled + new kernel modules being built + new kernel waiting for first boot This is precisely the moment where configuration debt becomes visible. Old signing keys fail. Kernel modules stop compiling. Services depend on Java versions you forgot about. Python environments point at interpreters that no longer exist. Configuration files have changed upstream. Ansible itself starts warning about patterns that were perfectly normal several years ago. And eventually you reach the point where the only way to know whether the new system actually works is to reboot it. That is a much more meaningful test of maintainability than another 100 days of uptime. Preparing Bookworm for Trixie I did not want APT simultaneously solving Debian's entire distribution transition and half a dozen external repositories. So the first objective was to simplify the package layer. I replaced the old-style Debian repository configuration with a modern deb822 source: Types: deb URIs: https://deb.debian.org/debian Suites: trixie trixie-updates Components: main contrib non-free non-free-firmware Signed-By: /usr/share/keyrings/debian-archive-keyring.gpg Types: deb URIs: https://security.debian.org/debian-security Suites: trixie-security Components: main contrib non-free non-free-firmware Signed-By: /usr/share/keyrings/debian-archive-keyring.gpg I also moved from a fixed German Debian mirror to deb.debian.org. There is nothing inherently wrong with a good national mirror, but Debian's CDN-backed endpoint is a nicer default for infrastructure that should work regardless of where I eventually deploy it. Bookworm backports were removed. Old source-package entries I did not use disappeared. Third-party repositories were temporarily taken out of the equation. Then came the normal staged transition: apt upgrade --without-new-pkgs followed by installing the new kernel and headers, and ultimately: apt full-upgrade The important thing was not to rush the reboot. As long as the old kernel and the existing SSH session remained alive, I still had a very comfortable recovery environment. The first real snag: ZFS and DKMS The new Debian 13 kernel initially failed to configure. Not because the kernel was broken, but because DKMS attempted to rebuild an old ZFS module against it and failed. That failure propagated upward: zfs DKMS build fails ↓ kernel postinst fails ↓ linux-image remains unconfigured ↓ headers remain unconfigured ↓ kernel meta packages remain unconfigured This is exactly why I prefer understanding dependency chains instead of treating an APT error as an opaque wall of red text. The kernel was fine. ZFS was the problem. Updating the ZFS stack to the Trixie-compatible version via apt install zfs-dkms zfsutils-linux and rerunning DKMS solved it. The kernel configured normally afterward. No reinstall. No recovery environment. No mystery. NodeSource met Debian's stricter crypto policy Another interesting failure appeared when bringing Node.js back. Debian 13 rejected the existing NodeSource repository signature because the signing key still had an old SHA-1 certification signature in its chain. APT was quite explicit about it: Policy rejected non-revocation signature because SHA1 is not considered secure Again, the system was doing exactly what it should. The fix was not to weaken APT's security policy. It was to replace the obsolete NodeSource key and repository definition with their current one. I used the opportunity to install Node.js 24 cleanly, together with Yarn 4, rather than preserving an old development runtime for historical reasons. That is one recurring benefit of major upgrades: they create a natural point to ask whether old infrastructure should actually survive the migration. Sometimes the correct migration strategy is deletion. Java got simpler too openHAB had accumulated a small JVM museum over time: Zulu 11 Zulu 17 Zulu 21 OpenJDK 11 OpenJDK 17 ... Most of that was history, not architecture. Debian 13 provides a perfectly suitable OpenJDK 21 environment, so openHAB now simply uses: JAVA_HOME=/usr/lib/jvm/java-21-openjdk-amd64 The old Azul repository could disappear.