Digital Audio on the ZX Spectrum’s 1-Bit Beeper
Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,658 words · 1 segments analyzed
It’s finally time for me to wrap up this ZX Spectrum series, with a look at how to get the simple 1-bit beeper on every Spectrum model to emit high-quality digital audio. The results are quite good: it’s the last sound played on last week’s demo reel, and the difference between it and the simple PCM playback system before it is night and day. I have also, at this point, gotten hardware verification of the demo program on both period hardware and modern clones: the vagaries of the Internet mean that identification is usually by handle, but “gmc” via Mastodon was able to run it on a 48K Spectrum, and Tom Harte got results from the Omni 128 system, which is a modern rebuild of the system. Emulation support is also quite widespread; I used Fuse to capture the demo above, but it also works fine on, at minimum, EightyOne, Zesarux, Spectaculator, and Clock Signal. The practicality of the technique is a bit limited; the sound is very soft on hardware without an external amplifier, and the RAM required for samples of the length and quality I’m experimenting is effectively “all of it.” But it does work. Getting it to work, though, was a journey. It took me two complete rewrites before I landed on something that worked at all, and even then it barely fit into the timing constraints I’d set for myself. I had to pull out several techniques I haven’t used before on this system. I mostly work in assembly language here, just because it’s generally more comfortable to work there on the systems I’m playing with—as long as there’s suitable compiler support like there was on DOS you usually can do just fine in C. Today, I would not have done just fine in C. Even leaving aside how I needed exactly timed instruction sequences for the audio effects, the timing constraints were so tight and the register pressure so fierce that I needed direct chip access to make it work. This isn’t even the hardest form of the problem, either, and there are some straightforward ways to improve the technique I use here which I don’t need for my tests. I’ll be wrapping up today by looking at Dmitry Milk’s Spectrum sound engines as well as Michael J. Mahon’s RTSynth system for the Apple IIe; they’re solving a slightly different problem from the one I faced, so even though our sound output approaches are similar, they end up more similar to each other than to my playback system despite being on totally different CPUs. Prior Work I first ran into 1-bit digital samples by way of early-1990s DOS games. In particular, the illustrated text adventure Eric the Unready could play sound effects through the PC speaker with a technology its manual called “RealSound,” and Star Control II: The Ur-Quan Masters used Amiga-style tracker music for its soundtrack and could deliver it out your PC speaker just as cheerfully as it could your SoundBlaster. The PC’s 1-bit speaker is markedly easier to program than the Spectrum or the Apple II; I went through both simple tone generation and PWM playback about a decade ago. Everything there is still functional and accurate as far as I know, but we don’t really need it for what we’re doing today. Its approaches do motivate our design and strategy, though. The fundamental physical principles here are that circuits have capacitance and speakers have mass, so it takes time for a 1-bit signal to actually travel from 0 to 1 or vice versa, and then it takes more time for the speaker’s cone to actually move across its own space to actually produce the sound waves we hear. If we interrupt the signal partway through, the signal and the speaker won’t make it all the way across the space and we’ll get a more finely-varying sound wave. This was unreasonably easy on the PC—while the speaker’s setting can be directly commanded via an I/O port the way the Spectrum’s is, it could also be driven from a highly-programmable hardware timer that ran independent of the system clock interrupt. This meant that instead of having to constantly baby-sit the beeper with cycle-counted code, you could just configure timing information only when you needed to change the signal being sent. For normal beeping that’s just when you’re changing frequencies, but you could also configure one-shot pulse widths with microsecond precision. That meant that you could set a CPU timer interrupt at 16kHz, or whatever your sample rate was, and then set the sound timer to run for a time based on that sample’s PCM value. The timer chip itself thus ended up serving as a sort of digital-analog converter. We don’t have a timer on the Spectrum, either for enforcing a sample rate or for configuring a pulse width, but we should be able to do both of those things in software directly. The general idea will be to have a loop that lasts as long as each sample, and it will turn the speaker on then off each loop, with two short but variable-length delays on either side of “turn the speaker off again.” The first delay sets the pulse width, and the second makes sure that the total block of code takes the same amount of time no matter how wide the pulse was. Misguided Attempt 1: Actual Delay Loops My first implementation attempt started with some napkin math. My PC playback system accepted pulse widths that ranged from 1 to 127 microseconds. An 8kHz sample rate means that we need a pulse every 125 microseconds. This is the first point where I lift an eyebrow; my test program used a 16kHz sample rate which means I would be overrunning myself if the sample were really using the whole range. Perhaps my sample was softer than I thought. All the same, it’s broadly the same sample here so we should be in reasonable shape. The Spectrum CPU is about 3.5 MHz, so 125 microseconds is 438 cycles. 438 cycles is comfortable enough that it fits 15 iterations of our default delay loop: We need 18 cycles to turn the speaker off after turning it on, so our delay range goes up to 13*15+2+18=215 cycles. That’s close enough to the 16kHz mark itself that I’m motivated to stay in 8kHz mode for my first experiments. I already have a version of the audio clip that’s 7-bit audio at 8kHz; I used it for the NES awhile ago. There’s plenty of time in this 8kHz window to just use that data directly and convert it into 4-bit PCM on the fly. The pseudocode for the first draft turned into this: Read a sample byte, N, and truncate it to 4 bits. Turn on the speaker. Run N iterations of a do-nothing loop. Turn off the speaker. Run 16-N iterations of the same do-nothing loop. Increment sample pointer and quit if we’re at the end of the sample. Wait an appropriate amount of time until we hit our 125 microsecond mark and then return to step 1. The key insight here is that while steps 3 and 5 might take variable amounts of time individually, together they would always take the same amount. Unfortunately for me, this build didn’t work at all; all I got was a piercing high-pitched whistle. I concluded at the time that I was delaying too long—we were not succeeding in interrupting the speaker mid-transition and were instead just producing an 8kHz tone. Truncating to shorter bit widths didn’t really help, either. Misguided Attempt 2: Dynamic Delays My first thought was that I was having a “constant factor” problem—the sample was pretty soft which meant mostly samples in the 6-10 range, and maybe I needed to get those numbers down. I didn’t really know what I should be aiming for, though, and jumps of 13 cycles seemed a bit excessive. I decided to rework it to be finer grained. Way back at the start of the blog I looked at fine-tuned programmable delays on the 6502, borrowing a technique I first saw on the Atari 2600. There were some CPU-specific shenanigans there that we can’t replicate on the Sinclair’s Z80, but we also don’t really need to. Those tricks allowed the 1MHz 6502 to get cycle-exact precision, and that means 4-cycle precision will be good enough for us. The general technique was to lay down a long stream of delay instructions and achieve a variable delay by doing a computed jump into the middle of the sequence. If we have that “long stream of delay instructions” just all be NOP instructions, we’ll get 1.14-microsecond precision in the changes of our delay. The result was still mostly that 8kHz tone, but now I could just barely hear a whisper of the sound clip playing behind the tone. This was a hint that the technique would work, but still no sense that it would function acceptably or even that any emulators would produce good results, or if success or failure on emulation could mean anything about what real hardware would do. At that point I was stuck and started asking more knowledgeable enthusiasts. This was where I learned about Dmitry Milk’s demos, and that they worked both on hardware and in the Fuse emulator (which was the one I was using). I also learned that the physical speaker on the Spectrum’s beeper was much smaller and lighter than the on on the PC. That at least lent some credibility to the idea that I was wildly overshooting the length of my puzzles. This also meant that having each value in my 4-bit signal correspond to an extra 4 cycles gave us a dynamic range of about 20 microseconds on our pulses, and that might be fine as is. I now knew that the goal was possible. What I needed now were some actual numbers to aim at. Finding a Path