Skip to content
HN On Hacker News ↗

Optimizing x264 Settings and Per-title ladders - Streaming Learning Center

▲ 30 points • 13 comments • by dabinat • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

2 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,654
PEAK AI % 2% · §1
Analyzed
Sep 29
backend: pangram/v3.3
Segments scanned
1 windows
avg 1654 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,654 words · 1 segments analyzed

Human AI-generated
§1 Human · 2%

If you’re producing solely H.264-encoded video, optimizing your encodes can save you bandwidth costs and deliver a higher quality experience to your viewers. If you’re considering adding HEVC or AV1 to the mix, you have another reason. Before you attempt to compute the bandwidth savings and quality improvements these codecs deliver, you should know just how much quality and bitrate you can get from your current H.264 encoder. Otherwise, some of the gain you credit to the new codec is really the gain you left on the table with H.264. To illustrate this, I produced two series of tests with x264 and FFmpeg. First was to identify the optimal encoding parameters for x264 using 1080p clips, then to apply those parameters to a full encoding ladder and compute the overall benefits. For you TL/DR aficionados, the results come first. The rest of the article explains how we got them. You can download a Zip file with much of the results of this testing here. Details of what’s in the Zip are below. Contents The Results (TL/DR)Encoding Cost and BreakevenWhat’s it MeanHow we testedStage One: Testing 1080 to Identify the Optimal Configuration SettingsStage Two: Comparing Fixed and Optimized LaddersOur Fixed Bitrate LadderOur Per-Title LadderOther FindingsWhat’s in the PDF The Results (TL/DR) You’re going to have a ton of questions, but this is the TL/DR section; we’ll answer most of them below. Figure 1. What each step adds to average quality, average bitrate, and encoding time, from x264 defaults on the Netflix 2015 ladder to tuned x264 on a per-title ladder. Click the image to view at full resolution. The TL/DR bottom line? Optimizing x264 boosts VMAF by .66 VMAF points while reducing bandwidth from 4.9 Mbps to 3.27 Mbps, or 33%.  Encoding time is more than tripled. As you can see if you peek ahead to Figure 4, the bandwidth savings exceed encoding costs at 186 hours of viewing, making the decision to optimize a no brainer for all but  the tiniest of encoding shops. Figure 1 starts with x264 at mostly default settings (with a two-second GOP) using a fixed H.264 ladder, and adds one optimization at a time: the veryslow preset, a ten-second GOP, two-pass encoding, three reference frames, and finally a per-title ladder with the same settings. Figure 2. The top-heavy distribution pattern. How did we compute the VMAF and bitrates? It’s complicated. Take a deep breath. Exhale slowly. Repeat. OK. We computed the numbers assuming the distribution pattern shown in Figure 2, which is called the Top-Heavy distribution pattern in the Streaming Learning Center Bitrate Explorer (SBE) shown in Figure 3. Why didn’t we just average the bitrates and VMAF values for all rungs? For several related reasons. First, if we average the rung values, we give equal weight to the top rung and the bottom rung. That doesn’t match the reality of most streaming services, particularly in the US, Canada, Europe, and other high-bandwidth regions where the top two or three rungs predominate. In this case, since we’re comparing a fixed bitrate ladder to an optimized per title ladder, the per-title bottom rungs have much higher quality than the fixed ladder because they’re encoded at higher resolutions with optimized parameters. You see this in Figure 3, where the bottom rung of the optimized ladder (red) has a VMAF score that’s 46 points higher than the fixed ladder, much greater than the modest bitrate differential. If we averaged the rung values, the bottom rung would significantly boost the overall average, which makes no sense given that only a very tiny percentage of viewers would actually watch that rung. Using the Top-heavy distribution pattern, we weight this rung by a much more appropriate .34%. Figure 3. Fixed ladder vs. per-title ladder for the test clip Meridian. While we’re staring at Figure 3, let’s look at the other end of the spectrum, the top rung. Here we see that the top rung of the fixed ladder, encoded at around 5.6 Mbps, has a VMAF score of 96. In contrast, the optimized ladder is encoded at about 1.8 Mbps with a VMAF score of 93. Which looks better to the viewer? Actually, neither, most viewers will consider them about the same. There’s substantial research showing that once a video achieves a VMAF score of 93, additional quality isn’t discernable by the viewer. One key downside of a fixed ladder is that easy-to-encode files often produce one or more rungs that exceed VMAF 93, which wastes storage and bandwidth. This is why we tuned our optimized per-title ladders to a top rung value of 93. In comparing the fixed ladder vs per title ladders, using the top-heavy distribution assumes that 71.6 % of viewers watch that top rung, which significantly boosts the VMAF average for that ladder, even though viewers can’t actually perceive the quality difference between 93 and 96. Just remember a couple of things when you look at the .66 VMAF differential between the baseline and optimal ladders. First, had we averaged the rungs rather than assuming the Top-heavy distribution pattern, the difference would have been much higher (6.5 VMAF points, actually). Second, the .66 delta using this distribution pattern is understated because we’re not accounting for the fact that values over 93 aren’t perceivable by the viewer. OK, breath again. Math lesson over. Encoding Cost and Breakeven As you can see in Figure 1, the optimized ladder increased encoding time by 3.2x. To convert that to dollars, we priced the encoding on an AWS c7i.8xlarge instance at $1.428 per hour, which has the same 32 threads as our test workstation, and bandwidth at $0.02 per GB. Our times were measured on the workstation, so treat the encoding costs as estimates. Scaled from our two-minute clips, the baseline ladder cost about $1.25 to encode per hour of content, and the optimized ladder, at 3.2 times the encoding time, cost about $3.99, an extra $2.74 per hour of content. The Breakeven tab in SBE enables users to compute the breakeven on encoding decisions like the optimized ladder by factoring in bandwidth savings, distribution cost, and encoding cost, using multiple distribution profiles. Figure 4 compares the baseline and optimized ladder using the Top-heavy distribution. Not surprisingly, BBE shows the same numbers as Figure 1 for bitrate  (4.90 Mbps to 3.269 Mbps) and VMAF  (90.44 to 91.10). At this bitrate differential, and assuming a $0.02/GB bandwidth cost, that saved about $0.0147 in bandwidth for every hour viewed. At that rate, each hour of encoded content must be watched about 186 times before the bandwidth savings cover the extra encoding cost. After a million hours of viewing bandwidth saved equaled $14,682. Figure 4. Breakeven between the baseline and optimized ladders in SLC Bitrate Explorer. The encoding cost is recovered in 186 hours of viewing, while overall ladder quality improved by .66 VMAF (at a 33% lower bandwidth). For an audience with more viewers on slower connections, the quality gain would be larger. If we switched the audience to the mobile preset, the VMAF differential jumps to 8.61. However, because this group retrieved much lower bitrate videos, the bandwidth savings dropped to $0.0023/hour and the breakeven increased to 1,169 hours. What’s it Mean As a codec researcher, you hate to spend weeks of testing only to reach the obvious answer; while there are multiple configuration options you can adjust for minor gain, the most effective adjustment you can make is to implement some form of per-title encoding. Two caveats. First, the benefit from per-title encoding depends upon the efficiency of your current encoding ladder. If the top rung is 5.5 Mbps to 6.5 Mbps or higher for 1080p, you should see substantial bitrate savings. If your current top rung is in the 4.5 Mbps range and below, you may achieve some bandwidth savings, but the primary benefit will be higher quality, particularly in the lower rungs. Second, clip complexity also matters. If you have a single ladder for all content, from soccer matches to the 6:00 news, a per title ladder should deliver substantial bandwidth savings for news and other easy-to-encode clips and and higher quality at similar or slightly higher bitrates for the hardest content. If you’re encoding primarily sports and other hard-to-encode content, bandwidth savings will be much less. Now that you know how the story ends, let’s return to the beginning and tell you what we did and why. How we tested We tested in two stages. The first stage encoded 13 test clips at 1080p and changed one parameter at a time to find the optimal setting for encoding time and quality.  The second stage used those results to build full encoding ladders. In that stage, we started with a fixed ladder using mostly x264 defaults, changing to the optimal configuration one option at a time, and finishing with a per-title ladder that used all the optimal settings. The 13 clips cover animation, sports, primetime drama and music, news and education, and screen content. All quality scores are VMAF, measured at 1080p. All encoding times are wall-clock time on a single workstation (an Intel Core i9-14900 with 32 logical cores). Stage One: Testing 1080 to Identify the Optimal Configuration Settings In this stage we encoded each clip using the different configuration options to identify the optimal settings. We encoded each clip at a unique bitrate that produced a VMAF score of ~93 using the x264 slow preset and otherwise mostly default FFmpeg configurations. This kept each clip at a realistic top-rung quality level instead of a fixed bitrate that would be too high for some clips and too low for others. The baseline configuration for the first series of tests was the veryslow preset, two-pass encoding, and a two-second GOP. These tests also used eight encoding threads, so their encoding times are relative to that configuration. As you’ll see, the baseline for the second series of tests used the medium preset. Why the difference?