Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,637 words · 1 segments analyzed
30 Sep 2026 In early September, NVIDIA released an update to their vk_lod_clusters open-source sample, that - among other things - featured a new impressive Zorah scene as a glTF file. The demo showcased the application of new driver-exposed raytracing features, specifically clustered raytracing, in combination with Nanite-like clustered LOD pipeline - allowing to stream and display a very highly detailed scene with full pathtracing. Naturally, this piqued my curiosity - and led me to spend some time to improve support for hierarchical clustered LOD in meshoptimizer. … Wait, this sounds familiar. Indeed, the first version of a standalone Zorah scene, as well as the foundational driver components for clustered RT, actually released last year! I’ve written about the journey to get that scene to process quickly in Billions of triangles in minutes. However, the vk_lod_clusters sample has made significant progress since last year, the scene got updated with full attribute set and texture data, the renderer is using path-tracing and looks beautiful, and the processing mostly uses meshoptimizer’s clusterlod.h. Because I did spend some time to improve aspects of this in meshoptimizer, I figured it would make sense to do a writeup, focusing a bit less on processing time now and more on other aspects. Reading the previous article is not required but would probably help understand this one better. glTF compression The first thing that you would notice is that the new scene is on a similar scale (1.6B triangles) and comes with a hefty 70 GB download - which should be expected, as it now has a lot of textures! But if you look at just the glTF scene data, the 2025 version was ~36 GB, and the 2026 version is ~32 GB. The overall amount of geometry did not change very much, so this should be surprising as the old scene had just the position data on most of the meshes and should have been noticeably smaller. Indeed, NVIDIA now provides a new version of the original, attribute-less and texture-less, scene, which is just 9 GB. This reduction is due to the new glTF scenes using meshoptimizer compression extensions, namely EXT_meshopt_compression. This extension exposes support for vertex and index compression; note that vertex data is compressed using an older, “v0” vertex format. The vertex codec has been upgraded with v1 since, which provides further gains in compression ratios and decompression performance, and is supported in a new extension, KHR_meshopt_compression. Since KHR variant was only standardized in early 2026, it makes sense that the assets use the previous version - but we’ll look at both for completeness. The extensions expose the vertex and index compression but leave a lot of room for the processing tools to organize the data or choose the compression characteristics at will; unlike some black-box “one mesh at a time” formats, meshoptimizer formats are defined on arrays of bytes representing vertex or index data. For vertex data, the compressor itself is lossless and will perfectly reproduce the original byte sequence; however, you can also use additional layers underneath it that would make it lossy, either by using pre-quantized data when glTF geometry accessors support it, or by using “filters” which compress the data with a little bit of lost precision which allows to gain extra compression ratios. The way the meshes reference the arrays is also up to the tools that prepare the data: some tools may decide to use separate buffers for individual meshes, and some may choose to combine them. Typically, for the last-mile delivery scenarios - e.g. producing a glTF file that will be rendered directly - you’d want to use the lossy compression (see the appendix at the end) and tune the precision to the level that is required for display; that would give you the smallest output file. But in case of the Zorah scene specifically, the scene we’re dealing with is merely an intermediate file - the actual clustered LOD processing will significantly change the precision characteristics in a way that can be tuned later. So it makes sense that the Zorah glTF scene only uses the lossless option. Let’s look at how effective it is. Because the files are using v0 (EXT) compression, I’ve run stats on the scene as well as what switching to v1 would achieve, and got the following numbers; “v0” is what is stored on disk in the downloaded file at the moment, “raw” is what the scene would look like once decompressed1. Category Raw v0 v1 (recompressed) Positions 14.3 GB 9.1 GB (63.5%) 8.9 GB (62.1%) Attributes 30.1 GB 21.6 GB (71.8%) 20.5 GB (68.2%) Indices 18.9 GB 1.9 GB (10.1%) 1.9 GB (10.1%) Total 63.3 GB 32.6 GB (51.5%) 31.3 GB (49.5%) So overall on the entire file we get about 2x compression for geometry; indices compress quite well, and the positions and attributes compress a bit less well2. Using v1 vertex format would gain a few extra percentage points of compression ratio (and would decode faster). Now, you may ask - do we even need to compress this file, given that it’s not the final delivery format? The file we download is compressed with a general-purpose lossless compressor, so why bother? There’s two things to say about this. First, from the processing time point of view, this allows us to reduce the cost of reading the scene from disk compared to either a raw or a hypothetical representation where we use a compressor like Zstandard to keep the files compressed on disk. Decoding vertex or index data runs at multiple gigabytes per second of output data on a single core; as I’ve written in the previous article, the processing is massively multi-core, so our aggregate decompression throughput on 16 cores is “a lot”. The way this is integrated into the processing is that the vertex/index streams for each mesh are compressed independently and decompressed before processing each mesh; as a result, decompression takes under 1% of total processing time, runs fully in parallel, and reduces the size we need to read from disk 2x. Second, this might result in a smaller file compared to using general-purpose compressors. If we instead used Zstandard (at level 6 in this case), we’d get a larger file, that would also decompress more slowly - either decoder is noticeably faster than Zstandard. What’s more, if you wanted to further reduce the file size, you could additionally compress the encoded bytes with Zstandard again - for an even smaller size! You would have to decompress the file twice - first decompress the encoded bytes using Zstandard, and then decompress the original bytes using meshopt_decode*. However, because Zstandard would need to decompress less data, generally the resulting decoding would still be faster than just using Zstandard alone3. Raw v1 Raw + Zstd v1 + Zstd 63.3 GB 31.3 GB (49.5%) 38.0 GB (60.1%) 27.4 GB (43.3%) This seems like an unusual property; normally, you should not be able to compress compressed representations! But this is an explicit design point. Instead of implementing the entire featureset of a modern lossless compressor, meshoptimizer codecs focus on parts that lossless compressors, especially faster ones, often don’t model, but try to keep the output stream in a format that is amenable to both fast decoding and further post-compression. This is nice from the versatility point of view - for example, maybe you want to use LZMA as a post-compression layer, which is impractically slow to decompress at game load time, but you could decompress it at install time instead, and just leave the raw meshopt-compressed files on disk. Ultimately the best balance depends on the data characteristics, and, importantly, the speed of the underlying media. Part of the motivation for the meshoptimizer codecs originally was just that SSDs were becoming unreasonably fast. If your SSD can read many gigabytes per second, you might need many cores to actually save time when streaming assets from disk when using slower compression algorithms; and some compression algorithms simply can’t get there at all as they are too slow. Instead of aiming for a maximum possible compression ratio, meshoptimizer codecs are designed for extremely fast decompression, and are flexible enough that the resulting size can be dialed further with layered compression. It probably feels like all of this is not directly related to clustered LOD… and indeed it’s not :) it was just an excuse to write about some aspects of mesh compression that happened to be used in Zorah! Let’s talk about the more relevant parts of the processing and runtime. Processing improvements The rest of the processing proceeds largely as I’ve written about it last year. A slightly modified version of clusterlod.h is used in vk_lod_clusters and the pipeline itself is as before - split the input mesh into clusters, partition clusters into groups, simplify each group preserving boundary, split the new group into more clusters, rinse and repeat. For runtime LOD selection, clusterlod.h now provides helper utilities to build a BVH over the DAG (clodBuildHierarchy), which can be used to quickly select the clusters that are at the correct level of fidelity, as described in the original Nanite presentation. Because the scene now has attributes, they need to be communicated to the simplifier so that it can compute the error that fully reflects the appearance change. Each attribute has a scalar weight, and is weighted according to its impact over the surface that it’s changing across. This is using attribute-aware simplification (meshopt_simplifyWithAttributes) which has been available for a few years, and the initial version of clusterlod.h exposed it too; one new improvement to the metric is noted below, but otherwise not a lot has changed with attributes themselves. When doing cluster partitioning, clusterlod.h uses spatial information by default to improve partitioning and also has an additional option, partition_sort, which sorts partition by Morton order, which makes the entire