Pangram verdict · v3.3
We believe that this text is a mix of AI, AI-assisted, and human-written content.
AI likelihood · overall
MixedArticle text · 1,054 words · 5 segments analyzed
We were a little impatient. Qwen 3.8 27B was released on August 14, 2026. Four days later, on August 18, we published our first set of GGUFs. We called them ShapeLearn-Lite for a reason: they were produced using a much smaller optimization budget, fewer checks, and much less waiting. Now the full ShapeLearn models are done, and we have benchmarked them alongside the original Lite set and competing quants. The good news: ShapeLearn-Lite held up pretty well. We will come back to that later in “ShapeLearn-Lite, in retrospect”. The better news: the full ShapeLearn models are even better. Quick start with llama.cpp The MTP draft head is bundled in every GGUF. DFlash2 uses a separate 1.1 GB draft model. Both commands use GPU-5. Swap the tag for any other model in the release. MTP Embedded draft. Works with image inputs.
llama-server \ -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw \ --mmproj-auto \ --spec-type draft-mtp --spec-draft-n-max 3 DFlash2 External draft. Fastest option, text only. llama-server \ -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw \ -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \ --spec-type draft-dflash --spec-draft-n-max 7 \ --no-mmproj DFlash2 needs llama.cpp b10658 or newer.
Ready-to-run commands for every model, with the recommended sampling settings, are in the run tool and on the model card. TL;DR Full ShapeLearn moves the measured quality-speed frontier beyond Lite. All five models in the new release sit on the frontier in each of our six GPU comparisons. GPU-5 is our default recommendation wherever it fits, reaching 99.63% of BF16’s aggregate benchmark score. If it does not fit with the context you need, GPU-4 is still very competitive: it reaches 98.72% of BF16 at a much smaller size (11.0 GB instead of 13.1 GB), and it is faster.
ShapeLearn-Lite also performed better than its KLD ranking suggested: three of its six models sit on the frontier in the Lite-versus-Unsloth Dynamic v3 comparison. Speculative Decoding with MTP or DFlash2 increases throughput across every ShapeLearn model and GPU tested. DFlash2 is usually faster but requires more memory and does not support image inputs with llama.cpp. Choose DFlash2 for maximum text-only throughput when memory allows, and MTP when VRAM or multimodal support matters more. Full ShapeLearn moves the frontier We are releasing the full ShapeLearn run for Qwen 3.8 27B. Within this release, larger models yield higher aggregate scores, while smaller models deliver higher throughput. That ordering holds across all six GPUs tested. Because this is a dense model and memory transfers are the bottleneck, lower BPW translates more directly into higher TPS than it does for MoEs. The per-GPU comparisons also include AtomicChat, Bartowski, ISTA-DASLab, and Unsloth Dynamic v3. Bartowski’s latest models were released after our testing and are not included. Full ShapeLearn is labelled ByteShape in the figures. All five ShapeLearn models remain on the measured frontier, with GPU-5 achieving the highest aggregate score among the plotted quants. Other teams also contribute competitive points. Notably, ISTA-DASLab’s excellent model (the yellow “d” on the graph below) also sits on the frontier. By “frontier,” we mean that no other plotted model is both faster and more accurate. 96 GB: RTX Pro 6000 RTX PRO 6000, the GPU with the most memory, lets us show the full range of models tested. RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend #ModelAccTPSBPW ByteShape GPU-1IQ2_XXS-2.56bpw0.9304116.112.56 GPU-2IQ3_XXS-2.88bpw0.9656108.012.88 GPU-3IQ3_XS-3.01bpw0.9726105.443.01 GPU-4IQ3_S-3.23bpw0.9872101.113.23 GPU-5IQ4_XS-3.84bpw0.996390.423.84 Unsloth AUD-IQ2_S0.8633114.812.49 BUD-Q2_K_XL0.9572106.952.82 CUD-IQ3_XXS0.9359100.513.14 DUD-IQ3_S0.955595.533.47 EUD-Q3_K_XL0.976090.883.80 FUD-IQ4_XS0.992086.734.13 GUD-Q4_K_S0.987782.214.46 HUD-Q4_K_M0.970378.584.79 IUD-Q4_K_XL0.987174.755.12 JUD-Q5_K_S0.987871.005.44 KUD-Q5_K_M0.989767.715.77 LUD-Q5_K_XL0.990565.526.10 ISTA-DASLab aGSQ-RCO-IQ2_XS0.8647112.772.50 bGSQ-RCO-IQ2_S0.9364108.222.75 cGSQ-RCO-IQ3_XXS0.9438103.393.00 dGSQ-RCO-IQ3_S0.994394.773.50 Bartowski aIQ2_XXS0.7986112.212.72 bIQ2_S0.9296106.332.99 cQ2_K0.961696.083.45 dIQ3_XXS0.959492.433.68 eIQ3_XS0.958288.363.89 fIQ3_M0.966786.174.06 AtomicChat aAD-IQ2_XXS0.8385116.282.58 bAD-IQ2_XS0.9296108.672.85 cAD-IQ2_S0.9061100.283.22 dAD-IQ3_XXS0.964495.293.50 eAD-IQ3_S0.973088.524.04 GPU-5 is our default wherever you can fit it: it reaches 90.4 tok/s at 99.63% of the BF16 baseline. 32 GB: RTX 5090 The RTX 5090 tells a similar story, leading to the same recommendations. RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend #ModelAccTPSBPW ByteShape GPU-1IQ2_XXS-2.56bpw0.9304119.132.56 GPU-2IQ3_XXS-2.88bpw0.9656110.782.88 GPU-3IQ3_XS-3.01bpw0.9726108.083.01 GPU-4IQ3_S-3.23bpw0.9872103.573.23 GPU-5IQ4_XS-3.84bpw0.996393.663.84 Unsloth AUD-IQ2_S0.8633115.342.49 BUD-Q2_K_XL0.9572108.122.82 CUD-IQ3_XXS0.9359102.873.14 DUD-IQ3_S0.955597.873.47 EUD-Q3_K_XL0.976093.463.80 FUD-IQ4_XS0.992089.694.13 GUD-Q4_K_S0.987785.284.46 HUD-Q4_K_M0.970381.634.79 IUD-Q4_K_XL0.987177.515.12 JUD-Q5_K_S0.987873.605.44 KUD-Q5_K_M0.989769.875.77 LUD-Q5_K_XL0.990567.596.10 ISTA-DASLab aGSQ-RCO-IQ2_XS0.8647114.112.50 bGSQ-RCO-IQ2_S0.9364110.192.75 cGSQ-RCO-IQ3_XXS0.9438105.673.00 dGSQ-RCO-IQ3_S0.994395.493.50 Bartowski aIQ2_XXS0.7986115.772.72 bIQ2_S0.9296109.392.99 cQ2_K0.961699.293.45 dIQ3_XXS0.959495.833.68 eIQ3_XS0.958291.143.89 fIQ3_M0.966789.084.06 AtomicChat aAD-IQ2_XXS0.8385119.532.58 bAD-IQ2_XS0.9296112.392.85 cAD-IQ2_S0.9061103.463.22 dAD-IQ3_XXS0.964498.223.50 eAD-IQ3_S0.973091.724.04 Once again GPU-5 is our default choice, reaching 93.7 tok/s. Choose GPU-4 for slightly more context length or slightly better TPS. 24 GB: RTX 4090 and RTX 3090 Both 24 GB cards fit all five ShapeLearn models. We plot them separately because their throughput differs, but the ordering is the same on both. RTX 4090 The RTX 4090 keeps the same pattern: GPU-5 is the default, reaching 59.2 tok/s. RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend #ModelAccTPSBPW ByteShape GPU-1IQ2_XXS-2.56bpw0.930478.672.56 GPU-2IQ3_XXS-2.88bpw0.965672.892.88 GPU-3IQ3_XS-3.01bpw0.972671.123.01 GPU-4IQ3_S-3.23bpw0.987267.773.23 GPU-5IQ4_XS-3.84bpw0.996359.163.84 Unsloth AUD-IQ2_S0.863378.232.49 BUD-Q2_K_XL0.957272.792.82 CUD-IQ3_XXS0.935967.243.14 DUD-IQ3_S0.955563.023.47 EUD-Q3_K_XL0.976059.133.80 FUD-IQ4_XS0.992055.624.13 GUD-Q4_K_S0.987752.444.46 HUD-Q4_K_M0.970349.774.79 IUD-Q4_K_XL0.987147.085.12 JUD-Q5_K_S0.987844.805.44 KUD-Q5_K_M0.989742.455.77 LUD-Q5_K_XL0.990540.846.10 ISTA-DASLab aGSQ-RCO-IQ2_XS0.864777.802.50 bGSQ-RCO-IQ2_S0.936473.342.75 cGSQ-RCO-IQ3_XXS0.943869.263.00 dGSQ-RCO-IQ3_S0.994362.413.50 Bartowski aIQ2_XXS0.798675.282.72 bIQ2_S0.929671.272.99 cQ2_K0.961663.743.45 dIQ3_XXS0.959461.033.68 eIQ3_XS0.958258.203.89 fIQ3_M0.966756.164.06 AtomicChat aAD-IQ2_XXS0.838579.512.58 bAD-IQ2_XS0.929674.192.85 cAD-IQ2_S0.906167.403.22 dAD-IQ3_XXS0.964463.603.50 eAD-IQ3_S0.973057.384.04 RTX 3090 Older, but still fast in these measurements. RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend #ModelAccTPSBPW ByteShape GPU-1IQ2_XXS-2.56bpw0.930453.032.56 GPU-2IQ3_XXS-2.88bpw0.965651.202.88 GPU-3IQ3_XS-3.01bpw0.972650.753.01 GPU-4IQ3_S-3.23bpw0.987249.493.23 GPU-5IQ4_XS-3.84bpw0.996345.793.84 Unsloth AUD-IQ2_S0.863352.822.49 BUD-Q2_K_XL0.957250.372.82 CUD-IQ3_XXS0.935948.163.14 DUD-IQ3_S0.955546.393.47 EUD-Q3_K_XL0.976046.283.80 FUD-IQ4_XS0.992046.124.13 GUD-Q4_K_S0.987744.494.46 HUD-Q4_K_M0.970343.304.79 IUD-Q4_K_XL0.987141.375.12 JUD-Q5_K_S0.987839.375.44 KUD-Q5_K_M0.989737.385.77 LUD-Q5_K_XL0.990536.166.10 ISTA-DASLab aGSQ-RCO-IQ2_XS0.864751.562.50 bGSQ-RCO-IQ2_S0.936450.072.75 cGSQ-RCO-IQ3_XXS0.943848.663.00 dGSQ-RCO-IQ3_S0.994347.193.50 Bartowski aIQ2_XXS0.798653.632.72 bIQ2_S0.929651.212.99 cQ2_K0.961645.343.45 dIQ3_XXS0.959447.463.68 eIQ3_XS0.958244.683.89 fIQ3_M0.966743.424.06 AtomicChat aAD-IQ2_XXS0.838553.952.58 bAD-IQ2_XS0.929651.912.85 cAD-IQ2_S0.906148.243.22 dAD-IQ3_XXS0.964447.103.50 eAD-IQ3_S0.973047.224.04 GPU-4 reaches 49.5 tok/s, compared with 45.8 tok/s for GPU-5.
Moving to the larger model costs about 7.5% in throughput, while the aggregate score rises from 98.72% to 99.63% of BF16. That makes GPU-5 the default here as well. 16 GB: RTX 4080 and RTX 5060 Ti With a tighter VRAM budget, the pragmatic choice is to leave room for the context you need, not just the model weights. These plots contain fewer competing configurations, but all five ShapeLearn models are represented.