Systems · · 2,346 words · 11 min read
What a quantisation format actually costs, derived from the source
Q4_K is not four bits per weight. Exact costs derived from ggml-common.h, joined to llama.cpp's measured size and perplexity tables, and the dominated formats named.
quantisation GGUF llama.cpp inference
Q4_K does not store four bits per weight. It stores 4.5, and the file you download stores 4.8944, and if you sized a GPU on the number in the name you are out by twenty per cent at the low end. This article derives the real figure for every GGUF block type directly from llama.cpp's own header, joins it to the project's two published measurement tables, and names the formats that are strictly worse than something else.
Everything here is computed or quoted, nothing estimated. The listing is code/gguf-bits-per-weight.py; it fetches the three upstream files at run time, so re-running it tomorrow reports tomorrow's upstream rather than this article's snapshot. The joined table is datasets/gguf-bits-per-weight.csv. No model was run and no GPU was touched: the cost arithmetic is exact, and every quality number is llama.cpp's measurement with its hardware and revision attached.
Where the cost actually comes from
A GGUF quantised tensor is not a flat array of n-bit integers. It is an array of fixed-size blocks, and each block carries its quantised weights plus the metadata needed to reconstruct them: a scale, sometimes a minimum, sometimes a second tier of scales for sub-blocks. The block layout is pinned in ggml/src/ggml-common.h by a static_assert on each struct's size, which means the exact byte cost is a compile-time fact you can read off rather than estimate.
block_q4_K asserts its size as:
static_assert(sizeof(block_q4_K) == 2*sizeof(ggml_half) + K_SCALE_SIZE + QK_K/2,
"wrong q4_K block size/padding");
With QK_K 256, K_SCALE_SIZE 12 and ggml_half two bytes, that is 4 + 12 + 128 = 144 bytes for 256 weights, which is 4.5 bits per weight. The 128 bytes are the four-bit weights. The other 16 are a super-block scale and minimum in half precision, plus twelve bytes of six-bit packed scales and minimums for the eight sub-blocks. Half a bit per weight of bookkeeping, or an eighth of the total.
The script evaluates every such assertion, so the table below is not typed in by hand:
| Format | Block bytes | Weights | Bits per weight |
|---|---|---|---|
| iq1_s | 50 | 256 | 1.5625 |
| tq1_0 | 54 | 256 | 1.6875 |
| iq1_m | 56 | 256 | 1.7500 |
| iq2_xxs | 66 | 256 | 2.0625 |
| tq2_0 | 66 | 256 | 2.0625 |
| iq2_xs | 74 | 256 | 2.3125 |
| iq2_s | 82 | 256 | 2.5625 |
| q2_K | 84 | 256 | 2.6250 |
| iq3_xxs | 98 | 256 | 3.0625 |
| iq3_s | 110 | 256 | 3.4375 |
| q3_K | 110 | 256 | 3.4375 |
| iq4_xs | 136 | 256 | 4.2500 |
| mxfp4 | 17 | 32 | 4.2500 |
| iq4_nl | 18 | 32 | 4.5000 |
| q4_0 | 18 | 32 | 4.5000 |
| q4_K | 144 | 256 | 4.5000 |
| nvfp4 | 36 | 64 | 4.5000 |
| q4_1 | 20 | 32 | 5.0000 |
| q5_0 | 22 | 32 | 5.5000 |
| q5_K | 176 | 256 | 5.5000 |
| q5_1 | 24 | 32 | 6.0000 |
| q6_K | 210 | 256 | 6.5625 |
| q8_0 | 34 | 32 | 8.5000 |
Three things fall out of this table immediately.
The small block types pay the most for their metadata. q4_0 packs 32 weights at four bits and adds one half-precision scale: 16 bytes of data, two bytes of overhead, 4.5 bits per weight. q4_K reaches the same 4.5 bits while carrying a two-level scale hierarchy, because it spreads that hierarchy over 256 weights instead of 32. Same cost, more structure, and the quality measurements below show what the structure buys.
The legacy formats got more expensive as they got more careful. q4_1 adds a minimum alongside the scale, which takes it from 4.5 to 5.0 bits. q5_1 does the same on top of five-bit weights and lands at 6.0, which is more than q5_K at 5.5 and within half a bit of q6_K.
Nothing is named after its cost. The only format in the table whose name matches its bits per weight is none of them. q8_0 is 8.5, q6_K is 6.5625, iq1_s is 1.5625.
Four block types in the header are not weight quantisation targets and are excluded above: q8_1 at 9.0 and q8_K at 9.125 are activation-side formats used during matrix multiplication, and q1_0 at 1.125 and q2_0 at 2.25 are not offered by the quantise tool.
The second overhead: the file is bigger than the format
The per-block figure is the cost of a tensor stored in that format. A model file is not one tensor in one format. The quantise tool mixes types across tensors, and it keeps the token embedding and output projection at higher precision than the rest, because those two are where low-bit quantisation does the most damage.
llama.cpp publishes the whole-model figure for Llama-3.1-8B. Subtracting gives the mixture-and-embedding overhead:
| Format | Block type | Per-block bits | Whole-model bits | Gap | Gap % | File GiB |
|---|---|---|---|---|---|---|
| IQ1_S | iq1_s | 1.5625 | 2.0042 | +0.4417 | +28.3% | 1.87 |
| IQ1_M | iq1_m | 1.7500 | 2.1460 | +0.3960 | +22.6% | 2.01 |
| IQ2_XXS | iq2_xxs | 2.0625 | 2.3824 | +0.3199 | +15.5% | 2.23 |
| IQ2_XS | iq2_xs | 2.3125 | 2.5882 | +0.2757 | +11.9% | 2.42 |
| IQ2_S | iq2_s | 2.5625 | 2.7403 | +0.1778 | +6.9% | 2.56 |
| Q2_K | q2_K | 2.6250 | 3.1593 | +0.5343 | +20.4% | 2.95 |
| IQ3_XXS | iq3_xxs | 3.0625 | 3.2548 | +0.1923 | +6.3% | 3.04 |
| IQ3_S | iq3_s | 3.4375 | 3.6606 | +0.2231 | +6.5% | 3.42 |
| IQ4_XS | iq4_xs | 4.2500 | 4.4597 | +0.2097 | +4.9% | 4.17 |
| IQ4_NL | iq4_nl | 4.5000 | 4.6818 | +0.1818 | +4.0% | 4.38 |
| Q6_K | q6_K | 6.5625 | 6.5633 | +0.0008 | +0.0% | 6.14 |
| Q8_0 | q8_0 | 8.5000 | 8.5008 | +0.0008 | +0.0% | 7.95 |
The two rows at the bottom are the ones that prove the method. Q6_K and Q8_0 are applied nearly uniformly, and for both of them the figure derived from a C header matches the figure measured on a real file to four decimal places, with eight ten-thousandths of a bit left over for GGUF metadata. The derivation is not an approximation of the measurement; it is the measurement minus the mixture.
Everything above those two rows is the mixture. The pattern is the useful part: the overhead grows as the quantisation gets more aggressive, from four per cent at IQ4_NL to twenty-eight per cent at IQ1_S. The embedding and output tensors are a roughly fixed number of bits, so as the rest of the model shrinks they become a larger share of the file. Choosing a one-bit format does not give you a one-bit model.
Names carrying a _S, _M or _L suffix, and IQ3_XS, IQ2_M and IQ3_M, are mixtures the tool assembles from more than one block type. They have no single per-block cost, so the script reports them as mixtures rather than inventing a comparison.
What that costs in gigabytes
The F16 row lets the weight count be recovered rather than assumed: 14.96 GiB at 16.0005 bits per weight is 8.0313 billion weights, which is the model on the label. Multiplying that count by the per-block cost gives the file size a naive reader would predict:
| Format | Naive GiB | Actual GiB | Extra |
|---|---|---|---|
| IQ1_S | 1.46 | 1.87 | +0.41 |
| IQ2_XXS | 1.93 | 2.23 | +0.30 |
| Q2_K | 2.45 | 2.95 | +0.50 |
| IQ3_XXS | 2.86 | 3.04 | +0.18 |
| IQ4_XS | 3.97 | 4.17 | +0.20 |
| Q6_K | 6.14 | 6.14 | +0.00 |
| Q8_0 | 7.95 | 7.95 | +0.00 |
Half a gigabyte on Q2_K is the difference between fitting and not fitting on a small card, and it is entirely predictable once you stop reading the name as a cost.
The quality table, with its provenance
llama.cpp publishes a measured scoreboard for Llama 3 8B at revision f364eb6f, on a CUDA backend with an AMD Epyc 7742 and a 1x NVIDIA RTX 4090, sorted by KL divergence against FP16. Forty-six rows, covering the same type at several importance-matrix sizes.
Read the first row before any of the others. The f16 entry shows a perplexity delta of 0.001524, not zero, and the README says why: the FP16 logits are cached between runs as scaled 16-bit unsigned integers, so that row measures only the downcast. That is the noise floor of the whole table. Q8_0's delta of 0.002650 is less than twice it. Whatever Q8_0 does to this model, this experiment cannot resolve it.
| Format | imatrix | GiB | PPL | ΔPPL | KLD |
|---|---|---|---|---|---|
| f16 | None | 14.97 | 6.233160 | 0.001524 | 0.000551 |
| q8_0 | None | 7.96 | 6.234284 | 0.002650 | 0.001355 |
| q6_K | None | 6.14 | 6.253382 | 0.021748 | 0.005452 |
| q5_K_M | None | 5.33 | 6.288607 | 0.056974 | 0.010762 |
| q4_K_M | WT 10m | 4.58 | 6.382937 | 0.151303 | 0.028152 |
| q4_K_M | None | 4.58 | 6.407115 | 0.175482 | 0.031273 |
| q4_0 | None | 4.34 | 6.700147 | 0.468514 | 0.071940 |
| q3_K_M | WT 10m | 3.74 | 6.734255 | 0.502622 | 0.084358 |
| q2_K | WT 10m | 2.96 | 8.647825 | 2.416191 | 0.332223 |
| q2_K | None | 2.96 | 9.751568 | 3.519934 | 0.445132 |
| iq2_XXS | WT 10m | 2.24 | 14.091782 | 7.860148 | 0.812022 |
| iq1_S | WT 1m | 1.88 | 58.097760 | 51.866126 | 2.211278 |
The bottom of that table is not a gentle degradation. Between 2.24 GiB and 1.88 GiB, perplexity goes from fourteen to fifty-eight. Whatever IQ1_S is useful for, it is not a smaller version of the same model.
What an importance matrix is worth
The scoreboard measures several types both with and without an importance matrix at identical file size, which makes the imatrix a free variable: same bytes, different calibration.
| Format | PPL without | PPL with WT 10m | Saved | Share of the gap closed |
|---|---|---|---|---|
| q2_K | 9.7516 | 8.6478 | 1.1037 | 31.4% |
| q3_K_S | 7.8638 | 7.6029 | 0.2609 | 16.0% |
| q3_K_M | 6.8885 | 6.7343 | 0.1542 | 23.5% |
| q3_K_L | 6.7879 | 6.6712 | 0.1167 | 21.0% |
| q4_K_S | 6.5005 | 6.4097 | 0.0908 | 33.8% |
| q4_K_M | 6.4071 | 6.3829 | 0.0242 | 13.8% |
Only three types are measured at more than one calibration size, so every row above uses the same one, WT 10m, rather than each type's best.
The absolute saving tracks how much damage there was to undo: 1.1037 of perplexity at q2_K, 0.0242 at q4_K_M. The share column says it differently, and it is the more useful reading: an imatrix recovers between 13.8 and 33.8 per cent of what quantisation cost, everywhere it was measured. It is not a fix. It is a third of a fix, for free, and it matters most exactly where you are most tempted to skip it.
More calibration tokens is not monotonically better, which is why picking each type's best row would have been cherry-picking. q2_K is measured at all five sizes, and ordered by perplexity they run WT 100k at 8.641993, WT 10m at 8.647825, WT 10k at 8.652290, WT 1m at 8.674365 and WT 1k at 8.682605: a span of 0.040612, four hundredths of a point, in an order that has nothing to do with token count. iq1_S is measured at the same five sizes and the spread is far larger and just as unordered, running from WT 1m at 58.0978 to WT 10k at 63.2213, with WT 10m fourth of five at 60.6946. Ten times the calibration data can make the result worse.
The formats nothing should use
Joining size against perplexity gives a Pareto frontier: the rows for which no other row is both smaller and more accurate. Twenty-one of the forty-six rows are on it. Every legacy format is off it.
| Dominated | GiB | PPL | Beaten by | GiB | PPL |
|---|---|---|---|---|---|
| q4_0 | 4.34 | 6.7001 | iq4_XS (WT 10m) | 4.14 | 6.4597 |
| q4_1 | 4.78 | 6.6827 | q4_K_M (WT 10m) | 4.58 | 6.3829 |
| q5_0 | 5.21 | 6.3632 | q5_K_S (None) | 5.21 | 6.3366 |
| q5_1 | 5.65 | 6.3379 | q5_K_M (None) | 5.33 | 6.2886 |
| q3_K_S (WT 10m) | 3.41 | 7.6029 | iq3_XS (WT 10m) | 3.28 | 7.1630 |
| q2_K_S (WT 10m) | 2.96 | 9.3238 | iq2_M (WT 10m) | 2.74 | 8.6008 |
| q2_K (WT 10m) | 2.96 | 8.6478 | iq2_M (WT 10m) | 2.74 | 8.6008 |
q4_1 is the clearest case: it is 0.20 GiB larger than q4_K_M and 0.30 perplexity worse. It costs 5.0 bits per weight against q4_K's 4.5 and spends the extra half bit on a per-32-weight minimum that a two-level scale hierarchy over 256 weights does better. The four legacy formats are kept for compatibility, and on this model on this measurement there is no size at which any of them is the right choice.
Below about 3.5 GiB the frontier is entirely I-quants with an importance matrix; from 3.74 to 4.58 GiB it alternates between K and I; above 5 GiB it is K-quants with no imatrix, though the scoreboard does not measure imatrix versions up there, so that last stretch is a gap in the data rather than a finding.
Where INT8, INT4 and FP8 sit
GGUF is one ecosystem. The formats named in most quantisation discussions belong to others, and the distinction that matters is not the bit count but what the bits mean.
Integer formats store a signed or unsigned integer per weight and a floating-point scale per group. INT8 and INT4 are the family GGUF's q8_0 and q4_0 belong to. The design space is entirely in the grouping: per-tensor, per-channel, per-block of 32, per-block of 256 with a second tier. That choice is what separates q4_0 from q4_K at the same nominal width.
Floating-point formats store an exponent and a mantissa per weight. FP8 has two standard shapes, E4M3 with four exponent and three mantissa bits and E5M2 with five and two, trading range against precision. They are hardware formats first: their appeal is that recent accelerators multiply them natively, so the saving is arithmetic throughput as well as memory.
ggml now carries two four-bit floating point block types, and their costs are in the first table: mxfp4 at 17 bytes per 32 weights is 4.25 bits per weight, with a single shared byte-sized exponent; nvfp4 at 36 bytes per 64 weights is 4.5, spending four bytes on UE4M3 scales, one per 16-weight sub-block. The same pattern as the integer formats, then: the finer the scale grouping, the more metadata, and the arithmetic is identical.
What the tables above cannot tell you about these is quality, because the scoreboard does not measure them. That is a real limit and not a small one.
What this does not establish
- One model, one measurement. Everything quality-related is Llama 3 8B at one revision on one machine. Quantisation damage is model-dependent, and a 70B model or a mixture-of-experts model may rank these formats differently.
- Perplexity is not capability. The scoreboard also publishes KLD and token-probability percentiles precisely because perplexity alone hides where the damage lands. A format can hold perplexity and still change which token comes out on top: q4_K_M agrees with FP16's top token 91.901% of the time, and q2_K only 71.138%.
- No throughput trade-off is modelled here. The quantize README's own measurements show text generation at 79.73 tokens per second for IQ1_S against 50.93 for Q8_0 and 29.17 for F16, so smaller is faster to generate with, while prompt processing runs the other way, fastest at Q8_0 and F16. A frontier drawn on size and perplexity ignores that axis.
- The mixture is not reverse-engineered. This article measures the gap between per-block and whole-model cost; it does not claim to know which tensor got which type. That is in the tool's source, and it is a different article.
- Upstream moves. These figures were read on 22 September 2026. Re-run the listing rather than trusting the tables here.
Code and data
- gguf-bits-per-weight.py — the complete listing used in this article.
- gguf-bits-per-weight.csv — the data behind the numbers here.
Sources
- ggml-org/llama.cpp, ggml/src/ggml-common.h on master, read 2026-09-22 (the block struct definitions and their static_assert size expressions for 27 block types; QK_K 256, QK4_0 and QK5_0 and QK8_0 32, K_SCALE_SIZE 12, IQ3S_N_SCALE QK_K/64, QK_MXFP4 32, QK_NVFP4 64 with QK_NVFP4_SUB 16)
- ggml-org/llama.cpp, tools/quantize/README.md on master, read 2026-09-22 (whole-model bits/weight, file size in GiB and prompt-processing and text-generation throughput for 25 quantisation types on meta-llama/Llama-3.1-8B; the Llama 3.1 memory table giving 8B at 32.1 GB original and 4.9 GB at Q4_K_M; the description of quality loss as measured in perplexity and Kullback-Leibler divergence and minimised by a suitable imatrix file)
- ggml-org/llama.cpp, tools/perplexity/README.md on master, read 2026-09-22 (the LLaMA 3 8b Scoreboard: revision f364eb6f, CUDA backend, AMD Epyc 7742, 1x NVIDIA RTX 4090; 46 rows of quantisation type, imatrix, model size, perplexity, delta-perplexity, KLD, mean and RMS change in correct-token probability, sorted by KLD relative to FP16; the note that the stored FP16 logits are downcast to 16-bit unsigned integers so the f16 row measures only that downcast; the Wikitext importance matrices at varying token counts)
- Meta, Llama-3.1-8B model card, the model the quantize README measures, read 2026-09-22
- the importance matrix file the perplexity README links for its WT rows, read 2026-09-22