Matrix Cognition

Systems · · 2,346 words · 11 min read

What a quantisation format actually costs, derived from the source

Q4_K is not four bits per weight. Exact costs derived from ggml-common.h, joined to llama.cpp's measured size and perplexity tables, and the dominated formats named.

quantisation GGUF llama.cpp inference

Q4_K does not store four bits per weight. It stores 4.5, and the file you download stores 4.8944, and if you sized a GPU on the number in the name you are out by twenty per cent at the low end. This article derives the real figure for every GGUF block type directly from llama.cpp's own header, joins it to the project's two published measurement tables, and names the formats that are strictly worse than something else.

Everything here is computed or quoted, nothing estimated. The listing is code/gguf-bits-per-weight.py; it fetches the three upstream files at run time, so re-running it tomorrow reports tomorrow's upstream rather than this article's snapshot. The joined table is datasets/gguf-bits-per-weight.csv. No model was run and no GPU was touched: the cost arithmetic is exact, and every quality number is llama.cpp's measurement with its hardware and revision attached.

Where the cost actually comes from

A GGUF quantised tensor is not a flat array of n-bit integers. It is an array of fixed-size blocks, and each block carries its quantised weights plus the metadata needed to reconstruct them: a scale, sometimes a minimum, sometimes a second tier of scales for sub-blocks. The block layout is pinned in ggml/src/ggml-common.h by a static_assert on each struct's size, which means the exact byte cost is a compile-time fact you can read off rather than estimate.

block_q4_K asserts its size as:

static_assert(sizeof(block_q4_K) == 2*sizeof(ggml_half) + K_SCALE_SIZE + QK_K/2,
              "wrong q4_K block size/padding");

With QK_K 256, K_SCALE_SIZE 12 and ggml_half two bytes, that is 4 + 12 + 128 = 144 bytes for 256 weights, which is 4.5 bits per weight. The 128 bytes are the four-bit weights. The other 16 are a super-block scale and minimum in half precision, plus twelve bytes of six-bit packed scales and minimums for the eight sub-blocks. Half a bit per weight of bookkeeping, or an eighth of the total.

The script evaluates every such assertion, so the table below is not typed in by hand:

Format Block bytes Weights Bits per weight
iq1_s 50 256 1.5625
tq1_0 54 256 1.6875
iq1_m 56 256 1.7500
iq2_xxs 66 256 2.0625
tq2_0 66 256 2.0625
iq2_xs 74 256 2.3125
iq2_s 82 256 2.5625
q2_K 84 256 2.6250
iq3_xxs 98 256 3.0625
iq3_s 110 256 3.4375
q3_K 110 256 3.4375
iq4_xs 136 256 4.2500
mxfp4 17 32 4.2500
iq4_nl 18 32 4.5000
q4_0 18 32 4.5000
q4_K 144 256 4.5000
nvfp4 36 64 4.5000
q4_1 20 32 5.0000
q5_0 22 32 5.5000
q5_K 176 256 5.5000
q5_1 24 32 6.0000
q6_K 210 256 6.5625
q8_0 34 32 8.5000

Three things fall out of this table immediately.

The small block types pay the most for their metadata. q4_0 packs 32 weights at four bits and adds one half-precision scale: 16 bytes of data, two bytes of overhead, 4.5 bits per weight. q4_K reaches the same 4.5 bits while carrying a two-level scale hierarchy, because it spreads that hierarchy over 256 weights instead of 32. Same cost, more structure, and the quality measurements below show what the structure buys.

The legacy formats got more expensive as they got more careful. q4_1 adds a minimum alongside the scale, which takes it from 4.5 to 5.0 bits. q5_1 does the same on top of five-bit weights and lands at 6.0, which is more than q5_K at 5.5 and within half a bit of q6_K.

Nothing is named after its cost. The only format in the table whose name matches its bits per weight is none of them. q8_0 is 8.5, q6_K is 6.5625, iq1_s is 1.5625.

Four block types in the header are not weight quantisation targets and are excluded above: q8_1 at 9.0 and q8_K at 9.125 are activation-side formats used during matrix multiplication, and q1_0 at 1.125 and q2_0 at 2.25 are not offered by the quantise tool.

The second overhead: the file is bigger than the format

The per-block figure is the cost of a tensor stored in that format. A model file is not one tensor in one format. The quantise tool mixes types across tensors, and it keeps the token embedding and output projection at higher precision than the rest, because those two are where low-bit quantisation does the most damage.

llama.cpp publishes the whole-model figure for Llama-3.1-8B. Subtracting gives the mixture-and-embedding overhead:

Format Block type Per-block bits Whole-model bits Gap Gap % File GiB
IQ1_S iq1_s 1.5625 2.0042 +0.4417 +28.3% 1.87
IQ1_M iq1_m 1.7500 2.1460 +0.3960 +22.6% 2.01
IQ2_XXS iq2_xxs 2.0625 2.3824 +0.3199 +15.5% 2.23
IQ2_XS iq2_xs 2.3125 2.5882 +0.2757 +11.9% 2.42
IQ2_S iq2_s 2.5625 2.7403 +0.1778 +6.9% 2.56
Q2_K q2_K 2.6250 3.1593 +0.5343 +20.4% 2.95
IQ3_XXS iq3_xxs 3.0625 3.2548 +0.1923 +6.3% 3.04
IQ3_S iq3_s 3.4375 3.6606 +0.2231 +6.5% 3.42
IQ4_XS iq4_xs 4.2500 4.4597 +0.2097 +4.9% 4.17
IQ4_NL iq4_nl 4.5000 4.6818 +0.1818 +4.0% 4.38
Q6_K q6_K 6.5625 6.5633 +0.0008 +0.0% 6.14
Q8_0 q8_0 8.5000 8.5008 +0.0008 +0.0% 7.95

The two rows at the bottom are the ones that prove the method. Q6_K and Q8_0 are applied nearly uniformly, and for both of them the figure derived from a C header matches the figure measured on a real file to four decimal places, with eight ten-thousandths of a bit left over for GGUF metadata. The derivation is not an approximation of the measurement; it is the measurement minus the mixture.

Everything above those two rows is the mixture. The pattern is the useful part: the overhead grows as the quantisation gets more aggressive, from four per cent at IQ4_NL to twenty-eight per cent at IQ1_S. The embedding and output tensors are a roughly fixed number of bits, so as the rest of the model shrinks they become a larger share of the file. Choosing a one-bit format does not give you a one-bit model.

Names carrying a _S, _M or _L suffix, and IQ3_XS, IQ2_M and IQ3_M, are mixtures the tool assembles from more than one block type. They have no single per-block cost, so the script reports them as mixtures rather than inventing a comparison.

What that costs in gigabytes

The F16 row lets the weight count be recovered rather than assumed: 14.96 GiB at 16.0005 bits per weight is 8.0313 billion weights, which is the model on the label. Multiplying that count by the per-block cost gives the file size a naive reader would predict:

Format Naive GiB Actual GiB Extra
IQ1_S 1.46 1.87 +0.41
IQ2_XXS 1.93 2.23 +0.30
Q2_K 2.45 2.95 +0.50
IQ3_XXS 2.86 3.04 +0.18
IQ4_XS 3.97 4.17 +0.20
Q6_K 6.14 6.14 +0.00
Q8_0 7.95 7.95 +0.00

Half a gigabyte on Q2_K is the difference between fitting and not fitting on a small card, and it is entirely predictable once you stop reading the name as a cost.

The quality table, with its provenance

llama.cpp publishes a measured scoreboard for Llama 3 8B at revision f364eb6f, on a CUDA backend with an AMD Epyc 7742 and a 1x NVIDIA RTX 4090, sorted by KL divergence against FP16. Forty-six rows, covering the same type at several importance-matrix sizes.

Read the first row before any of the others. The f16 entry shows a perplexity delta of 0.001524, not zero, and the README says why: the FP16 logits are cached between runs as scaled 16-bit unsigned integers, so that row measures only the downcast. That is the noise floor of the whole table. Q8_0's delta of 0.002650 is less than twice it. Whatever Q8_0 does to this model, this experiment cannot resolve it.

Format imatrix GiB PPL ΔPPL KLD
f16 None 14.97 6.233160 0.001524 0.000551
q8_0 None 7.96 6.234284 0.002650 0.001355
q6_K None 6.14 6.253382 0.021748 0.005452
q5_K_M None 5.33 6.288607 0.056974 0.010762
q4_K_M WT 10m 4.58 6.382937 0.151303 0.028152
q4_K_M None 4.58 6.407115 0.175482 0.031273
q4_0 None 4.34 6.700147 0.468514 0.071940
q3_K_M WT 10m 3.74 6.734255 0.502622 0.084358
q2_K WT 10m 2.96 8.647825 2.416191 0.332223
q2_K None 2.96 9.751568 3.519934 0.445132
iq2_XXS WT 10m 2.24 14.091782 7.860148 0.812022
iq1_S WT 1m 1.88 58.097760 51.866126 2.211278

The bottom of that table is not a gentle degradation. Between 2.24 GiB and 1.88 GiB, perplexity goes from fourteen to fifty-eight. Whatever IQ1_S is useful for, it is not a smaller version of the same model.

What an importance matrix is worth

The scoreboard measures several types both with and without an importance matrix at identical file size, which makes the imatrix a free variable: same bytes, different calibration.

Format PPL without PPL with WT 10m Saved Share of the gap closed
q2_K 9.7516 8.6478 1.1037 31.4%
q3_K_S 7.8638 7.6029 0.2609 16.0%
q3_K_M 6.8885 6.7343 0.1542 23.5%
q3_K_L 6.7879 6.6712 0.1167 21.0%
q4_K_S 6.5005 6.4097 0.0908 33.8%
q4_K_M 6.4071 6.3829 0.0242 13.8%

Only three types are measured at more than one calibration size, so every row above uses the same one, WT 10m, rather than each type's best.

The absolute saving tracks how much damage there was to undo: 1.1037 of perplexity at q2_K, 0.0242 at q4_K_M. The share column says it differently, and it is the more useful reading: an imatrix recovers between 13.8 and 33.8 per cent of what quantisation cost, everywhere it was measured. It is not a fix. It is a third of a fix, for free, and it matters most exactly where you are most tempted to skip it.

More calibration tokens is not monotonically better, which is why picking each type's best row would have been cherry-picking. q2_K is measured at all five sizes, and ordered by perplexity they run WT 100k at 8.641993, WT 10m at 8.647825, WT 10k at 8.652290, WT 1m at 8.674365 and WT 1k at 8.682605: a span of 0.040612, four hundredths of a point, in an order that has nothing to do with token count. iq1_S is measured at the same five sizes and the spread is far larger and just as unordered, running from WT 1m at 58.0978 to WT 10k at 63.2213, with WT 10m fourth of five at 60.6946. Ten times the calibration data can make the result worse.

The formats nothing should use

Joining size against perplexity gives a Pareto frontier: the rows for which no other row is both smaller and more accurate. Twenty-one of the forty-six rows are on it. Every legacy format is off it.

Dominated GiB PPL Beaten by GiB PPL
q4_0 4.34 6.7001 iq4_XS (WT 10m) 4.14 6.4597
q4_1 4.78 6.6827 q4_K_M (WT 10m) 4.58 6.3829
q5_0 5.21 6.3632 q5_K_S (None) 5.21 6.3366
q5_1 5.65 6.3379 q5_K_M (None) 5.33 6.2886
q3_K_S (WT 10m) 3.41 7.6029 iq3_XS (WT 10m) 3.28 7.1630
q2_K_S (WT 10m) 2.96 9.3238 iq2_M (WT 10m) 2.74 8.6008
q2_K (WT 10m) 2.96 8.6478 iq2_M (WT 10m) 2.74 8.6008

q4_1 is the clearest case: it is 0.20 GiB larger than q4_K_M and 0.30 perplexity worse. It costs 5.0 bits per weight against q4_K's 4.5 and spends the extra half bit on a per-32-weight minimum that a two-level scale hierarchy over 256 weights does better. The four legacy formats are kept for compatibility, and on this model on this measurement there is no size at which any of them is the right choice.

Below about 3.5 GiB the frontier is entirely I-quants with an importance matrix; from 3.74 to 4.58 GiB it alternates between K and I; above 5 GiB it is K-quants with no imatrix, though the scoreboard does not measure imatrix versions up there, so that last stretch is a gap in the data rather than a finding.

Where INT8, INT4 and FP8 sit

GGUF is one ecosystem. The formats named in most quantisation discussions belong to others, and the distinction that matters is not the bit count but what the bits mean.

Integer formats store a signed or unsigned integer per weight and a floating-point scale per group. INT8 and INT4 are the family GGUF's q8_0 and q4_0 belong to. The design space is entirely in the grouping: per-tensor, per-channel, per-block of 32, per-block of 256 with a second tier. That choice is what separates q4_0 from q4_K at the same nominal width.

Floating-point formats store an exponent and a mantissa per weight. FP8 has two standard shapes, E4M3 with four exponent and three mantissa bits and E5M2 with five and two, trading range against precision. They are hardware formats first: their appeal is that recent accelerators multiply them natively, so the saving is arithmetic throughput as well as memory.

ggml now carries two four-bit floating point block types, and their costs are in the first table: mxfp4 at 17 bytes per 32 weights is 4.25 bits per weight, with a single shared byte-sized exponent; nvfp4 at 36 bytes per 64 weights is 4.5, spending four bytes on UE4M3 scales, one per 16-weight sub-block. The same pattern as the integer formats, then: the finer the scale grouping, the more metadata, and the arithmetic is identical.

What the tables above cannot tell you about these is quality, because the scoreboard does not measure them. That is a real limit and not a small one.

What this does not establish

Code and data

Sources

  1. ggml-org/llama.cpp, ggml/src/ggml-common.h on master, read 2026-09-22 (the block struct definitions and their static_assert size expressions for 27 block types; QK_K 256, QK4_0 and QK5_0 and QK8_0 32, K_SCALE_SIZE 12, IQ3S_N_SCALE QK_K/64, QK_MXFP4 32, QK_NVFP4 64 with QK_NVFP4_SUB 16)
  2. ggml-org/llama.cpp, tools/quantize/README.md on master, read 2026-09-22 (whole-model bits/weight, file size in GiB and prompt-processing and text-generation throughput for 25 quantisation types on meta-llama/Llama-3.1-8B; the Llama 3.1 memory table giving 8B at 32.1 GB original and 4.9 GB at Q4_K_M; the description of quality loss as measured in perplexity and Kullback-Leibler divergence and minimised by a suitable imatrix file)
  3. ggml-org/llama.cpp, tools/perplexity/README.md on master, read 2026-09-22 (the LLaMA 3 8b Scoreboard: revision f364eb6f, CUDA backend, AMD Epyc 7742, 1x NVIDIA RTX 4090; 46 rows of quantisation type, imatrix, model size, perplexity, delta-perplexity, KLD, mean and RMS change in correct-token probability, sorted by KLD relative to FP16; the note that the stored FP16 logits are downcast to 16-bit unsigned integers so the f16 row measures only that downcast; the Wikitext importance matrices at varying token counts)
  4. Meta, Llama-3.1-8B model card, the model the quantize README measures, read 2026-09-22
  5. the importance matrix file the perplexity README links for its WT rows, read 2026-09-22