Systems
Inference, quantisation, GPUs, cost, latency: the engineering under the model.
-
Systems ยท
KV cache arithmetic: memory per token, batch size, and context cost
Bytes of KV cache per token from a config.json, worked for Qwen3-8B and Qwen3-32B at four context lengths on an 80 GB device, with GQA and int8 KV as the two levers.