Matrix Cognition

Paper Readings · · 1,840 words · 8 min read

Where should a document live: context, KV cache, or weights?

A reading of arXiv 2609.17346v1, which compares in-context documents, KV-cache injection and fine-tuning on five benchmarks. Tables re-derived, code included.

RAG long context KV cache fine-tuning

A document the model has never seen can reach it three ways: pasted into the context window, compressed into a key-value cache the model attends over, or trained into the weights. arXiv 2609.17346v1, "Where Should a Document Live: Context, Representations, or Parameters?", posted on 15 September 2026 by Nathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias and Adrià de Gispert of Amazon AGI, runs that comparison under one protocol on five knowledge-intensive benchmarks. We read the abstract page and extracted the PDF locally, because the numbers worth having are in the tables rather than the abstract. Everything quoted below is the paper's measurement; the arithmetic layered on top is ours, it lives in code/where-should-a-document-live.py, and every input value ships in datasets/where-should-a-document-live.csv so the derivations can be checked without re-reading the PDF.

What the paper claims

Four claims, taken from the abstract. In the oracle setting, where the model is handed exactly the gold document, Cartridges (a trained KV-cache prefix) are the most accurate injection method at nearly every storage budget, ahead of the parametric methods by 10 points. Compaction, the other KV method, matches Cartridges only at low compression and falls 10 points behind the parametric methods above 50x. In multi-document retrieval, Cartridges alone keep pace with in-context learning, leading parametric methods by 29 points and Compaction by 15. And Cartridges are one of three methods that damage general ability, degrading control benchmarks by 6%, with 13% on coding.

The paper treats in-context learning as the accuracy target rather than the thing to beat: its question is which cheaper representation of a document gets closest to having that document in the prompt.

The setup

Five benchmarks, described in Table 1: LongHealth (clinical records averaging 11,700 tokens), QASPER (research papers), QuALITY (narratives), T2-RAGBench/FinQA (earnings reports) and TechQA (IBM support technotes). Metrics differ per set: accuracy on the two multiple-choice sets, token F1 on QASPER, exact match on FinQA after evaluating the generated formula to within 1%, and a DeepSeek-Distilled-Qwen-32B judge on TechQA. Four control benchmarks (GSM8K, HumanEval, IFEval, MMLU) measure whether injection damages the base model.

All methods train on the same LLM-generated self-study data, under a distillation objective that matches the teacher's next-token distribution. The base model is Qwen3-8B, with Gemma-3-12B as a replication. At fixed size, the KV methods use 2x compression, LoRA rank 64 on the feed-forward projections, and MLP adapters a bottleneck of 512; the sweep runs from 2x to 100x, with the parametric adapters resized to match the KV footprint at each rate. Retrieval uses 1,024-token chunks, concatenating KV caches for the representation methods and averaging weights for the parametric ones.

One oracle document

Table 2, Qwen3-8B, scores averaged over three runs:

Method LongHealth QuALITY QASPER FinQA TechQA Avg.
No context 37.5 43.6 19.2 2.9 21.1 24.9
ICL 87.4 82.5 56.7 66.8 74.7 73.6
Cartridge 81.1 78.6 54.9 62.7 75.8 70.6
Compaction 87.7 82.1 54.8 66.4 76.0 73.4
LoRA 75.3 73.7 50.3 49.0 74.3 64.5
MLP adapters 74.8 72.3 47.5 44.1 72.5 62.2
Full fine-tuning 69.0 72.8 39.8 41.6 76.0 59.8

The no-context row is the honesty check: 24.9 average says the knowledge is genuinely absent from the base model, so the other rows measure injection rather than recall. Re-deriving the summary column takes six lines:

def check_averages() -> None:
    """Recompute each printed Avg. column from its five per-benchmark cells."""
    for table, model, data in [("Table 2", "Qwen3-8B", TABLE2),
                               ("Table 4", "Gemma-3-12B", TABLE4),
                               ("Table 6", "Qwen3-8B", TABLE6)]:
        for method, vals in data.items():
            cells = [v for v, _ in vals[:5]]          # the five benchmark scores
            recomputed = sum(cells) / len(cells)
            printed = vals[5]                          # the Avg. column as published
            print(f"{table} {model} {method}: printed {printed}, "
                  f"recomputed {recomputed:.2f}, diff {recomputed - printed:+.2f}")

Every printed average reproduces to within 0.04, which is what rounding a mean of five one-decimal numbers should cost.

Two things in this table do not match the abstract's summary, and both are visible without leaving the page. First, at the fixed 2x setting the best injection method is Compaction at 73.4, effectively tied with in-context learning at 73.6, while Cartridges sit at 70.6. The claim that Cartridges are the most accurate injection method is a statement about the storage sweep in Figure 1, not about this table. Second, the 10-point gap over parametric methods is not this table's gap either: Cartridges lead the best parametric method, LoRA, by 6.1 points, lead the mean of the three parametric methods by 8.4, and reach 10.8 only against full fine-tuning, the weakest row. Reading 10 points as the Cartridge-to-LoRA distance at a fixed budget would overstate it by four points.

The sweep values quoted in section 5.1 show why the two KV methods separate. Compaction falls from 66.4 to 19.3 on FinQA between 2x and 20x, and from 87.7 to 46.5 on LongHealth by 100x, while Cartridges stay nearly flat across the same range (LongHealth 81.1 to 77.3, QuALITY 78.6 to 76.4, TechQA 75.8 to 76.9). One method degrades gracefully under compression and the other does not.

More than one document

The composition result is the part with the clearest engineering consequence, and it is severe. Feeding the top-k retrieved chunks and composing their per-document artifacts, only the KV caches survive: from k=1 to k=10 Cartridges rise on LongHealth from 70.8 to 83.2 and hold on TechQA (70.7 to 70.9) and QuALITY (72.5 to 74.6). Everything else decays monotonically. Compaction drops 16.6 points on LongHealth and 39.0 on TechQA over the same range. Merged LoRA adapters fall 33.9 points on TechQA between k=1 and k=3, and merging collapses FinQA from 34.5 at k=1 to 4.9 at k=3, a 29.6-point loss caused by adding two documents.

Table 3 rules out the obvious explanation. Swapping the merge operator between mean, concatenation, TIES and DARE moves the k=5 result by at most 0.8 points on any dataset and not at all on QuALITY. The failure is in combining independently trained adapters at all, not in the choice of arithmetic. Joint training on all documents beats merging as soon as a distractor appears (FinQA 13.0 against 4.9 at k=3), but it still trails composed KV caches and gives up the per-document modularity that made the approach attractive.

We could not re-derive the abstract's 29-point and 15-point multi-document gaps. Those summarise Figure 2, and the full per-point values behind that figure are not printed in the text, so what the shipped CSV holds is the subset of values the results section quotes. The direction is unambiguous; the exact margin is not checkable from the paper alone.

Two-panel chart: left, horizontal bars of the five-benchmark average per method on Qwen3-8B and Gemma-3-12B; right, line segments showing each method's score as the number of retrieved documents rises
Left: the method ranking from Table 2 does not survive the change of base model in Table 4, where Cartridges fall to 37.3. Right: the multi-document values quoted in section 5.2, where only the KV caches hold up as documents are added. Generated by code/where-should-a-document-live.py from the shipped CSV.

The second base model disagrees

Table 4 repeats the single-document protocol on Gemma-3-12B, and the ranking changes. Compaction reaches 65.2 against an in-context oracle of 68.9, LoRA trails at 50.8, and Cartridges collapse to 37.3, which is 33.3 points below their Qwen3-8B result. On the first model Cartridges beat LoRA by 6.1 points; on the second they lose to it by 13.5. The paper attributes this to forgetting, and Table 5 supports that: against a Gemma-3-12B baseline of 88.7 on GSM8K, 83.5 on HumanEval, 78.4 on IFEval and 72.6 on MMLU, the 2x Cartridge scores 36.1, 54.4, 39.2 and 56.9, with standard deviations as wide as 23.9. The headline ordering therefore holds on one of the two models tested.

On Qwen3-8B the same failure is milder: section 6 reports Cartridges degrading control benchmarks by 6% on average, concentrated in code generation, where HumanEval drops 16 points at high compression. LoRA keeps the base model's general ability at every rank, while the widest MLP adapters fall from about 92 to about 60 on GSM8K. The authors read this as evidence that the low-rank constraint, rather than the parameter count, is what preserves general capability, since full fine-tuning and full-rank adapters forget while rank-64 LoRA does not.

What it costs

In-context learning re-processes every retrieved document on every query, from roughly 1k tokens on FinQA to 11k on LongHealth for the document alone. Injection removes that prefill, and because attention is quadratic the paper notes that a 10x token reduction buys about 100x fewer prefill FLOPs. Their measurement, on Qwen3-8B with H200 GPUs: loading 10 cartridges at 20x compression (6K KV tokens, 850 MiB) takes 30 to 50 ms, against 400 to 800 ms to prefill the equivalent 120K raw tokens. The parametric methods are cheapest at inference and do not scale with document length.

A second paper pointing the other way

The oracle framing above treats a full context window as the accuracy ceiling. arXiv 2608.25655v1, "Reconstructing the Right Episode" by Zhexi Feng, Ruiyi Zhang, Yongbo Yang and Pengtao Xie, accepted to EMNLP 2026 and posted on 26 August 2026, measures a setting where that ceiling sags: its SCALE-QA benchmark is 3,000 audited four-way questions over flat mixed-topic threads, where answering means finding which earlier episode still governs a later request.

On GPT-4o-mini, section 6.1 reports full-context accuracy falling from 62.5% at 16k to 29.8% at 128k while the paper's memory system stays at 73.8% using about 1k retrieved tokens. The stress test in section 6.4 is the sharper number: at a 1M-token budget on Gemini 2.5 Flash, feeding the whole thread scores 87.2% at 1.05M prompt tokens and 23.87 seconds per question, against 96.5% from roughly 1.3k retrieved tokens in 2.16 seconds. Even a strong chunk-retrieval baseline, tuned with BM25, dense reranking, HyDE and reciprocal-rank fusion, reaches only 56.2% with three times the context of the winning system.

These two papers do not contradict each other, and it is worth being precise about why. The first measures a single gold document or a short retrieved set against an oracle prompt, where more context is strictly more evidence. The second measures a 128k thread thick with distractors, where more context is mostly more noise. Taken together they say the context window is the target to match when the right document is already identified, and a liability when it is not, which is an argument about retrieval quality rather than about context length.

What neither establishes

The injection paper tests two mid-sized instruction-tuned models and says so; nothing here transfers automatically to frontier-scale models, and it covers knowledge-intensive question answering only, not skill acquisition. Its composition finding is the behaviour of one retrieve-then-compose pipeline using mean merging and prompt concatenation, which the authors flag that reranking or smarter merge schemes could change. All methods train on the same synthetic self-study data, so the comparison inherits that generator's quality. SCALE-QA carries its own caveat, stated in its limitations: it is counterfactually constructed rather than sampled from real logs, and four-way multiple choice cannot measure partial or hedged answers. Neither paper licenses a general claim that retrieval beats long context, or the reverse.

What survives is narrower and more useful. If per-document artifacts are on the table, test composition before committing: merged weights lost 29.6 points on FinQA between one document and three, and no merge operator recovered it. Run the control benchmarks after any injection, since the same method scored 70.6 on one base model and 37.3 on another while forgetting fifty points of grade-school maths. And when a prompt grows past 100k tokens, measure accuracy at that length rather than assuming the window is free; one of these papers found it falling by half.

Code and data

Sources

  1. Rakotonirina, Hardalov, Iglesias, de Gispert (Amazon AGI), "Where Should a Document Live: Context, Representations, or Parameters?" (arXiv 2609.17346v1, 15 Sep 2026; Tables 1-6, sections 5.1, 5.2, 6, 7 and Limitations)
  2. The same paper's PDF, from which the tables quoted here were extracted
  3. Feng, Zhang, Yang, Xie, "Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context" (arXiv 2608.25655v1, 26 Aug 2026, EMNLP 2026 main; Table 2, Table 3, sections 6.1, 6.4, 6.6 and Limitations)