Facets KV cache Tricks that shrink the per-token attention cache during generation. 3 techniques tagged with this facet. Attention Mechanisms Multi-Query Attention MQA 2019-11 One K, V projection shared by every query head. H× smaller KV cache than MHA; small but real quality drop that motivated GQA. The first move in the MHA → MQA → GQA → MLA evolution. Grouped-Query Attention GQA 2023-05 Group query heads so each group reads one shared K, V pair. 4-8× KV-cache reduction with quality near MHA; the dominant attention layout for dense LLMs from 2023 onward. Multi-Head Latent Attention MLA 2024-05 Compress K and V to a small per-token latent; reconstruct heads at attention time. ~5–7× smaller KV cache than MHA on DeepSeek-V2 ablations.