Facets Hardware-aware Designs that exploit modern accelerator memory hierarchies. 2 techniques tagged with this facet. Attention Mechanisms FlashAttention FlashAttention 2022-05 Tile attention so QK^T and the softmax stay in SRAM rather than round-tripping through HBM. Exact (not approximate) attention, 2–4× faster, 5–20× less peak memory. The universal kernel under every modern transformer trainer. Lightning Attention Lightning 2024-01 Tile and fuse linear attention's prefix-sum recurrence so it runs faster than FlashAttention at long context. Interleaved 7:1 with softmax attention in MiniMax-01 to recover what linear attention loses on absolute quality while keeping its linear-in-T scaling.