Browse
Learning paths
Each path is a hand-picked ordering of three or four technique entries that, taken together, tell one coherent story. Short, chronological, and not a substitute for the full category index.
- From MHA to MLA Intermediate
Walk the four canonical attention layouts in chronological order — the path from full multi-head to grouped to latent compression.
4 steps
- Norm placement and type Intro
Where to put the normalization layer, and which variant to use. Four short entries to grasp the 2024+ consensus.
4 steps
- Sparse MoE foundations Intermediate
The four entries that built today's sparse mixture-of-experts FFN: the original sparse routing idea, Switch's k=1 simplification, Mixtral's open instantiation, and DeepSeek's shared-expert refinement.
4 steps