Efficiency
Techniques that reduce parameters, FLOPs, or memory at fixed quality.
14 techniques tagged with this facet.
Attention Mechanisms
- Sparse Transformer Sparse Transformer
- Multi-Query Attention MQA
- Reformer — LSH Attention Reformer
- Sliding Window Attention SWA
- Linear Attention Linear Attention
- Linformer — Low-Rank Attention Projection Linformer
- BigBird BigBird
- Performer — Random Feature Softmax Approximation Performer
- FlashAttention FlashAttention
- Grouped-Query Attention GQA
- Lightning Attention Lightning
- Multi-Head Latent Attention MLA
- DeepSeek Sparse Attention DSA