Browse
References
Every paper cited as the introducing source for a technique in this knowledge base. 48 unique papers across 10 years (2015–2025). All links go to arXiv abstracts.
2015
-
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · Deep Residual Learning for Image Recognition arXiv:1512.03385
Used in: The Residual Stream (Residual)
2016
-
Dan Hendrycks, Kevin Gimpel · Gaussian Error Linear Units (GELUs) arXiv:1606.08415
Used in: Gaussian Error Linear Unit (GELU)
-
Jimmy Lei Ba, Jamie Ryan Kiros, Geoffrey E. Hinton · Layer Normalization arXiv:1607.06450
Used in: Layer Normalization (LayerNorm)
2017
-
Noam Shazeer et al. · Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer arXiv:1701.06538
Used in: Sparsely-Gated MoE (Sparse MoE)
-
Ashish Vaswani et al. (Google Brain) · Attention Is All You Need arXiv:1706.03762
Used in: Multi-Head Attention (MHA) · FFN with ReLU (FFN-ReLU) · Sinusoidal Position Encoding (Sinusoidal)
2019
-
Rewon Child, Scott Gray, Alec Radford, Ilya Sutskever (OpenAI) · Generating Long Sequences with Sparse Transformers arXiv:1904.10509
Used in: Sparse Transformer (Sparse Transformer)
-
Biao Zhang and Rico Sennrich · Root Mean Square Layer Normalization arXiv:1910.07467
Used in: Root Mean Square Layer Normalization (RMSNorm)
-
Noam Shazeer · Fast Transformer Decoding: One Write-Head is All You Need arXiv:1911.02150
Used in: Multi-Query Attention (MQA)
-
Jack W. Rae et al. (DeepMind) · Compressive Transformers for Long-Range Sequence Modelling arXiv:1911.05507
Used in: Compressive Transformer (Compressive)
2020
-
Nikita Kitaev, Łukasz Kaiser, Anselm Levskaya (Google Research) · Reformer: The Efficient Transformer arXiv:2001.04451
Used in: Reformer — LSH Attention (Reformer)
-
Noam Shazeer · GLU Variants Improve Transformer arXiv:2002.05202
Used in: GELU-Gated Linear Unit (GeGLU) · ReLU-Gated Linear Unit (ReGLU) · Swish-Gated Linear Unit (SwiGLU)
-
Ruibin Xiong et al. · On Layer Normalization in the Transformer Architecture arXiv:2002.04745
Used in: Pre-Norm, Post-Norm, and Sandwich Placement (Norm Placement)
-
Thomas Bachlechner et al. (UCSD) · ReZero is All You Need: Fast Convergence at Large Depth arXiv:2003.04887
Used in: ReZero — Residual With Learnable Skip Scale (ReZero)
-
Iz Beltagy, Matthew E. Peters, Arman Cohan · Longformer: The Long-Document Transformer arXiv:2004.05150
Used in: Sliding Window Attention (SWA)
-
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François Fleuret · Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention arXiv:2006.16236
Used in: Linear Attention (Linear Attention)
-
Sinong Wang et al. (Facebook AI) · Linformer: Self-Attention with Linear Complexity arXiv:2006.04768
Used in: Linformer — Low-Rank Attention Projection (Linformer)
-
Dmitry Lepikhin et al. (Google Research) · GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding arXiv:2006.16668
Used in: GShard (GShard)
-
Manzil Zaheer et al. (Google Research) · Big Bird: Transformers for Longer Sequences arXiv:2007.14062
Used in: BigBird (BigBird)
-
Krzysztof Choromanski et al. (Google Research) · Rethinking Attention with Performers arXiv:2009.14794
Used in: Performer — Random Feature Softmax Approximation (Performer)
-
Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, Yuxuan Chen · Query-Key Normalization for Transformers arXiv:2010.04245
Used in: Query-Key Normalization (QK-Norm)
2021
-
William Fedus, Barret Zoph, Noam Shazeer · Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity arXiv:2101.03961
Used in: Switch Transformer (Switch)
-
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, Yunfeng Liu · RoFormer: Enhanced Transformer with Rotary Position Embedding arXiv:2104.09864
Used in: Rotary Position Embedding (RoPE)
-
Ming Ding et al. (Tsinghua University — CogView) · CogView: Mastering Text-to-Image Generation via Transformers arXiv:2105.13290
Used in: Sandwich-LN (Sandwich-LN)
-
Ofir Press, Noah A. Smith, Mike Lewis · Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation arXiv:2108.12409
Used in: Attention with Linear Biases (ALiBi)
-
Sam Shleifer, Jason Weston, Myle Ott (Facebook AI) · NormFormer: Improved Transformer Pretraining with Extra Normalization arXiv:2110.09456
Used in: NormFormer — Extra Normalization in the Residual (NormFormer)
2022
-
Yuhuai Wu, Markus N. Rabe, DeLesley Hutchins, Christian Szegedy · Memorizing Transformers arXiv:2203.08913
Used in: Memorizing Transformers (Memorizing)
-
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, Omer Levy · Transformer Language Models without Positional Encodings Still Learn Positional Information arXiv:2203.16634
Used in: No Position Encoding (NoPE)
-
Hongyu Wang et al. (Microsoft Research) · DeepNet: Scaling Transformers to 1,000 Layers arXiv:2203.00555
Used in: DeepNet — Scaling Transformers to 1000 Layers (DeepNet)
-
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré · FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness arXiv:2205.14135
Used in: FlashAttention (FlashAttention)
2023
-
Joshua Ainslie et al. (Google Research) · GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints arXiv:2305.13245
Used in: Grouped-Query Attention (GQA)
-
Amirkeivan Mohtashami, Martin Jaggi (EPFL) · Landmark Attention: Random-Access Infinite Context Length for Transformers arXiv:2305.16300
Used in: Landmark Attention (Landmark)
-
Shouyuan Chen, Sherman Wong, Liangjian Chen, Yuandong Tian (Meta AI) · Extending Context Window of Large Language Models via Positional Interpolation arXiv:2306.15595
Used in: Position Interpolation (PI)
-
Jiayu Ding et al. (Microsoft Research) · LongNet: Scaling Transformers to 1,000,000,000 Tokens arXiv:2307.02486
Used in: LongNet — Dilated Attention (LongNet)
-
Anonymous (Reddit user 'bloc97'); later formalized by Peng et al. · Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning arXiv:2309.00071
Used in: NTK-Aware RoPE Scaling (NTK-Aware) · YaRN — Yet Another RoPE eXtensioN (YaRN)
-
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis · Efficient Streaming Language Models with Attention Sinks arXiv:2309.17453
Used in: StreamingLLM and Attention Sinks (Attention Sinks)
2024
-
Zhen Qin et al. (MiniMax) · Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models arXiv:2401.04658
Used in: Lightning Attention (Lightning)
-
Damai Dai et al. (DeepSeek-AI) · DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models arXiv:2401.06066
Used in: DeepSeekMoE (DeepSeekMoE)
-
Mistral AI · Mixtral of Experts arXiv:2401.04088
Used in: Mixtral-Style Coarse MoE (Mixtral MoE)
-
Peitian Zhang, Zheng Liu, Shitao Xiao, Ninglu Shao, Qiwei Ye, Zhicheng Dou · Soaring from 4K to 400K: Extending LLM's Context with Activation Beacon arXiv:2401.03462
Used in: Activation Beacon (Activation Beacon)
-
Yiran Ding et al. (Microsoft Research) · LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens arXiv:2402.13753
Used in: LongRoPE — Per-Dimension RoPE Search (LongRoPE)
-
DeepSeek-AI · DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model arXiv:2405.04434
Used in: Multi-Head Latent Attention (MLA) · Decoupled RoPE (Decoupled RoPE)
-
Lean Wang, Huazuo Gao, Chenggang Zhao, Xu Sun, Damai Dai (DeepSeek-AI) · Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts arXiv:2408.15664
Used in: Auxiliary-Loss-Free Load Balancing (Aux-Loss-Free)
-
Defa Zhu et al. (ByteDance Doubao) · Hyper-Connections arXiv:2409.19606
Used in: Dynamic Hyper-Connections (DHC / mHC) · Hyper-Connections (HC)
-
Ilya Loshchilov et al. (NVIDIA) · nGPT: Normalized Transformer with Representation Learning on the Hypersphere arXiv:2410.01131
Used in: nGPT — Normalized Transformer on the Hypersphere (nGPT)
-
Allen Institute for AI (AI2) · 2 OLMo 2 Furious arXiv:2501.00656
Used in: OLMo 2 Reordered Post-Norm (OLMo 2 Post-Norm)
2025
-
Google DeepMind (Gemma 3 team) · Gemma 3 Technical Report arXiv:2503.19786
Used in: Gemma 3 Norm-Everywhere (Norm-Everywhere)
-
Moonshot AI · Kimi K2: Open Agentic Intelligence arXiv:2507.20534
Used in: Kimi K2 MoE (K2 MoE)
-
DeepSeek-AI · DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models arXiv:2512.02556
Used in: DeepSeek Sparse Attention (DSA)