Facets Parameter-free Reformulations that add no learnable parameters. 1 technique tagged with this facet. Positional Encoding Attention with Linear Biases ALiBi 2021-08 Bias attention scores by a per-head linear function of the query-key distance. No learnable position parameters; extrapolates beyond training length without any fine-tune. Lost the dominance race to RoPE for dense decoders but is mechanically illuminating.