Facets Inference-only Modifications that touch decoding without retraining. 5 techniques tagged with this facet. Positional Encoding Position Interpolation PI 2023-06 Divide position values by the extension factor s before applying RoPE. Position t becomes t/s; trained rotation angles never extrapolate. Simple, parameter-free, works — but loses resolution uniformly across all frequency bands. NTK-Aware RoPE Scaling NTK-Aware 2023-07 Multiply RoPE's base b by s^(d_h/(d_h-2)) for extension factor s. Fast dimensions (low index) are nearly untouched; slow dimensions get linearly interpolated. Zero fine-tuning needed at small extensions; the inspiration for YaRN's per-band approach. YaRN — Yet Another RoPE eXtensioN YaRN 2023-08 Treat RoPE's rotation bands as three regimes — preserve the fast ones, linearly interpolate the slow ones — and rescale the softmax temperature. The 2023 long-context extension recipe of choice. LongRoPE — Per-Dimension RoPE Search LongRoPE 2024-02 Use evolutionary search to find per-dimension RoPE rescaling factors. Generalizes YaRN's closed-form frequency-band recipe to arbitrary non-monotone schedules. Demonstrated 2M+ context extension on Llama-2 with a short fine-tune. Long Context StreamingLLM and Attention Sinks Attention Sinks 2023-09 The first 1-4 tokens of any pretrained decoder act as attention sinks — they absorb the softmax mass that has nowhere else to go. Pin them in the KV cache and you can slide the rest of the window over arbitrarily long input without quality collapse.