Rotary position embedding (RoPE)
On 20 April 2021 Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen and Yunfeng Liu of Zhuiyi Technology in Shenzhen posted RoFormer: the position of a token is encoded with a rotation matrix that turns the vector through an angle depending on position, and self-attention naturally acquires a dependency on relative distance. The experiments of the first version are Chinese only: pre-trained on about 34 GB of Chinese text, on the CAIL2019-SCM legal case-matching task RoFormer reached 69.79 per cent on test at length 1,024, against 68.10 for WoBERT at 512.
Why it matters
Position became a property of the attention product itself rather than a vector added to the input, and, as the paper shows, this is compatible with linear attention. LLaMA in 2023 removed absolute positional embeddings and took the rotary ones introduced by Su et al. (2021); the Falcon model card names them too.
The first version says evaluation on English data is still under way. The WMT 2014 English-German comparison usually quoted with RoPE, 27.5 BLEU against 27.3 for the base Transformer, appears only in the second version of 9 October 2021 and is not this record's result. CAIL2019-SCM holds 8,964 triplets of cases published by the Supreme People's Court of China; test accuracy at length 512 is 67.77 per cent for BERT, 68.10 for WoBERT and 68.29 for RoFormer.