Weights with three states instead of numbers
On 27 February 2024 Microsoft Research and the University of Chinese Academy of Sciences showed a language model in which every weight is −1, 0 or 1. At 3 billion parameters it matched a full-precision LLaMA on perplexity while using 3.55 times less GPU memory and running 2.71 times faster.
Why it matters
Until then lowering precision was compression applied to an already-trained model, and it was paid for in quality. Here precision became a property of the training itself, and multiplication in the linear layers collapsed into addition and subtraction — so what can run outside a data centre stopped depending on having something to compress.
Three states per weight need log₂(3) ≈ 1.58 bits, which is where the name comes from. Activations stay at 8 bits. The comparison is against the authors' own reproduced FP16 LLaMA, both trained on 100 billion RedPajama tokens. From Table 1 of the first version: parity on perplexity arrives at the 3B size — 2.71 times faster and 3.55 times less GPU memory. The 3.9B configuration gives 2.40 times faster and 3.32 times less memory. Separately, on 7nm chips the arithmetic of matrix multiplication consumes 71.4 times less energy. The record does not claim that 3.9B wins everywhere. By Table 2 it beats FP16 LLaMA 3B on six zero-shot tasks out of seven and loses on OpenbookQA (24.2 against 24.6); the average is 51.2 against 49.7. The claim "on all tests" is not supported by the source. The largest model in the paper is 3.9 billion parameters; how the scheme behaves at tens of billions is not measured here.