Mamba
The authors proposed a sequence model with a selective state whose cost grows linearly with length rather than quadratically.
Why it matters
Six years of transformer dominance got their first technically serious challenge, precisely where it is most expensive.
Attention compares every element with every other, so doubling the length quadruples the cost. Mamba instead keeps a state, deciding selectively what to remember, and runs linearly. On long sequences, in genomics and audio among others, that is decisive. It did not displace the transformer but produced hybrid architectures combining both mechanisms.