A million tokens of context
Google announced Gemini 1.5 Pro with a context window of one million tokens: an hour of video, eleven hours of audio, or whole code bases in a single request.
Why it matters
The amount a model holds in view grew tens of times over, and the question of searching documents turned into the question of reading them all.
Until then a typical window was a few tens of thousands of tokens, and any work with a large corpus meant splitting it and retrieving the relevant pieces. A million tokens made that step unnecessary for most documents. In tests that hide a fact inside a long text, the model found it reliably across the whole window. That moved engineering attention away from building retrieval pipelines and towards what to do with the answer.