Genie 3
DeepMind showed a model that generates from a text description a world one can walk through in real time: 24 frames a second, several minutes without losing consistency.
Why it matters
Generated video stopped being a recording and became an environment: the frame depends on where the viewer just turned.
Sora made a minute-long clip that stayed the same every time. Here each frame is generated with the user's action in view, and the model has to remember what it has already shown: turn away and back, and the wall must still be there. That consistency over time is the main technical difference. DeepMind presented the work as a step towards training agents in generated environments, where experience is not limited to recorded data.