Sora 2
The second generation of the video model generated audio together with the image — speech that matches the lips, ambient sound, music — and shipped as a separate application with a feed.
Why it matters
Video generation stopped being a tool and became a social network, with everything that implies for trust in a recording.
Technically the main thing is synchronised audio: until then video was generated silent and scored separately. Here one model holds picture and sound together, and speech matches the movement of the lips. The product side turned out to matter more than the technical one: the application let people insert real faces into a clip with their permission, and the feed filled with convincing invented scenes featuring recognisable people.