Back to timeline

Availability · May 13, 2024

GPT-4o

OpenAI released a model trained on text, audio and image at once; it answered in voice with a latency of about a third of a second.

Why it matters

Spoken conversation with a machine first ran at human tempo, because the three separate models that lost tone and timing between them were gone.

Until then voice mode was a recogniser, a language model and a synthesiser. Each stage added delay, and more importantly the recogniser passed on words alone: tone, pauses, laughter, several voices in the room never reached the model. A single network trained on audio directly hears all of it and can answer with intonation. The o in the name stands for omni. The model also became free for most ChatGPT users, putting the frontier within reach without payment for the first time.

Event record

Event date
May 13, 2024
Timeline date
Event date
Verification
Sources gathered automatically · September 18, 2026
Lines
ID
evt-0299

Announcement of 13 May 2024.

Sources

Related events

Records that link to this one