GPT-4o
OpenAI released a model trained on text, audio and image at once; it answered in voice with a latency of about a third of a second.
Why it matters
Spoken conversation with a machine first ran at human tempo, because the three separate models that lost tone and timing between them were gone.
Until then voice mode was a recogniser, a language model and a synthesiser. Each stage added delay, and more importantly the recogniser passed on words alone: tone, pauses, laughter, several voices in the room never reached the model. A single network trained on audio directly hears all of it and can answer with intonation. The o in the name stands for omni. The model also became free for most ChatGPT users, putting the frontier within reach without payment for the first time.