gpt-realtime
The speech model that takes in and produces audio directly came out of preview, with tool calling, images on input and telephone calls.
Why it matters
A line that began with the ARPA programme of 1971 ended in a service where there is no text between the person and the model at all.
Speech recognition held as a separate problem for half a century: audio in, text out, and something else after that. Here there is no text in the path; the model hears and speaks, and the text stream remains only as an internal trace. The practical part: support for remote tool servers, image input and connection to the telephone network. A voice agent could take a call, look at a document and carry out an action, all in one session.