Moshi
The French laboratory Kyutai showed a conversational model that listens and speaks at the same time, with about 200 milliseconds of latency, and released the weights openly.
Why it matters
Conversation with a machine stopped being an exchange of turns: the model could interrupt and be interrupted for the first time.
Every earlier voice system waited for the person to finish. Moshi holds two audio streams in parallel, its own and the other party's, so it can begin answering earlier, fall silent when interrupted, drop in an acknowledgement. Inside it runs the Mimi audio codec and a transformer that generates a text stream alongside the audio, an inner monologue that carries the meaning. Eight researchers built it in half a year, and unlike the GPT-4o voice mode the weights were published openly. For the speech recognition line this is the end of the problem statement itself: there is nothing left to recognise, the audio goes into the model as it is.