Back to timeline

Availability · July 3 – September 17, 2024

Moshi

The French laboratory Kyutai showed a conversational model that listens and speaks at the same time, with about 200 milliseconds of latency, and released the weights openly.

Why it matters

Conversation with a machine stopped being an exchange of turns: the model could interrupt and be interrupted for the first time.

Every earlier voice system waited for the person to finish. Moshi holds two audio streams in parallel, its own and the other party's, so it can begin answering earlier, fall silent when interrupted, drop in an acknowledgement. Inside it runs the Mimi audio codec and a transformer that generates a text stream alongside the audio, an inner monologue that carries the meaning. Eight researchers built it in half a year, and unlike the GPT-4o voice mode the weights were published openly. For the speech recognition line this is the end of the problem statement itself: there is nothing left to recognise, the audio goes into the model as it is.

Event record

Event date
July 3, 2024 – September 17, 2024
Timeline date
Event date
Verification
Sources gathered automatically · September 18, 2026
Lines
ID
evt-0304

The model was shown in Paris on 3 July 2024; the technical report and weights were released in September.

Sources

Related events

Records that link to this one