Gemini 2.0 Flash
Google released a model that takes in and produces text, images and audio, calls tools by itself, and runs twice as fast as the previous generation's large model.
Why it matters
Multimodality on both input and output converged in one cheap model, and the company named it preparation for agents outright.
Until then image and audio generation lived in separate models and a system was assembled from parts. Here a single pass takes in and gives back all three kinds of data, including a live audio stream in dialogue. The model can also call search and run code without external orchestration. Google framed the release around an agentic era and showed experimental prototypes alongside it, a browser assistant and a coding agent.