Imagen: a large language encoder for generation
On 23 May 2022 Google Research described Imagen, a diffusion model that generates images from text encoded by a frozen T5-XXL language model. Without training on COCO it reached an FID of 7.27. Google released neither code nor a public demo.
Why it matters
The paper showed that for text-to-image generation, scaling a language encoder trained on text alone helps more than scaling the diffusion model itself. Google also declined an open release because of misuse risks.
The paper also proposes DrawBench, a set of prompts for comparing models by human raters. It explains the decision to release neither code nor a demo by the use of generative methods for harassment and misinformation and by concerns about bias; TechCrunch confirmed on 20 July 2022 that Google would not release Imagen because of misuse risks.