GPT-3 in megawatt-hours
On 21 April 2021 Google and the University of California, Berkeley published an energy account of five large models. Training GPT-3 - 10,000 V100 accelerators over 14.8 days, 3.14e23 operations - cost 1,287 MWh and 552.1 tonnes of CO2 equivalent. That is 0.01055 per cent of Google's total energy use for 2019.
Why it matters
Until then the only figure for a large model's carbon footprint was an estimate made from outside. Here the total is stated alongside what it depends on: the efficiency coefficient of a particular data centre and the carbon intensity of a particular grid in a particular month. The quantity became visible not as a property of the model but as a property of the place and time it was trained in.
For the GPT-3 run the data centre efficiency coefficient was 1.10 and the carbon intensity 0.429 kg per kilowatt-hour; the run took place in 2020 in a Microsoft data centre. The table compares five models: T5, Meena, GShard, Switch Transformer and GPT-3. GPT-3 has the largest footprint, the smallest is 3.2 tonnes. The abstract adds two findings: sparsely activated networks can use under a tenth of the energy of dense ones at the same accuracy, and the choice of network, data centre and processor together moves the footprint by 100 to 1000 times. What the record does not claim. There is no 4M framework in this paper: the strings 4M, Mechanization and Map occur in it zero times, and the mnemonic appeared in later work by the same authors. The GPT-3 figures are a reconstruction from data OpenAI supplied to the authors, not a meter reading.