Trainium reaches general availability
On 10 October 2022 Amazon opened general availability of EC2 Trn1 instances built on its own AWS Trainium silicon. The largest configuration carries 16 chips, 512 gigabytes of high-bandwidth memory, up to 3.4 petaFLOPS in TF32, FP16 and BF16, and up to 800 gigabits a second of network.
Why it matters
The largest landlord of compute stops being only a reseller of other people's accelerators. This is where the branching starts: from here the question of what training costs has different answers depending on whose hardware it runs on, and an operator can compare its own line against the market's.
The figures come from the announcement: trn1.32xlarge carries 16 Trainium chips and 128 vCPUs, up to 512 gigabytes of high-bandwidth memory, up to 3.4 petaFLOPS at TF32/FP16/BF16, four 2-terabyte NVMe drives, and NeuronLink between chips. These are the first EC2 instances with Elastic Fabric Adapter networking up to 800 gigabits a second. The comparison is against P4d instances: 1.4x the teraFLOPS at BF16, 2.5x at TF32, 5x at FP32, 4x the inter-node network bandwidth and up to 50 per cent cost-to-train savings. The preview was announced at re:Invent 2021. One detail that is easy to get wrong. Trainium is AWS's second-generation machine-learning chip, the one after Inferentia, not a second generation of Trainium itself; the announcement says so in as many words. A second Trainium comes later. The record gives no per-chip memory bandwidth: the 820 gigabytes a second the research report quoted is not on the announcement page.