Ascend 910 and MindSpore
On 23 August 2019 Huawei released the Ascend 910 accelerator - 256 teraFLOPS at FP16, 512 teraOPS at INT8, 310W - together with its own framework, MindSpore. The two are meant to run together: the company claimed ResNet-50 training about twice as fast as other mainstream cards on TensorFlow.
Why it matters
The first attempt at a whole training stack outside the American one: own die, own framework, own tooling. It happened three years before the first export restrictions, which means the closed computing loop was being built ahead of the ban rather than in answer to it.
The figures come from the release: 256 teraFLOPS at half precision, 512 teraOPS on 8-bit integer operations, a maximum draw of 310W, "much lower than its planned specs (350W)". Huawei had announced the planned specs a year earlier at Huawei Connect 2018 and, in the words of Eric Xu, then Rotating Chairman, the chip came out better than expected. MindSpore, by the same release, needs 20 per cent fewer lines of core code on a typical natural-language network and was to go open source in the first quarter of 2020. What this record does not claim. The release does not say where the die was printed: the words TSMC and 7nm appear zero times on the page. The report's statement that the Ascend 910 was fabricated on TSMC's 7nm line does not enter here; the only available source for it is a TechInsights teardown, and techinsights.com answers from this country with a redirect to an "unable to access platform" page. The foundry question stays an empty part, and it is not a small one: whether this stack was in fact closed depends on the answer. Two of the release's numbers are comparisons whose other side it does not name. "About two times faster than other mainstream training cards using TensorFlow" does not say which cards; "20 per cent fewer lines of core code" does not say which frameworks were the baseline.