Chainer: the graph is built as it computes
On 5 June 2015 Preferred Networks of Tokyo released Chainer 1.0, a Python deep learning library in which the computation graph is not declared in advance but recorded by itself as the forward computation runs. In December that year Seiya Tokui, Kenta Oono, Shohei Hido and Justin Clayton named the approach Define-by-Run at the LearningSys workshop of NIPS.
Why it matters
A model with loops and branches, such as a recurrent network over inputs of varying length, could be written in ordinary code and debugged with an ordinary debugger. Before, frameworks such as Caffe first read the model's description from a separate file and only then ran it. PyTorch names Chainer in its README among the work it drew on.
The paper sets Define-by-Run against Define-and-Run, in which the model is given as a Protobuf file for Caffe or YAML for PyLearn2, and names three drawbacks of the latter: memory held for the whole graph, limited extensibility and no use of the language's debugger. For GPUs the library has its own CuPy, partly compatible with NumPy. Against Caffe, Chainer's first mini-batch is up to 7.0 times slower than later ones, the forward pass 1.1 to 1.8 times slower, the backward pass comparable or faster. The 2019 PyTorch paper says Chainer implements the same approach at the cost of performance. What the record does not claim: that PyTorch was built on Chainer's architecture, as the research report behind the candidate put it - the README names it alongside autograd and torch-autograd.