The gated recurrent unit (GRU)
On 3 June 2014 Kyunghyun Cho, Yoshua Bengio and co-authors at the Universite de Montreal, the University of Gothenburg and the Universite du Maine posted an RNN Encoder-Decoder: one recurrent network encodes a phrase into a fixed-length vector and another decodes it into a phrase in the other language. For it they proposed a new hidden unit with a reset gate and an update gate, motivated by the LSTM but much simpler. As a feature in a phrase-based English-French system it raised test BLEU from 29.33 to 29.96, and to 31.18 together with a neural language model and a word penalty.
Why it matters
Recurrent networks got a gated unit much simpler than the LSTM: the reset gate lets it drop the previous state, the update gate decides how much of it to carry forward. The attention paper of September 2014 built its translation model on this unit, and in December 2014 Chung and co-authors, comparing it with the LSTM, gave it the name GRU.
The name GRU does not occur in the preprint (zero matches): it calls the unit a new type of hidden unit. The name appears in Chung, Gulcehre, Cho and Bengio, arXiv:1412.3555, 11 December 2014, which found the GRU comparable to the LSTM on music and speech modelling. Encoder and decoder had 1,000 hidden units each; vocabularies were limited to the 15,000 most frequent words, covering about 93 per cent of the data. Table 1 of the first version, read from the rendered page because the text layer misorders the rows: baseline 27.63 dev and 29.33 test BLEU, with a continuous-space language model 28.33 and 29.58, with the RNN Encoder-Decoder 28.48 and 29.96, with both 28.60 and 30.64, with both and a word penalty 28.93 and 31.18. What the record does not claim: a gain of about one BLEU point from the recurrent network alone, which was expected and is 0.63.