Deep networks in YouTube recommendations
On 7 September 2016 Paul Covington, Jay Adams and Emre Sargin of Google described, for the RecSys conference, YouTube's recommender system built on deep networks in two stages: candidate generation narrows a corpus of millions of videos to hundreds, and a separate ranking network orders them by predicting expected watch time rather than the probability of a click. The models have about a billion parameters and are trained on hundreds of billions of examples.
Why it matters
One of the largest recommender systems was described by its developers: deep learning replaced matrix factorisation in it, and the ranking objective moved from the click to watch time. What the system rewards depends on that objective, and the paper names it plainly.
The paper: YouTube's recommendations help 'more than a billion users'; the system is built on Google Brain, recently open-sourced as TensorFlow; the candidate stage's predecessor was matrix factorisation under a rank loss, which the new model outperformed. In the table of hidden-layer configurations the deepest (1024, 512 and 256 ReLU units) gives the lowest weighted pairwise loss on next-day data, 34.6 per cent. There is no independent confirmation of scale: this is the company's own account. The record does not claim what share of views recommendations bring: the paper has no such figure.