Predicting clicks on new ads
On 8 May 2007 Matthew Richardson, Ewa Dominowska and Robert Ragno of Microsoft published, in the proceedings of the WWW conference, a model that predicts the click-through rate of search ads with no history of impressions. Logistic regression on features of the ad, its keywords and its advertiser cut the estimation error by 30 per cent against a baseline that simply predicts the average click-through rate.
Why it matters
The order of ads in paid search depends on the predicted probability of a click, and the paper showed that this probability can be learned for a new ad instead of waiting for it to gather impressions. Here machine learning directly sets a search engine's revenue; Facebook's paper of 2014 on click prediction cites this one.
The data: '10,000 advertisers, with over 1 million ads for over a half million keywords'. The measure is the average KL-divergence between predicted and true click-through rate: term click-through features gave 13 per cent, with related terms nearly 20, the full model 30. The record does not claim that the model was deployed in Microsoft's ad system: the paper's conclusion says only that it is efficient enough for any of the major search engines to use. The paid-search market of $5.75 billion and Google's $1.63 billion of search ad revenue in the third quarter of 2006 are cited from the press, not measured by the paper.