Back to timeline

Research · October 2001

Gradient boosting

In The Annals of Statistics for October 2001 (volume 29, issue 5, pages 1189-1232) Jerome Friedman of Stanford presented boosting as gradient descent in function space: each new component of an additive model is fitted to the negative gradient of any chosen loss. He gave algorithms for least squares, least absolute deviation, Huber loss and multiclass logistic likelihood, special versions for regression trees (TreeBoost) and tools for interpreting such models. The paper is the 1999 Reitz Lecture.

Why it matters

Boosting stopped being tied to one loss function. AdaBoost's reweighting of examples became a special case of a general recipe that works for regression and classification with any differentiable criterion, and the paper itself connects it to the boosting of Freund and Schapire.

The paper introduces shrinkage: each update is scaled by a learning rate nu between 0 and 1. In a simulation with 5,000 training observations smaller values of nu gave better accuracy and larger optimal numbers of iterations, with a diminishing return below 0.125; the author calls the improvement usually dramatic and its cause a mystery under investigation. Read from the full PDF on Project Euclid in the browser pane, because the site answers the shell with a script page; the Stanford copies did not answer. What the record does not claim: a recommended value such as nu of 0.1 or less (expected, and not in the text), or use of the method in the Netflix Prize, which the source of that record does not show.

Event record

Event date
October 2001
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0661

The Annals of Statistics volume 29, issue 5, October 2001. The paper is the 1999 Reitz Lecture: received May 1999, revised April 2001.

Sources

Related events

Records that link to this one