The UCI Machine Learning Repository
In 1987 David Aha, a doctoral student at the University of California, Irvine, opened an FTP archive of data sets for testing machine learning algorithms. By 4 December 1995 it held 111 databases and domain theories, 36 megabytes in all.
Why it matters
Machine learning gained a shared proving ground: an algorithm could be tested on the same data as someone else's, and someone else's experiment repeated. In 2001 Breiman compared random forests with AdaBoost on thirteen data sets from it.
Each data set is a file of attribute-value records, one per line, with a separate documentation file. Access was by anonymous FTP, over the web and, for those without FTP, by e-mail through an archive server. Most data sets were donated by researchers outside Irvine; three oncology databases from Ljubljana were available only on request. The 1995 README asked users to acknowledge the repository so that others could obtain the same data and replicate the experiments; the home page in May 1997 spoke of over one hundred entries. The year 1987 is borne out by a document of 1990: in Aha's dissertation, dated 27 November 1990, his curriculum vitae names him organizer of the UCI repository of machine learning databases in 1987-90. What the record does not claim. No document from 1987 was found, so the record does not know the month. How many data sets there were at the start is unknown.