NIST's shared test on handwritten characters
On 27-28 May 1992 NIST, for the US Census Bureau, convened in Gaithersburg the participants of the first Census optical character recognition systems conference: 26 organisations from North America and Europe recognised the same roughly 85,000 handwritten digits and letters, unseen before. About half the systems recognised over 95 percent of the digits; a human, about 98.5.
Why it matters
Over forty-five handwriting recognition systems, from neural networks to rules, were compared on the same data with one count of errors. The conference's training disc, NIST Special Database 3, later became half of MNIST.
Participants received a training disc of over 300,000 images of separate characters and a test disc, with results due by 27 April; 26 organisations returned over 115 submissions from over 45 systems, while three, HNC among them, returned none. The lowest error on digits with no rejection was 1.56 percent (the OCRSYS system); AT&T's four systems ranged from 3.16 to 4.84. Half the systems read upper-case letters with under 10 percent error and lower-case with under 20. SD-3 was written by Census Bureau employees, and MNIST mixed it with a database written by high-school students because SD-3 is much cleaner. What the record does not claim. The report itself warns that the test was not proctored and that its figures must not be a basis for purchasing; it tested only characters already cut out, not whole forms.