INTERNIST-I falls short of experts
On 19 August 1982 NEJM published a systematic evaluation of INTERNIST-I, a program for multiple diagnoses in internal medicine: on 19 of the journal's clinicopathological cases it performed qualitatively like the hospital clinicians but worse than the specialists who discussed the cases.
Why it matters
The program, its authors say, differed from others in the generality of its approach and the size of its knowledge base, covering internal medicine rather than one disease, and they concluded themselves that it was not yet reliable for the clinic. The evaluation named what was missing: reasoning about anatomy and time.
The cases are the Case Records of the Massachusetts General Hospital, which NEJM printed as exercises. The authors named four deficiencies: the program cannot reason anatomically or temporally, cannot build a differential diagnosis spanning several areas, sometimes attributes findings to the wrong causes, and cannot explain its "thinking". Only the abstract was read; the full text is paywalled. The record does not claim the size of the knowledge base (about 500 diseases, 3,500 manifestations) or the number of diagnoses in the cases: neither is in the abstract.