GUHA: a machine that proposes hypotheses from data
In 1966 Petr Hájek, Ivan Havel and Metoděj Chytil of Prague published the GUHA method (General Unary Hypotheses Automaton): a computer systematically generates every hypothesis of a given form about relations between properties of objects and outputs those that hold in the data. The Czech paper in Kybernetika was tried on a MINSK 2 computer with data on 230 epileptic patients, 36 properties each.
Why it matters
Finding hypotheses in data became a program rather than only a researcher's intuition; the authors call the method 'an offering of hypotheses'. More than thirty years later GUHA's authors pointed out that the association rules of data mining are in substance already in their paper of 1966. This is an editorial assessment resting on their account.
There are two papers, both of 1966 and by the same authors: a Czech one in Kybernetika 2, No. 1, pp. 31-47 (received 19 September 1965; Hájek and Havel at the Mathematical Institute of the Czechoslovak Academy of Sciences, Chytil at the Academy's Physiological Institute), and an English one in Computing 1, No. 4, pp. 293-308 (received 31 March 1966), where the program ran on the IBM 7040 of the Technische Hochschule Wien. On the MINSK 2 one block of main memory, 4,096 words of 36 bits, sufficed. The English paper itself says that mathematically the method contains no innovations; what is new is the complete machine search. On association rules: technical report No. 867 of the Institute of Computer Science of the Academy of Sciences of the Czech Republic (2002) writes that Agrawal's notion of an association rule 'occurs in fact' in the 1966 paper and that this long went unnoticed; the report cites the 1996 chapter by Agrawal and co-authors. This is a claim by the method's own authors. What the record does not claim: that Agrawal and co-authors drew on GUHA, which no document read says; how many hypotheses the Prague experiment found.