On 24 April 1939 Homer Dudley of Bell Telephone Laboratories read a paper on the automatic synthesis of speech before the US National Academy of Sciences; it appeared in PNAS that July. An analyser broke a talker's voice into 11 slow signals: one carried the pitch, ten the power in bands from 0 to 2950 cycles, each passed through a filter of up to 25 cycles. A synthesiser rebuilt speech from them out of two sources: a buzz from a relaxation oscillator for voiced sounds and a hiss from a gas tube for unvoiced ones.
September 1940 — February 1, 1942Research · FoundationsLines: Learning theory · Defence
In September 1940 Norbert Wiener offered his services to Vannevar Bush, who ran American war research, and took up the anti-aircraft gunner's problem: predicting where an aircraft will be when the shell reaches it. On 1 December the NDRC let MIT a contract, and with the engineer Julian Bigelow Wiener built a predictor that computed the target's future position from its past course.
March 1946Milestone · FoundationsLines: Learning theory
The Macy Foundation convened its first conference on feedback mechanisms and circular causality in biological and social systems. The group drew physicists, mathematicians, electrical engineers, physiologists, neurologists, psychiatrists, sociologists and anthropologists.
In July 1946 W. Koenig, H. K. Dunn and L. Y. Lacy of Bell Labs published the design of the sound spectrograph, a wave analyzer that produces a permanent visual record of how a sound's energy is distributed in frequency and in time at once. The paper describes how the device works, gives the mechanical arrangement and the circuits of one particular model, and shows spectrograms of voice, animal and bird sounds, music and frequency modulations.
February 20, 1947Milestone · FoundationsLines: Learning theory
Turing told a mathematical audience that a machine should learn from experience, and that a machine expected to be infallible cannot also be intelligent.
In May 1951 Franklin S. Cooper, Alvin M. Liberman and John M. Borst of Haskins Laboratories published, in the Proceedings of the National Academy of Sciences, a paper on the interconversion of audible and visible patterns as a basis for research in the perception of speech. It had been read before the Academy on 10 October 1950. In November 1952 the same group reported what came of it: synthetic syllables allowed a systematic exploration of the acoustic cues by which people perceive consonants.
November 6, 1950 – December 25, 1951Milestone · FoundationsLines: Compute
In S. A. Lebedev's laboratory at the Institute of Electrical Engineering of the Ukrainian Academy of Sciences, at Feofaniia near Kyiv, the prototype of the Small Electronic Calculating Machine (MESM) was first started on 6 November 1950, and on 25 December 1951 a government commission accepted it into regular operation. That day the machine solved its first real problem: 585 values of a probability distribution function, nearly 250,000 operations in 2.5 hours.
In November 1952 K. H. Davis, R. Biddulph and S. Balashek of Bell Labs described a circuit that automatically recognises ten telephone-quality digits spoken at normal speech rates by a single individual, with an accuracy varying between 97 and 99 percent. The spectrum was split into two bands, one below and one above 900 cycles per second, axis-crossing counts were made in each, and the result was compared against ten built-in standards, one per digit, with the best match selected.
June 1953Research · FoundationsLines: Learning theory
The paper described how to sample from a distribution that cannot be computed directly, by accepting or rejecting a random step on a ratio of probabilities.
1953Research · FoundationsLines: Search and reasoning
In a chapter of Faster Than Thought, Turing set out rules for choosing a chess move that "could without difficulty be made into a machine programme", and printed a game played by those rules against a weak player.
In November 1956 Harry F. Olson and Herbert Belar of RCA Laboratories described a phonetic typewriter, a simplified but nevertheless complete system that types in response to words spoken into a microphone. The paper works through five problems: the form in which the words are typed, the means of analysing the sounds of speech, the identification of the analysed sounds, the encoding and decoding that drives the actuating mechanism, and the design of that mechanism.
On 9 December 1957 Russell Kirsch and three colleagues at the US National Bureau of Standards described a drum scanner attached to the SEAC computer. In 25 seconds it turned a 44 x 44 mm picture into 176 by 176 binary digits, and a program displayed the picture back on an oscilloscope.
July 1, 1958Milestone · FoundationsLines: Defence · Compute
On 1 July 1958 the first AN/FSQ-7 direction centre was declared operational at McGuire Air Force Base. The machine took in radar data, assembled a single picture of the air situation from it, and guided weapons onto targets.
November 24 – 27, 1958Research · FoundationsLines: Learning theory
Selfridge described recognition as several layers of simple agents: some see features of the image, others assemble guesses from them, and a last one picks the loudest.
June 1959Research · FoundationsLines: Search and reasoning · Symbolic AI
Newell, Shaw and Simon separated the solving method from the particular problem: the program reduces the difference between the current and the goal state.
July 3, 1959Research · FoundationsLines: Science · Learning theory
On 3 July 1959 Robert Ledley and Lee Lusted published 'Reasoning Foundations of Medical Diagnosis' in Science. They broke diagnosis into symbolic logic linking symptoms to diseases, probability by an adaptation of Bayes' formula, and value theory for choosing a treatment, and showed the computation on edge-notched cards sorted by hand.
October 1959Research · ImagesLines: Learning theory · Science
In October 1959 David Hubel and Torsten Wiesel described the fields of single cells in the cat's primary visual cortex: each field split into excitatory and inhibitory regions lying side by side, and a cell was driven only by a stimulus of a particular shape, size, position and orientation.
1959Research · FoundationsLines: Symbolic AI · Search and reasoning
In early spring 1959 an IBM 704 at the IBM Research Center, running a program of some 20,000 instructions, proved its first theorem of Euclidean plane geometry. Herbert Gelernter's machine accepted a subgoal only if it held in the diagram, and so rejected false steps before trying to prove them.
January 1960Research · FoundationsLines: Symbolic AI · Learning theory
In January 1960 Hao Wang described in the IBM Journal programs for the IBM 704 that proved over 200 theorems of the propositional calculus from the first five chapters of Principia Mathematica in under 3 minutes of proving time, and 139 of 158 theorems from its predicate part. He wrote the programs in the summer of 1958 at IBM's Poughkeepsie laboratory.
In spring 1960 engineers at the Cornell Aeronautical Laboratory described the Mark I, Rosenblatt's perceptron in hardware. A retina of 20 x 20 photoresistors, 512 association units, 8 response units, and weights that were potentiometers turned by electric motors. After sixteen exposures to each letter the machine identified all 26 letters of a standard font.
March 1960Research · Robotics and autonomyLines: Learning theory
In March 1960 R. E. Kalman published A New Approach to Linear Filtering and Prediction Problems in the Journal of Basic Engineering (Transactions of the ASME, volume 82, issue 1, pages 35-45). The classical filtering and prediction problem is re-examined through the state-transition description of a dynamic system: a nonlinear equation is derived for the covariance of the optimal estimation error, and from its solution the coefficients of the optimal linear filter follow without further calculation; the method applies unchanged to stationary and nonstationary statistics.
On 6 May 1960 Science printed Norbert Wiener's 'Some Moral and Technical Consequences of Automation'. Using IBM's checkers programs, which beat the man who programmed them after 10 to 20 hours of play and training, Wiener argued that a machine that learns can develop unforeseen strategies faster than a person can understand them and intervene.
June 8, 1961Research · FoundationsLines: Symbolic AI · Learning theory
In a New York University report of 8 June 1961 Martin Davis, George Logemann and Donald Loveland described a program for the IBM 704 that tests formulas by the Davis-Putnam method of 1960, with one of its rules replaced by a splitting rule. A formula on which Gilmore's program had run for 21 minutes without a result, it proved in under two minutes.
In 1961 the IBM engineer William C. Dersch introduced Shoebox at a press conference: a machine the size of a shoebox that recognised 16 words, the digits zero through nine and six commands among them plus, minus, total and subtotal. Dersch said "Seven plus three plus six plus nine plus five. Subtotal", the machine typed out each digit and command and returned 30. The circuit needed 31 transistors, fewer than two per word, where earlier machines, by IBM's account, needed as many as 200 transistors per word.
1961Milestone · FoundationsLines: Embodiment · Work
Unimate, the first industrial robot, went to work in the car industry: by the museum's record at a General Motors plant in New Jersey, lifting hot castings out of a die-casting machine.
January 1962Research · ImagesLines: Learning theory · Science
In January 1962 Hubel and Wiesel divided the fields of 303 cells in the cat's visual cortex into simple and complex: a complex cell answered to a correctly oriented slit, edge or bar anywhere in its field. The cortex turned out to be built of columns of cells sharing one orientation.
In 1962 the Computing Centre of the Academy of Sciences of the Ukrainian SSR, created in 1957, became the Institute of Cybernetics of the Ukrainian Academy of Sciences: the institute's statute names the Academy Presidium's resolution of 11 May 1962. Its director was Victor Glushkov, who had also headed the centre.
1961 — November 1963Research · FoundationsLines: Learning theory
In November 1963 Donald Michie described in The Computer Journal MENACE, a machine that learned noughts and crosses by trial and error: one matchbox for each of 287 essentially distinct positions, with coloured beads in it for the possible moves. After a draw each open box got one extra bead, after a win three, and after a defeat one was taken away. Michie had described the machine and its first tournament in 1961.
Lawrence Roberts showed how to recover the three-dimensional shape of a polyhedron from a single photograph, going from intensities to edges to a model.
In 1964 the publishing house of the Ukrainian Academy of Sciences in Kyiv issued Victor Glushkov's Vvedeniye v kibernetiku (Introduction to Cybernetics), 324 pages. It brings together the theory of algorithms and automata, Boolean algebra, self-organising systems and perceptron training, digital computers and Algol-60, predicate calculus and the automation of proofs. Two English translations appeared in 1966.
In 1965 the Kyiv publisher Naukova Dumka issued the surgeon and cybernetician Nikolai Amosov's book on modelling thinking and the mind, in Russian, 302 pages; in 1967 it appeared in English as Modeling of Thinking and the Mind. The book describes the cortex of the brain as 'a powerful modeling apparatus' with temporary and permanent models.
An MIT memo assigned a group of students to build a system over one summer that separates an image into objects and background and describes the scene.
November 1966Milestone · LanguageLines: Money and markets · Law and policy
A National Academy of Sciences committee concluded that machine translation was slower, costlier and worse than human translation, and recommended ending its funding.
In 1966 Petr Hájek, Ivan Havel and Metoděj Chytil of Prague published the GUHA method (General Unary Hypotheses Automaton): a computer systematically generates every hypothesis of a given form about relations between properties of objects and outputs those that hold in the data. The Czech paper in Kybernetika was tried on a MINSK 2 computer with data on 230 epileptic patients, 36 properties each.
In volume 6 of Advances in Computers, which the publisher dates to 1966, Irving John Good of Trinity College, Oxford published 'Speculations Concerning the First Ultraintelligent Machine'. He defined an ultraintelligent machine as one that far surpasses any man in all intellectual activities and reasoned that it could design still better machines.
January 1968Research · LanguageLines: Speech · Learning theory
In January 1968 the Ukrainian journal Kibernetyka published T. K. Vintsyuk's "Speech discrimination by dynamic programming", volume 4, issue 1, pages 81-88; the English translation appeared as Cybernetics volume 4, pages 52-57. It applies dynamic programming to aligning two utterances in time, together with algorithms for recognising connected words.
July 1968Research · FoundationsLines: Search and reasoning · Learning theory
Three SRI researchers described a search combining the cost so far with an estimate of the rest, and proved that with an admissible estimate the path found is optimal.
November 1968Research · FoundationsLines: Symbolic AI · Tools
In November 1968 N. G. de Bruijn issued a report at the Technological University Eindhoven on Automath, a language for writing mathematics in enough detail that a computer can check whether the text is correct. Such a text need not be the proof of a single theorem; it can hold an entire theory together with its rules of inference.
In Kibernetika in 1968 (No. 2, pp. 81-88) M. I. Schlesinger described a self-learning algorithm for pattern recognition: each iteration first recognises the patterns, computing the a posteriori probabilities of the classes, and then solves the learning problem with those probabilities. The paper proves that the likelihood strictly increases, and that the limiting parameter values are maximum likelihood estimates.
1965 — 1969Milestone · FoundationsLines: Compute · Symbolic AI
At the Institute of Cybernetics of the Ukrainian Academy of Sciences, under Victor Glushkov, the engineering-calculation machines MIR-1 and MIR-2 were built in 1965-1969. MIR-1's input language had signs for integral, sum and product, and the length of numbers was limited only by memory; the language Analitik for analytic transformations of formulas was interpreted by the machine itself. MIR-2 went into production in 1969.
Bryson and Ho showed how to get the gradient of a cost with respect to the control at every stage of a multistage system: Lagrange multipliers, which they also call influence functions, are computed backward from the final stage.
March – April 1970Research · FoundationsLines: Symbolic AI · Learning theory
In Kibernetika in 1970 (No. 2, March-April) Victor Glushkov published a paper on problems of automata theory and artificial intelligence in which, according to his students' retrospectives, he formulated the 'evidence algorithm' programme: a formal language close to that of mathematical papers, proof search built on a machine notion of an evident step that grows with the system's experience, and a person helping the search.
Around 1970, by its developer's account, Peter Toma's SYSTRAN, developed with US Air Force sponsorship, became an operational Russian-to-English translation system, and it was still one in March 1976, when Toma presented it at a seminar of the Foreign Broadcast Information Service. By his account the system, on an IBM 360 or 370, processed 300,000 words per hour of CPU time.
March – April 1971Research · LanguageLines: Speech
In Kibernetika in 1971 (No. 2, March-April) T. K. Vintsyuk posed the problem of recognising connected speech composed of words from a given dictionary and proposed a solution: the division of the signal into words is chosen in the recognition process itself by dynamic programming, using a new computing scheme with a 'potential-optimal index'. The result of recognition is the word sequence.
1971Research · LanguageLines: Symbolic AI · Evaluation
In spring 1971 Kenneth Colby, Sylvia Weber and Franklin Hilf described PARRY, a program simulating a paranoid patient in a psychiatric interview over teletype. They proposed to judge the model by indistinguishability tests, and in 1972 psychiatrists acting as judges tried to tell interviews with it from interviews with patients.
1971Research · FoundationsLines: Search and reasoning · Symbolic AI
Fikes and Nilsson described a planner in which an action is given by preconditions and by lists of what it adds to and deletes from the state of the world.
January 1972Research · LanguageLines: Learning theory
In January 1972 Karen Spärck Jones of Cambridge proposed weighting query terms by how many documents in the collection contain them, so that a match on a rare term counts for more than one on a common term. On three test collections, Cranfield, INSPEC and Keen, the weighting gave a substantial gain over simple counting of matches.
April 1, 1972Benchmark · FoundationsLines: Science · Evaluation
On 1 April 1972 the BMJ published a comparison that F. T. de Dombal's group at the University of Leeds ran from 1 January to 1 December 1971 on 304 patients with acute abdominal pain. The diagnosis to which the computer gave the highest probability matched the final one in 279 cases (91.8%); that of the most senior clinician to see the patient in 242 (79.6%).
May 1972Research · FoundationsLines: Symbolic AI · Tools
In May 1972 Robin Milner described LCF in a Stanford Artificial Intelligence Project memo: a program that checks formal proofs in the logic of computable functions that Dana Scott had proposed at Oxford in the autumn of 1969. A person states a goal and splits it into subgoals; the machine carries the proof and checks each step.
August 20 – 23, 1973Research · FoundationsLines: Symbolic AI · Tools
In August 1973, at IJCAI in Stanford, Robert Boyer and J Strother Moore of the University of Edinburgh described a program that proves theorems about recursive LISP functions by mathematical induction on its own; among them, that REVERSE is its own inverse and that a particular sorting program is correct. On an ICL 4130 each theorem took 8 seconds on average.
1973Milestone · Robotics and autonomyLines: Embodiment
In 1973 Ambler, Barrow, Brown, Burstall and Popplestone of Edinburgh described an assembly system at IJCAI-73: a moving table, a hand with sensors and two TV cameras, joined through a Honeywell 316 to an ICL 4130 running POP-2. The operator dumps parts in a heap; the machine breaks the heap up, recognises the parts with an overhead camera, lays them out and assembles a peg and rings, a toy car or a toy ship by feel. Showing it half a dozen new parts takes about two hours and programming the assembly about four; the assembly itself takes an hour or two and succeeds about four times out of five.
1973Research · Robotics and autonomyLines: Embodiment
At IJCAI-73 Boris Dobrotin of the Jet Propulsion Laboratory and Victor Scheinman of the Stanford AI Project described the manipulator of JPL's robot research programme, built on the design of Stanford's Hand-Eye project: six degrees of freedom (two rotary, one linear, three rotary), a reach of up to 52 inches, a lift of about 5 pounds, a full motion in about 5 seconds, and six DC torque motors with harmonic drives on the four inner rotary joints. The arm is entirely under computer control: a person gives only gross-level commands.
Minsky proposed representing knowledge as stereotyped situations: a frame has slots, some filled with defaults that hold until something contradicts them.
October 1974Research · FoundationsLines: Symbolic AI · Science
The system advised on antibiotics for blood infections, held a dialogue with the physician, explained every step and attached a measure of confidence to its conclusion.
Paul Werbos presented a Harvard thesis showing how to compute all the derivatives of a complex model's error in one backward pass down an ordered table of operations. He called the method dynamic feedback.
Baker built a system in which every level of knowledge about language is one probabilistic model, and recognition is a search for the most likely path.
September 1975Research · Robotics and autonomyLines: Embodiment
In September 1975, at IJCAI-75 in Tbilisi, Nikolai Amosov, Ernst Kussul and V. D. Fomenko of the Department of Biocybernetics at the Institute of Cybernetics in Kiev described TAIR, a transport robot controlled not by a program but by a neuron-like network of 100 nodes in six spheres, with a system of reinforcement and inhibition and 60 input channels from its sensors. The undercarriage has three wheels and measures 1600x1100x600 mm.
November 1975Research · LanguageLines: Learning theory
In November 1975 Gerard Salton, A. Wong and C. S. Yang of Cornell described documents and queries as vectors of term weights and showed that retrieval works better where documents lie further apart in that space. On three collections, 424 documents in aerodynamics, 450 in medicine and 425 Time articles, they used this measure to choose an indexing vocabulary.
December 9, 1975Availability · LanguageLines: Symbolic AI · Tools
On 9 December 1975 METEO, from the TAUM group at the University of Montreal, began translating Canada's public weather forecasts from English into French on an experimental basis, with full operation planned for 15 May 1976. Sentences the system accepted went to radio stations and newspapers without human revision.
January 13, 1976Milestone · MultimodalLines: Speech
On 13 January 1976 Raymond Kurzweil and James Gashel of the National Federation of the Blind showed reporters a machine that reads print aloud: a scanner with a camera under glass, a minicomputer that recognises the letters and turns them into speech, and Braille-marked keys with which a blind reader directs it.
April 1976Benchmark · LanguageLines: Speech · Evaluation
Lowerre's system recognised 1011 words of continuous speech by compiling every level of knowledge into one large transition network and searching it with a beam of best paths.
July 1976Milestone · FoundationsLines: Science · Symbolic AI
In July 1976 Kenneth Appel and Wolfgang Haken of the University of Illinois submitted a proof that every planar map can be coloured with four colours. It reduced the problem to a finite set of configurations, each of which had to be checked for reducibility; the authors' programs did the checking on IBM computers.
July 1976Research · FoundationsLines: Symbolic AI · Search and reasoning
Lenat's program started from basic set-theoretic notions and, guided by heuristics of interestingness, derived concepts such as prime numbers by itself.
December 23, 1976Research · MultimodalLines: Speech · Science
On 23 December 1976 Nature printed a paper by Harry McGurk and John MacDonald of the University of Surrey. When the sound [ba] was dubbed onto film of a woman saying [ga], adults heard [da]; with the reverse dubbing most heard [bagba] or [gaba]. Sound alone or the untreated film gave the syllables correctly.
Weizenbaum separated what a machine is able to do from what it should be given to do, and argued that some decisions cannot be handed to anything but a person.
August 22 – 25, 1977Benchmark · FoundationsLines: Science · Evaluation · Symbolic AI
At IJCAI-77 in Cambridge (22-25 August 1977) S. M. Weiss, C. A. Kulikowski and A. Safir described CASNET, a program for the long-term management of glaucoma that reasons through a causal network of disease states. At the 1976 meeting of the American Academy of Ophthalmology and Otolaryngology 49 ophthalmologists rated its advice: 95% found it clinically acceptable, 77% expert or very competent.
August 1977Research · FoundationsLines: Compute · Symbolic AI · Tools
In August 1977 the group at MIT's Artificial Intelligence Laboratory reported an almost complete system running on the Lisp machine prototype. It is a personal computer whose instruction set, the memo says, was designed specifically for Lisp, the language of AI research; a complete machine would cost about $80,000.
Schank and Abelson argued that understanding a text means supplying what is unsaid: a reader knows the restaurant script and so needs no mention of a waiter.
April – December 1978Research · FoundationsLines: Compute · Learning theory
In April 1978 H. T. Kung and Charles E. Leiserson of Carnegie-Mellon described systolic arrays: a network of identical processors through which data flow rhythmically, as blood through the heart, each processor doing one inner product step per beat. A hexagonal grid of w1w2 processors multiplies two n x n band matrices in 3n+min(w1, w2) units of time.
1977 — 1978Research · FoundationsLines: Symbolic AI · Science
In 1977 the mathematician Wu Wen-Tsun of the Chinese Academy of Sciences implemented, on a Great Wall 203 computer with 4K of memory, a method that turns a theorem of elementary geometry into a system of polynomials and checks it algebraically, and proved the Simson line theorem. The paper on the method was published in Scientia Sinica in 1978.
February 1980Research · ImagesLines: Learning theory · Science
In February 1980 David Marr and Ellen Hildreth published a theory of edge detection. An image is convolved with a Laplacian of a Gaussian and an edge is marked where the result crosses zero. The filter is applied at several widths — the paper shows w = 6, 12 and 24 pixels — and zero crossings count as an edge when they coincide across channels of adjacent sizes.
June 1980Research · LanguageLines: Speech · Symbolic AI
The system was built as a shared blackboard where independent knowledge sources post hypotheses and read each other's, never addressing one another directly.
July 1980Research · MultimodalLines: Speech · Tools
In July 1980 Richard Bolt of MIT's Architecture Machine Group described at SIGGRAPH a system in which a person seated before a wall-sized screen creates and moves shapes by saying "put that there" and pointing. An NEC DP-100 with a vocabulary of up to 120 words recognised the speech, and a Polhemus magnetic sensor on a watchband gave the direction of the arm.
1980Availability · Robotics and autonomyLines: Defence · Embodiment
Production began in 1978 and in 1980 the system was first installed aboard the carrier USS Coral Sea. Phalanx searches, detects, evaluates, tracks, engages and assesses the kill on its own, without outside input.
In April 1981 Bruce Lucas and Takeo Kanade presented a way of aligning two images that does not test candidate displacements, but uses the spatial intensity gradient to correct the current estimate at each step. Exhaustive search over an M by M range of displacements on an N by N picture costs O(M²N²); this method converges in O(M² log N) steps on the average.
In June 1981 Martin Fischler and Robert Bolles published RANSAC. Rather than averaging every measurement, the method draws the smallest number of points a hypothesis needs, builds a model, and counts how much of the rest agrees with it. In an experiment on twenty landmarks of which five were gross errors, it found the correct solution on the second triple of points and admitted none of the errors.
August 1981Research · ImagesLines: Learning theory
Berthold Horn and Brian Schunck showed that velocity cannot be computed locally: at each point there is one measurement and two unknowns. They added a second condition — that the velocity field vary smoothly almost everywhere — and reduced motion estimation to minimising a single functional over the whole image. The paper appeared in August 1981.
September 10, 1981Research · ImagesLines: Learning theory
On 10 September 1981 Christopher Longuet-Higgins published in Nature an algorithm that recovers the three-dimensional structure of a scene and the relative placement of two viewpoints from point correspondences alone. Eight pairs of corresponding points give a system of linear equations, and that is enough — with nothing known about where the cameras stood.
August 19, 1982Benchmark · FoundationsLines: Evaluation · Science · Symbolic AI
On 19 August 1982 NEJM published a systematic evaluation of INTERNIST-I, a program for multiple diagnoses in internal medicine: on 19 of the journal's clinicopathological cases it performed qualitatively like the hospital clinicians but worse than the specialists who discussed the cases.
September 3, 1982Milestone · FoundationsLines: Symbolic AI · Science
In Science of 3 September 1982 Campbell, Hollister, Duda and Hart reported that the expert system PROSPECTOR, given geological maps drawn before drilling and rules from a porphyry molybdenum specialist, identified on Mount Tolman in Washington State the location of previously unknown ore-grade mineralization.
September 1982Benchmark · FoundationsLines: Science · Evaluation · Symbolic AI
In September 1982 Stanford issued a report on PUFF, a program that interprets lung function tests at Pacific Medical Center in San Francisco. It had been working there since 1979, gave interpretations for about ten patients a day and had handled over 4,000 cases; about 85% of its reports were accepted without change. On 144 new cases it agreed with the physician who had taught it in 96% and with an independent physician in 89%.
October 1982Research · FoundationsLines: Learning theory
In October 1982 Zdzisław Pawlak of the Institute of Computer Science of the Polish Academy of Sciences published 'Rough Sets', on approximate operations on sets, approximate equality and approximate inclusion. The author presents the approach as an alternative to fuzzy set theory and tolerance theory and as a mathematical foundation for artificial intelligence: classification, inductive reasoning, pattern recognition.
In 1982 David Marr's book Vision appeared posthumously: vision in it is a sequence of representations running from a description of the image to a description of the three-dimensional objects around, and each problem has to be treated on three levels, the computational, the algorithmic and that of hardware implementation.
April 28, 1983Law and regulation · FoundationsLines: Law and policy · Money and markets
On 28 April 1983 the UK Secretary of State for Industry, Patrick Jenkin, told Parliament that the Government accepted the Alvey Committee's recommendation: a five-year programme of collaborative research by industry, universities and government laboratories in four areas, among them intelligent knowledge-based systems. The report put the programme at about 350 million pounds, some 200 million of it public.
May 1983Research · FoundationsLines: Learning theory · Search and reasoning
Three IBM researchers showed that a hard optimisation problem can be solved by sometimes accepting a worse step and gradually lowering the chance of doing so.
July 1983Availability · FoundationsLines: Symbolic AI · Money and markets
In July 1983 General Electric's research centre delivered to the company's Locomotive Operation CATS-1, the first field prototype of DELTA, an expert system that troubleshoots diesel-electric locomotives in running repair shops and suggests repair procedures to the maintenance staff. About 530 rules (330 for troubleshooting, 200 for help) ran in FORTH on a PDP 11/23 minicomputer.
October 28, 1983Law and regulation · FoundationsLines: Law and policy · Defence · Compute
On 28 October 1983 the US Defense Advanced Research Projects Agency (DARPA) set out its plan for a Strategic Computing programme: to create 'machine intelligence' within a decade and demonstrate it in three military applications — an autonomous vehicle, a pilot's associate and a battle management system for a carrier battle group. The first five years were estimated at approximately 600 million dollars.
In 1983 Digital Equipment Corporation announced DECtalk, a device that turns English text into speech by the rules of Dennis Klatt's laboratory system Klattalk at MIT, licensed in 1982. A digital formant synthesizer builds the speech, and any computer or terminal can drive it.
December 1983Research · Robotics and autonomyLines: Embodiment
In a Carnegie Mellon Leg Laboratory report dated 13 December 1983, Marc Raibert, Benjamin Brown and Michael Chepponis described a machine on one springy leg that hops and runs on an open floor without any physical support. Control is split into three independent parts: hopping height, forward running velocity and body attitude. The machine hops in place, travels at a specified rate, follows simple paths and keeps its balance when disturbed; the top recorded running speed was 2.2 m/s (4.8 mph).
February 28, 1984Law and regulation · FoundationsLines: Law and policy · Money and markets
On 28 February 1984 the Council of the European Communities adopted ESPRIT, a five-year programme of research in information technologies from 1 January 1984 with a Community contribution of 750 million ECU. One of its five areas is advanced information processing: expert systems, knowledge representation, inference and learning, pattern recognition.
1981 — March 1984Research · FoundationsLines: Compute
Igor Aleksander, Thomas Stonham and Bruce Wilkie at Brunel University built WISARD, a pattern recogniser whose neurons were 32,000 random access memory chips in a single-layer net. By Aleksander in 1983 the machine took 5-10 decisions a second on 512 x 512 pixel images; by March 1984 it had been turned into an industrial product.
November 1984Research · FoundationsLines: Learning theory
Valiant defined it: a class of concepts is learnable if an algorithm exists that, in polynomial time and from a reasonable number of examples, returns a nearly correct answer with high probability.
March 1985Research · Robotics and autonomyLines: Learning theory · Embodiment
At ICRA 1985 in St. Louis (25-28 March 1985) Oussama Khatib of Stanford showed that collision avoidance, until then treated as a high-level planning problem, can be distributed between levels of control: an obstacle raises an artificial potential field that repels the robot inside the control loop itself. Manipulator control is reformulated as direct control of motion in the space of the task rather than in joint space. The method was implemented in the COSMOS system on a PUMA 560 robot and demonstrated in real time on moving obstacles with visual sensing.
April 1985Research · FoundationsLines: Learning theory
Pearl proposed representing dependencies between events as a graph, and showed how to propagate new evidence through it locally, without recomputing everything.
February 1986Milestone · FoundationsLines: Compute · Autonomous driving
In February 1986 General Electric delivered to Carnegie Mellon the first ten-cell prototype of Warp, a systolic array of linearly connected programmable cells, each capable of 10 million floating-point operations a second, 100 MFLOPS in all. The machine was built for the vision of robot vehicles and rode in the Navlab van.
May 1986Research · FoundationsLines: Symbolic AI · Learning theory
In May 1986 Thierry Coquand and Gérard Huet of INRIA issued a report on the calculus of constructions, a higher-order formalism in which every proof is a lambda expression typed with the proposition it proves. Remove the types and what remains is the program corresponding to the proof. The authors proved strong normalisation: every computation terminates, so the logic is consistent.
May 1985 – October 1986Milestone · Robotics and autonomyLines: Autonomous driving · Defence
In May 1985 the Autonomous Land Vehicle, which Martin Marietta was building for DARPA's Strategic Computing programme, drove a road by itself in public for the first time, about 1 km. In June 1986 it covered the whole test track, 4.2 km, at up to 10 km/h, and in October 1986 it first steered around obstacles while staying on the road and reached 20 km/h on a straight road without obstacles.
November 1986Research · ImagesLines: Learning theory
In November 1986 John Canny wrote edge detection down as the optimisation of three quantities: signal-to-noise ratio, accuracy of the located response, and having only one response to a single edge. Solving it numerically, he showed there is an uncertainty principle between detection and localisation, and that the optimal operator is well approximated by the first derivative of a Gaussian.
June 1985 — 1986Research · FoundationsLines: Compute · Search and reasoning
In June 1985 Carnegie Mellon graduate student Feng-hsiung Hsu set out to build a chess move generator on a single chip in the 3-micron technology available to universities. The first working chips arrived about ten months later; the chip produced up to 2 million moves a second, ten times the 64-chip module of the Hitech machine.
December 1986Research · Robotics and autonomyLines: Learning theory
In the Winter 1986 issue of The International Journal of Robotics Research (volume 5, issue 4, pages 56-68) Randall Smith of SRI and Peter Cheeseman of NASA Ames gave a way to estimate the nominal relationship and expected error (covariance) between coordinate frames known only through a chain of uncertain relations. Two operations: compounding collapses a chain of uncertain transformations into one, merging combines parallel estimates into one with less uncertainty. The example is a mobile robot in three degrees of freedom (x, y, theta); the method generalises to six, and the estimates agree with an independent Monte Carlo simulation.
Rumelhart and McClelland's two volumes gathered connectionism into a coherent programme: representation is distributed, and knowledge lives in the weights rather than in symbols.
August 1986 — 1987Availability · FoundationsLines: Defence · Symbolic AI
In August 1986 Texas Instruments and btg delivered to the headquarters of the Commander-in-Chief, US Pacific Fleet, at Pearl Harbor the first prototype of FRESH, an expert system that tracks the status and employment of ships and helps plan their use. DARPA's budget of February 1987 says the prototype is installed and demonstrated at the headquarters along with a natural language interface.
1987Availability · FoundationsLines: Data · Evaluation
In 1987 David Aha, a doctoral student at the University of California, Irvine, opened an FTP archive of data sets for testing machine learning algorithms. By 4 December 1995 it held 111 databases and domain theories, 36 megabytes in all.
1987Milestone · Robotics and autonomyLines: Autonomous driving
In 1987 VaMoRs, a 5-ton Mercedes van that Ernst Dickmanns's group at the University of the Federal Armed Forces in Munich had fitted with cameras and processors, made autonomous runs on a free Autobahn at up to 96 km/h, with both steering and speed under computer vision. A year earlier the same vehicle had shown longitudinal and lateral guidance by vision on Daimler-Benz's skidpan in Stuttgart at up to 10 m/s (36 km/h). Dickmanns and Zapp reported the results at the 10th IFAC World Congress in Munich in 1987.
1987Milestone · FoundationsLines: Money and markets · Symbolic AI
The market for specialised AI hardware began to vanish: general-purpose workstations and personal computers came close to Lisp machines in performance at a far lower price.
February 9 – 12, 1988Research · LanguageLines: Data · Tools
At the ANLP conference in Austin in February 1988 Kenneth Church of Bell Laboratories presented a program that assigns parts of speech in unrestricted text by multiplying two probabilities estimated on the tagged Brown Corpus of about a million words: how often a word takes a part of speech, and how often that part follows the two before it. It was 95-99% correct, depending on the definition of correct.
May 1988Research · MultimodalLines: Speech · Evaluation
In May 1988 Eric Petajan and colleagues at AT&T Bell Laboratories, with Michael Brooke of the University of Bath, showed at CHI '88 a system that lipreads from video of the mouth and combines the result with an AT&T ASR1000 acoustic recogniser. On the English alphabet the combination cut its errors by a quarter to two thirds.
May 1988Research · Robotics and autonomyLines: Autonomous driving
In May 1988 Thorpe, Hebert, Kanade and Shafer of Carnegie Mellon described in IEEE Transactions on Pattern Analysis and Machine Intelligence (volume 10, issue 3, pages 362-373) the perception and navigation system of the Navlab van: a distributed architecture around the CODGER database, colour vision for road following and 3-D vision for detecting and avoiding obstacles; the vehicle drives continuously on roads in a real outdoor environment while avoiding obstacles. The CMU year-end report for 1987 lists among its milestones a Navlab demonstration on the Warp processor at 50 cm/s.
March – June 1988Milestone · FoundationsLines: Defence · Symbolic AI
In March 1988 the Lockheed and US Air Force team, and in June the McDonnell Aircraft team, gave the second demonstration of the Pilot's Associate, a DARPA and Air Force programme begun in February 1986 for Strategic Computing: expert systems assisting a pilot in a simulated cockpit. At McDonnell fully integrated software supported only one of six mission segments: a minute of flight took about two hours of processing.
March 1987 – November 1988Availability · FoundationsLines: Money and markets · Symbolic AI
In November 1988, according to its authors' paper at IAAI-89, American Express put the Authorizer's Assistant into full operation in all US centres: an expert system of 890 rules in ART from Inference Corporation prepares card authorization decisions for authorizers and approves millions of dollars of credit daily without human intervention. The pilot ran at one centre from March-April 1987.
1988Research · Robotics and autonomyLines: Autonomous driving
At the NIPS conference of 1988 Dean Pomerleau of Carnegie Mellon described ALVINN, a three-layer back-propagation network that takes a 30 by 32 camera image and an 8 by 32 laser range-finder image (1217 inputs in all) and, through 29 hidden units, outputs one of 45 travel directions. The network was trained on 1200 simulated road images for 40 epochs; on new simulated images it picks the curvature within two units about 90 percent of the time. On the Navlab van it drove the vehicle along a 400 metre path through a wooded part of the campus at half a metre per second.
January 1, 1989Milestone · FoundationsLines: Symbolic AI · Tools
On 1 January 1989 the first three articles were entered into the database of Andrzej Trybulec's Mizar project in Białystok, and the participants call this the official start of the Mizar Mathematical Library, a collection of mathematical texts every line of which a machine checks. By the end of 1989 there were 66 articles; from January 1990 they were printed in the journal Formalized Mathematics.
March 1987 – February 1989Benchmark · LanguageLines: Speech · Data · Evaluation
From March 1987 NIST, for DARPA, ran tests of speech recognition systems on a shared Resource Management database: read sentences for a naval resource management task, a vocabulary of about a thousand words, over 21,000 recordings from 160 talkers. The database was described at ICASSP-88 in April 1988; by February 1989 four tests had been held.
February 1989Research · Robotics and autonomyLines: Embodiment
In February 1989 Rodney Brooks described in MIT AI Lab memo 1091 a six-legged machine about 35 cm long and roughly 1 kg in mass, controlled by a network of 57 augmented finite state machines built incrementally: each network is a strict augmentation of the previous one, and at every step the robot remains a working system. The robot walks over rough terrain and follows a person it senses passively with six pyroelectric infrared sensors. The paper appeared in Neural Computation, volume 1, issue 2, in June 1989.
September 1990Research · LanguageLines: Learning theory · Data
In September 1990 Scott Deerwester, Susan Dumais, George Furnas, Thomas Landauer and Richard Harshman described in JASIS document retrieval through the singular value decomposition of a term-by-document matrix: terms and documents get vectors of about a hundred factors, and a query is compared with them by cosine. On the MED medical collection precision rose 13% over plain word matching.
October 1990Research · FoundationsLines: Compute · Learning theory
In October 1990 Carver Mead of the California Institute of Technology called analog chips built on the organising principles of the nervous system neuromorphic systems. His example was the silicon retina he had described with Misha Mahowald: about 100,000 devices doing 100 million operations a second on 1 mW, about 10^-11 J an operation against 10^-7 J in a digital design.
In 1989-1990 two chips appeared that were made for neural networks. Intel's analog ETANN had 64 neurons and 10,240 synapses, its weights held as charge on floating gates, and computed more than 1.3 billion connections a second. Adaptive Solutions' digital CNAPS had 64 processors at 25 MHz and learned on the chip itself.
December 1990Research · LanguageLines: Symbolic AI · Data
In December 1990 George Miller's Princeton group published WordNet, a database of English vocabulary built around meanings rather than the alphabet. Its unit is the synset, a set of words expressing one concept; the nouns are split into 25 independent hierarchies.
On 17 January 1991 the destroyer USS Paul F. Foster fired a salvo of Tomahawks in the opening barrage of the war, the missile's first operational use. On the final leg the missile was guided by DSMAC, which compared what its camera saw against stored reference scenes of the ground.
January 1991Research · ImagesLines: Learning theory
In January 1991 Matthew Turk and Alex Pentland described a system that locates a head in the frame and recognises the person in near real time. A face is projected into the space of the training set's eigenvectors — the "eigenfaces" — and recognition reduces to comparing a handful of coefficients rather than recovering the three-dimensional shape of a nose or an eye.
February 1991Research · FoundationsLines: Learning theory
In the February 1991 issue of Neural Computation (volume 3, issue 1, pages 79-87) Robert Jacobs and Michael Jordan of MIT and Steven Nowlan and Geoffrey Hinton of the University of Toronto described a system of several expert networks and a gating network that decides, for each training case, which expert takes it. On a task of telling apart four vowels from 75 speakers, mixtures of 4 or 8 simple linear experts reached the error criterion in about half the epochs that backpropagation networks needed, at the same 90 per cent on test.
November 8, 1991Benchmark · LanguageLines: Evaluation
On 8 November 1991 Boston's Computer Museum held the first contest for Hugh Loebner's prize, a restricted Turing test: ten judges from the public conversed on set topics with six programs and two people. Joseph Weintraub's program on whimsical conversation won; five judges rated it human.
November 1991Research · Robotics and autonomyLines: Learning theory · Embodiment
At the IROS '91 workshop in Osaka (3-5 November 1991) John Leonard of the NEC Research Institute and Hugh Durrant-Whyte of Oxford named an open problem of mobile robotics, simultaneous map building and localization, and defined it as long-term globally referenced position estimation without a priori information. The problem is hard because of a paradox: to move precisely a robot needs an accurate map, and to build an accurate map it must know precisely where it sensed from. For sonar they fit the vehicle with several servo-mounted sensors, learn a subset of environment features from the initial location and then track them.
December 1, 1991Benchmark · FoundationsLines: Evaluation · Science
On 1 December 1991 William Baxt published a prospective, blinded test of a neural network trained on 351 hospitalised patients: on 331 emergency patients with chest pain it recognised myocardial infarction with 97.2% sensitivity and 96.2% specificity, against 77.7% and 84.7% for the physicians.
July – December 1991Research · FoundationsLines: Compute
In 1991 Bernhard Boser, Yann LeCun and colleagues at AT&T Bell Labs described ANNA, a chip that performs over 2,000 multiplications and additions at once. It ran a convolutional network for handwritten digits with 133,000 connections: more than 1,000 characters a second at 5.3% error against 4.9% for the original floating-point network.
1991Milestone · FoundationsLines: Defence · Money and markets
A transport planner built in ten weeks during Operation Desert Shield was in use at US Transportation Command and US European Command through the winter of 1990–91. It reworked plans for moving troops and cargo at the speed at which those plans were changing.
March 31 – April 3, 1992Research · LanguageLines: Symbolic AI · Data
At the third ANLP conference in Trento (31 March to 3 April 1992) Eric Brill of the University of Pennsylvania presented a part-of-speech tagger that finds its own rules for correcting its errors. The 71 rules it found brought the error rate on a held-out 5% of the Brown Corpus down to 5.1%.
May 27 – 28, 1992Benchmark · ImagesLines: Data · Evaluation
On 27-28 May 1992 NIST, for the US Census Bureau, convened in Gaithersburg the participants of the first Census optical character recognition systems conference: 26 organisations from North America and Europe recognised the same roughly 85,000 handwritten digits and letters, unseen before. About half the systems recognised over 95 percent of the digits; a human, about 98.5.
July 1992Research · FoundationsLines: Learning theory
The authors showed that a maximum-margin linear classifier can be applied to nonlinear problems without ever constructing the high-dimensional space explicitly.
September 1992Availability · FoundationsLines: Money and markets
In September 1992 HNC of San Diego introduced Falcon, a system that scored the probability of fraud with a neural network while each credit card transaction was being authorised. According to the company's annual report, by August 1994 Falcon had been bought by 18 of the 20 largest US credit card issuers and monitored over 100 million accounts.
November 4 – 6, 1992Benchmark · LanguageLines: Evaluation · Data
On 4-6 November 1992 in Gaithersburg, NIST and DARPA held the first Text REtrieval Conference: 25 groups ran their systems on a shared collection of about two gigabytes of text and shared topics, and the results were scored by one procedure. 92 people attended.
November 1992Research · ImagesLines: Learning theory
In November 1992 Carlo Tomasi and Takeo Kanade showed that if the coordinates of P points tracked through F frames are collected into one 2F by P matrix, then under orthographic projection and without noise that matrix has rank exactly three. A singular value decomposition splits it into camera motion and scene shape — one computation over the whole sequence, with no error accumulating frame to frame.
December 1992Research · LanguageLines: Learning theory · Data
In December 1992 five IBM researchers led by Peter Brown described in Computational Linguistics how to divide a vocabulary of 260,741 words into 1,000 classes purely from which words stand next to each other in almost 366 million words of text, and how to build a language model on classes instead of single words.
June 1993Milestone · LanguageLines: Data · Evaluation
The University of Pennsylvania annotated millions of words of English for part of speech and syntactic structure, and released the corpus to the community.
On 22 October 1994, at the CSCW conference, Paul Resnick of MIT, John Riedl of the University of Minnesota and three co-authors described GroupLens, an open architecture for collaborative filtering of Usenet news. Readers rate articles; rating servers called Better Bit Bureaus gather the ratings and predict how another reader will rate an article, on the heuristic that people who agreed in the past will agree again.
1994Research · LanguageLines: Learning theory · Evaluation
In 1994 Stephen Robertson's group at City University, London, preparing its Okapi system for TREC-3, combined its two weighting functions BM11 and BM15 into one, BM25, in which a parameter b sets how far document length changes the contribution of term frequency. Run citya1 with this function headed the six best automatic ad hoc runs.
1994Milestone · Robotics and autonomyLines: Autonomous driving
In 1994, at the final demonstration of the European EUREKA project PROMETHEUS near Paris, the VaMP passenger car of the University of the Federal Armed Forces in Munich and Daimler-Benz's VITA-2 drove the A1 motorway in normally dense three-lane traffic at up to 130 km/h: free lane driving, convoy driving and lane changes for passing, with the decision whether a lane change was safe taken by the vehicle itself and a human safety pilot merely giving the go-ahead. VaMP watched the road with two camera pairs, front and rear, and tracked up to ten other vehicles out to about 100 m.
February 1995Research · MultimodalLines: Evaluation
Thad Starner's master's thesis at the MIT Media Lab (February 1995, supervised by Alex Pentland) described a hidden Markov model system that recognises sentences of American Sign Language in real time from one camera's video, with a forty-word lexicon. In November, at ISCV in Coral Gables, they reported 99.2% word accuracy.
March 1995Research · FoundationsLines: Learning theory
Freund and Schapire showed how to combine classifiers barely better than chance into one strong classifier, raising the weight of mistakes at each step.
May 9 – 12, 1995Research · LanguageLines: Learning theory · Speech
At ICASSP-95 in Detroit (9-12 May 1995) Reinhard Kneser of Philips' research laboratories in Aachen and Hermann Ney of Aachen proposed backing-off distributions for n-gram language models optimised for exactly the cases where the longer sequence had not been seen. Perplexity fell by about 10%, and the word error rate in speech recognition by 5%.
In September 1995 Myron Flickner and eleven other researchers at IBM Almaden described the QBIC system in IEEE Computer. An image and video database could be queried by example, by sketch, or by a chosen colour or texture rather than by words. Answers were ranked by similarity of colour, texture, shape and motion computed from the images themselves.
November 6 – 8, 1995Benchmark · LanguageLines: Evaluation · Data
On 6-8 November 1995 in Columbia, Maryland, the Sixth Message Understanding Conference, funded by DARPA, reviewed an evaluation that for the first time measured named entity recognition separately: names of people, organisations and places, dates, sums of money and percentages. The best of 20 systems from fifteen sites, SRA's, had an F-measure of 96.42; two human annotators scored against each other had 96.68.
In November 1995 Henry Rowley, Shumeet Baluja and Takeo Kanade of Carnegie Mellon described a face detector in a technical report: a neural network examines every 20 × 20 pixel window at several scales, and arbitration between two networks weeds out false detections. On 130 complex images with 507 faces the main system found 85.4% of them, with one false detection per 1,319,035 windows.
November 1995Milestone · Robotics and autonomyLines: Autonomous driving
In November 1995 VaMP drove more than 1,600 km from Munich to Odense for a project meeting and back with autonomous longitudinal and lateral guidance on black-and-white images: by the group's estimate about 95 percent of the distance was covered fully automatically. On a free stretch of Autobahn in the Lüneburg Heath the speed reached 180 km/h; five times the vehicle ran more than 100 km without any intervention, the longest stretch being about 160 km; more than 400 lane changes were performed autonomously on the safety driver's command; lateral deviation from the lane centre rarely exceeded 0.2 m.
March 1996Research · LanguageLines: Learning theory
In March 1996 Adam Berger and Stephen and Vincent Della Pietra described in Computational Linguistics how to build a probabilistic model of language on the principle of maximum entropy: of all the distributions consistent with chosen facts about the data, take the most uniform, and select the facts (features) automatically.
May 7 – 10, 1996Research · MultimodalLines: Speech · Generative media · Data
At ICASSP-96 in Atlanta (7-10 May 1996) Andrew Hunt and Alan Black of the ATR laboratories in Kyoto described how the CHATR synthesiser builds speech from phonemes cut out of a large database of one speaker's recordings. Each phoneme is chosen by a Viterbi search on two costs: how close the unit is to the target and how smoothly it joins its neighbour.
June 1996Milestone · ImagesLines: Money and markets
According to Bottou, Bengio and LeCun of AT&T Laboratories at CVPR-97, their system for reading the amount on cheques was built into NCR's cheque processing machines, first fielded in a bank in June 1996, and had since been reading millions of cheques a month. The amount was recognised by LeNet5, a convolutional network trained together with the whole system.
August 1996Research · Robotics and autonomyLines: Learning theory
In August 1996 Kavraki and Latombe of Stanford and Švestka and Overmars of Utrecht published in IEEE Transactions on Robotics and Automation (volume 12, issue 4, pages 566-580) a two-phase motion planning method: in a learning phase, randomly sampled collision-free configurations of the robot are joined by a simple, fast local planner into a graph, the roadmap; in a query phase the start and goal configurations are connected to the graph and a path is searched in it. For planar articulated robots with many degrees of freedom, planning takes a fraction of a second on a workstation of about 150 MIPS after a few dozen seconds of learning.
September 1996Benchmark · ImagesLines: Evaluation · Data · Defence
In September 1996 the third and final FERET evaluation was run. The protocol required matching a set of 3323 images against a set of 3816 — approximately 12.6 million matches — over a database of 14,126 images of 1199 individuals, gathered independently of the algorithm developers. Results were reported separately for images taken on the same day, on different days, and over a year apart.
October 10, 1996Milestone · FoundationsLines: Symbolic AI · Science
On 10 October 1996 the program EQP, built at Argonne National Laboratory, found a proof that every Robbins algebra is Boolean. The search took about eight days on an RS/6000 processor and used about 30 megabytes of memory. Herbert Robbins had posed the question shortly after 1933; neither he nor Huntington found a proof, and Tarski and his students later studied it.
December 20, 1996Milestone · Robotics and autonomyLines: Embodiment
On 20 December 1996 Honda disclosed ten years of closed research and showed for the first time a prototype of a self-standing two-legged humanoid robot with its battery on board: two legs, two arms, 180 cm tall, 210 kg, an operating time of about 15 minutes. The robot judges the surface itself and walks on flat floor, over steps, on slopes and up and down stairs, does not fall when pushed, changes direction on a wireless command, and with its two hands pushes a cart or tightens bolts under remote operation.
In the autumn of 1996 Erling Wold, Thom Blum, Douglas Keislar and James Wheaton of Muscle Fish in Berkeley described in IEEE MultiMedia a system that reduces any sound to loudness, pitch, brightness, bandwidth and harmonicity. With these features it retrieves similar sounds by example and assigns new ones to classes trained on examples, on a database of about 400 files.
July 7 – 12, 1997Research · LanguageLines: Learning theory · Data
At the joint ACL and EACL conference in Madrid (7-12 July 1997) Michael Collins of the University of Pennsylvania described three generative models of lexicalised context-free grammar. On section 23 of the Wall Street Journal in the Penn Treebank the best reached 88.1% constituent precision and 87.5% recall.
December 9, 1997Availability · LanguageLines: Tools
On 9 December 1997 Digital Equipment Corporation opened a free trial translation service on the AltaVista search engine, running on SYSTRAN: anyone could translate a web page, search results or pasted text from English into Spanish, French, German, Portuguese and Italian, and back.
June 26, 1998Availability · ImagesLines: Law and policy · Science
On 26 June 1998 the FDA approved R2 Technology's ImageChecker M1000, a system that marks suspicious regions on a screening mammogram for the radiologist to review after the initial reading.
July 1998Research · Robotics and autonomyLines: Embodiment
At AAAI-98 in Madison (26-30 July 1998) Cynthia Breazeal of the MIT AI Laboratory described the motivational system of the autonomous robot Kismet, which regulates interaction with a person on the model of a caretaker-infant pair: drives, emotions and facial expressions together keep the exchange at an intensity the robot can handle. The robot has active stereo vision with colour cameras in its eyeballs and a microphone on each ear; its moving eyebrows, ears, eyeballs and eyelids show anger, fatigue, fear, disgust, excitement, happiness, interest, sadness and surprise. Early human-robot interaction experiments are reported.
August 1998Research · LanguageLines: Evaluation · Learning theory
In August 1998 Stanley Chen and Joshua Goodman issued a Harvard technical report comparing the main smoothing methods for n-gram language models on the Brown, North American Business news, Switchboard and Broadcast News corpora across training set sizes. Their own modified variant of Kneser-Ney won every comparison.
In 1998 Mehran Sahami of Stanford and Susan Dumais, David Heckerman and Eric Horvitz of Microsoft Research showed that a naive Bayesian classifier trained on a user's mail filters junk. On 1,789 messages, with simple mail-specific features, it reached 100 per cent precision and 98.3 per cent recall on junk.
February 22, 1999Milestone · FoundationsLines: Symbolic AI · Data
On 22 February 1999 the W3C made the RDF specification a Recommendation. It set one way to write a statement about any web resource: subject, predicate, object, with subject and predicate named by a URI.
February 1999Research · MultimodalLines: Speech · Data
The Informedia Digital Video Library at Carnegie Mellon, begun in 1994 under the Digital Library Initiative (NSF, DARPA, NASA), transcribed the sound of video with the Sphinx recogniser and found stories by the words of the transcript. By May 1998 it held over 1,000 hours of news and 400 hours of documentary video; the account appeared in IEEE Computer in February 1999.
September 5 – 9, 1999Research · MultimodalLines: Speech · Generative media · Openness
At Eurospeech '99 in Budapest (5-9 September 1999) Takayoshi Yoshimura, Keiichi Tokuda and colleagues from the Nagoya and Tokyo Institutes of Technology described a synthesiser in which a hidden Markov model itself generates the spectrum, the pitch and the durations of sounds. On 25 December 2002 their group released the system openly as HTS 1.0.
September 20 – 27, 1999Research · ImagesLines: Generative media
In September 1999 at ICCV Alexei Efros and Thomas Leung of Berkeley showed texture synthesis without a model: a new image grows from a seed one pixel at a time, each pixel taken from a place in the sample whose neighbourhood resembles what has been made so far. There is one parameter, the window size.
September 1999Research · ImagesLines: Learning theory
In September 1999 David Lowe presented SIFT: stable points are found as extrema in a difference-of-Gaussians pyramid, and the neighbourhood of each is described by a vector of gradient orientation histograms. Over twenty images and about 15,000 keys, 85.4 percent of keys were still found after a 20 degree rotation, 85.1 percent after a scaling by 0.7, and 90.3 percent after adding 10 percent pixel noise.
October 11, 1999Availability · FoundationsLines: Compute
NVIDIA released a chip it called the world’s first graphics processing unit: transform, lighting and triangle rendering on a single chip for the first time.
April 12, 2000Benchmark · MultimodalLines: Speech · Evaluation · Data
In 1997-1999 NIST ran a spoken document retrieval track at TREC: systems recognised news broadcasts and searched the transcripts for queries. The collection grew from 50 hours in 1997 to 557 hours and 21,754 stories in 1999. NIST presented the summary at RIAO 2000 in Paris on 12-14 April 2000.
July 23 – 28, 2000Research · ImagesLines: Generative media
In July 2000 at SIGGRAPH Marcelo Bertalmío and Guillermo Sapiro (Minnesota) with Vicent Caselles and Coloma Ballester (Pompeu Fabra) described digital inpainting: a person roughly marks the damaged area, and the program continues inward the lines of equal brightness that reach its edge.
September 13 – 14, 2000Benchmark · LanguageLines: Evaluation · Data
On 13 and 14 September 2000 in Lisbon, eleven research systems cut the same text into groups of words, over the same 211,727 training and 47,377 test tokens, and were scored by one shared script. The best reached an F of 93.48 against 77.07 for a simple baseline.
January 21 – 28, 2001Milestone · ImagesLines: Harm · Law and policy
From 21 to 28 January 2001 in Tampa, at the stadium before Super Bowl XXXV and in the Ybor City district, cameras captured people’s faces and compared them with a database of criminals. The FaceTrac system, built on Viisage FaceFINDER, was supplied to the police by Graphco, Viisage, Raytheon and VelTek. 100,000 fans and workers passed through the turnstiles; according to the AP, 19 faces of petty offenders were flagged, and police said there were no arrests.
June 2001Research · LanguageLines: Learning theory
At ICML 2001 John Lafferty (Carnegie Mellon), Andrew McCallum and Fernando Pereira (WhizBang! Labs, Pereira also the University of Pennsylvania) presented conditional random fields: a model of the probability of a whole label sequence given an observation sequence, normalised over the sequence rather than at each state. On part-of-speech tagging of the Penn treebank the per-word error was 5.55 per cent against 5.69 for a hidden Markov model and 6.37 for a maximum entropy Markov model; with spelling features, 4.27 against 4.81.
August 16, 2001Research · ImagesLines: Generative media
On 16 August 2001 at SIGGRAPH Aaron Hertzmann and co-authors from NYU, Microsoft Research and the University of Washington showed how to learn a filter from a single example: a pair of images A and A’ defines a transformation, and the program applies it to a new B to produce B’.
October 2001Research · FoundationsLines: Learning theory
In The Annals of Statistics for October 2001 (volume 29, issue 5, pages 1189-1232) Jerome Friedman of Stanford presented boosting as gradient descent in function space: each new component of an additive model is fitted to the negative gradient of any chosen loss. He gave algorithms for least squares, least absolute deviation, Huber loss and multiclass logistic likelihood, special versions for regression trees (TreeBoost) and tools for interpreting such models. The paper is the 1999 Reitz Lecture.
October 2001Research · FoundationsLines: Learning theory
In Machine Learning (volume 45, issue 1, October 2001, pages 5-32) Leo Breiman of the University of California, Berkeley, defined random forests: an ensemble of trees, each grown from an independently drawn random vector, voting for the class. He proved that a forest's generalisation error converges almost surely to a limit as trees are added, bounded it by the strength of the individual trees and the correlation between them, and showed that choosing features at random at each split gives error rates that compare favourably with AdaBoost while being more robust to noise.
November 13 – 16, 2001Benchmark · MultimodalLines: Evaluation · Data
At TREC 2001 in Gaithersburg (13-16 November 2001) NIST ran its first video track: 12 groups from the US, Asia and Europe searched 11 hours of MPEG-1 video for 74 topics, each carrying a clip, image or sound besides its text, and detected shot boundaries. From 2003 the track became TRECVID, a separate annual evaluation.
April 29, 2002Research · MultimodalLines: Data · Learning theory
In the ECCV 2002 volume, published on 29 April 2002, Pinar Duygulu, Kobus Barnard, Nando de Freitas and David Forsyth of Berkeley and UBC cast recognition as translation: image regions are the words of one language, keywords those of another. On 4,500 Corel images with 371 words, EM learned a lexicon between 500 region types and the words.
July 2002Benchmark · MultimodalLines: Data · Evaluation
In July 2002 George Tzanetakis and Perry Cook of Princeton University published in IEEE Transactions on Speech and Audio Processing a system that assigns a music recording to a genre from features of timbre, rhythm and pitch. On a set of ten genres with a hundred 30-second excerpts each it was right 61% of the time, where chance would give 10%.
July 2002Research · LanguageLines: Evaluation · Tools
In July 2002 four IBM researchers showed a way to score machine translation without a person: count matches of runs of one to four words against human references, with a penalty for an answer that is too short. Across five systems the score correlated with human judges at 0.99.
In August 2002 Paul Graham published the essay 'A Plan for Spam': a filter combining, by Bayes' rule, the spam probabilities of a message's fifteen most telling words let through fewer than 5 spams in 1,000 on his own mail, with no false positives.
January 2003Research · FoundationsLines: Money and markets · Data
In the January-February 2003 issue of IEEE Internet Computing, Greg Linden, Brent Smith and Jeremy York of Amazon.com described item-to-item collaborative filtering, with which the store personalises its pages for each customer. The table of similar items is computed offline, and the online part depends only on how many items the customer has bought or rated, not on the size of the customer base or the catalogue.
February 2003Research · MultimodalLines: Data · Learning theory
In February 2003 the Journal of Machine Learning Research printed "Matching Words and Pictures" by Barnard, Duygulu, Forsyth, de Freitas, Blei and Jordan. Several models of the joint distribution of words and image regions predicted captions for whole images and named single regions; the data were images from 160 Corel CDs of 100 each.
July 2002 – March 2003Benchmark · ImagesLines: Evaluation
In March 2003 NIST published the FRVT 2002 report. In July and August 2002 ten commercial face recognition systems had been tested under the organisers’ supervision on 121,589 photographs of 37,437 people from the US State Department’s visa archive. On comparable tests the error had fallen by about half since FRVT 2000.
In July 2003 at SIGGRAPH Patrick Pérez, Michel Gangnet and Andrew Blake of Microsoft Research in Cambridge showed how to paste part of one image into another with no visible seam: solve a Poisson equation that takes brightness changes from the source and values from the boundary.
In October 2003 at ICCV Josef Sivic and Andrew Zisserman of Oxford carried the machinery of text retrieval over to images. Descriptors of regions in a frame were clustered into "visual words", the commonest and rarest were dropped by a stop list, frames were weighted by tf-idf and entered in an inverted file. Finding an outlined object among a film’s 4,000 keyframes took about 0.1 seconds.
October 26 – 30, 2003Research · MultimodalLines: Data
At ISMIR 2003 in Baltimore (26-30 October 2003) Avery Wang of Shazam Entertainment described an algorithm that recognises a recording from a few seconds of sound picked up by a phone's microphone. At the time of the paper the service held over 1.7 million tracks and was live in the UK and Germany; in the UK it had launched in August 2002, reached by dialling 2580.
October 2003Research · Robotics and autonomyLines: Embodiment · Learning theory
In October 2003 Andrew Davison presented a system in which a single moving camera at once fixes its own position and builds a map of its surroundings — at the rate frames arrive, thirty times a second. An extended Kalman filter holds the camera pose and the three-dimensional coordinates of landmarks in one state vector, and new points enter the map through a distribution over depth hypotheses.
February 10, 2004Milestone · FoundationsLines: Symbolic AI · Learning theory
On 10 February 2004 the W3C made OWL, a web ontology language, a Recommendation, and with it a reworked RDF specification. OWL added classes, restrictions on properties and three levels of expressiveness to the graph: Lite, DL and Full.
April 25, 2004Research · MultimodalLines: Data · Work
Luis von Ahn and Laura Dabbish of Carnegie Mellon posted the ESP game on 9 August 2003: two strangers see the same image and score when they type the same word. By 10 December 13,630 players had given 1,271,451 labels for 293,760 images. The paper was presented at CHI 2004 on 25 April 2004.
In August 2004 at SIGGRAPH Carsten Rother, Vladimir Kolmogorov and Andrew Blake of Microsoft Research in Cambridge described GrabCut: a person roughly draws a rectangle around an object, and the program separates it from the background by repeating a graph cut while refining its colour models.
October 2004Research · FoundationsLines: Symbolic AI · Data
In October 2004 Hugo Liu and Push Singh of the MIT Media Lab described ConceptNet 2.0, a base of everyday knowledge holding 1.6 million assertions over 300,000 nodes. It was not written by hand: it was extracted automatically from the 700,000 English sentences that more than 14,000 people had typed into the Open Mind Common Sense website.
Fei-Fei Li, Rob Fergus and Pietro Perona assembled a set of 101 object categories plus a background class, and presented a Bayesian method that learns from a handful of examples. Object class recognition had until then been tested mostly on one or two categories, on four by Weber and colleagues and on six by Fergus and colleagues; the authors call their dataset fifteen times larger than anything before it.
April 11, 2005Benchmark · ImagesLines: Evaluation · Data
On 11 April 2005 in Southampton the results of the first PASCAL Visual Object Classes challenge were announced. Four classes — motorbikes, bicycles, people and cars; two tasks — say whether an object is in the image, and put a box around it. Twelve teams entered and six presented at the workshop.
June 2005Research · ImagesLines: Learning theory · Data
In June 2005 Navneet Dalal and Bill Triggs showed that a dense grid of histograms of oriented gradients, normalised over overlapping blocks and fed to a linear support vector machine, detects people in images at least ten times more accurately than existing feature sets. Alongside the method they released the INRIA Person set of 1805 human images at 64 by 128 pixels.
September 4 – 8, 2005Benchmark · MultimodalLines: Speech · Evaluation
At Interspeech 2005 in Lisbon (4-8 September 2005) Alan Black (Carnegie Mellon) and Keiichi Tokuda (Nagoya Institute of Technology) reported the first contest of speech synthesisers on shared data: six teams from three continents built voices in a week or two from the same CMU ARCTIC databases and spoke the same 250 sentences. Listeners rated naturalness and typed in what they heard; the HMM-based system from Nagoya came first with every group of listeners.
On 8 September 2005 Bryan Russell, Antonio Torralba, Kevin Murphy and William Freeman described LabelMe in an MIT memo: a web tool in which anyone outlines objects in a photo with a polygon and names them in their own words. The tool came with 10,000 images, 3000 of them already labelled; by 21 December 2006 the database held 111,490 polygons.
November 2, 2005Availability · FoundationsLines: Data · Work · Money and markets
On the evening of 2 November 2005 Amazon opened beta testing of a marketplace where a requester posts a small task through a programmatic call and a person completes it. Developer documentation, a WSDL file and a Requester Tool Kit were already in place at launch.
November 9, 2005Milestone · Robotics and autonomyLines: Embodiment · Money and markets
iRobot went public at 24 dollars a share, and its registration statement disclosed that it had sold more than 1.2 million Roomba vacuuming robots over three years, and that they accounted for 73.8% of 2004 revenue.
In April 2006 Google opened its statistical Arabic-English translation system to everyone, as Franz Och wrote on 28 April. The system was trained on billions of words of monolingual text and aligned human translations; in NIST's 2005 evaluation it had the highest BLEU of all participants for Arabic and Chinese.
In May 2006 Herbert Bay, Tinne Tuytelaars and Luc Van Gool presented SURF. Gaussian kernels are replaced by rectangular box filters computed in constant time through an integral image, the detector is built on the determinant of the Hessian, and the descriptor is cut to 64 numbers. On the same 800 by 640 image, detection plus description took 354 ms against 1036 ms for SIFT.
June 8 – 9, 2006Benchmark · LanguageLines: Evaluation · Data
On 8 and 9 June 2006, fourteen teams from eleven institutions translated between English, French, German and Spanish on the Europarl corpus. For the first time in this competition the outputs were scored both by BLEU and by people - about 180 hours of the participants' own labour - and the two scorings disagreed.
July 30 – August 3, 2006Research · ImagesLines: Data
At SIGGRAPH 2006 Noah Snavely, Steven Seitz (University of Washington) and Richard Szeliski (Microsoft Research) presented Photo Tourism. The system worked out where each of thousands of random internet photos of a landmark had been taken from, built a sparse 3D model from them, and let one move through it from shot to shot. For Notre Dame, 597 of 2635 Flickr photographs were registered.
December 2006Research · Robotics and autonomyLines: Embodiment
Ashutosh Saxena, Justin Driemeyer, Justin Kearns and Andrew Ng trained an algorithm to find a grasp point directly in an image, using only synthetic pictures for training; on real objects absent from training the robot picked up 87.5% of attempts.
April 5, 2007Benchmark · ImagesLines: Evaluation · Harm
On 5 April 2007 NEJM published a comparison of 43 facilities in three states over 1998-2002: where computer-aided detection was adopted, the specificity of screening mammography fell from 90.2% to 87.2%, biopsies rose by 19.7%, and the rise in sensitivity from 80.4% to 84.0% was not significant.
May 8, 2007Research · FoundationsLines: Money and markets
On 8 May 2007 Matthew Richardson, Ewa Dominowska and Robert Ragno of Microsoft published, in the proceedings of the WWW conference, a model that predicts the click-through rate of search ads with no history of impressions. Logistic regression on features of the ad, its keywords and its advertiser cut the estimation error by 30 per cent against a baseline that simply predicts the average click-through rate.
May 8, 2007Research · FoundationsLines: Symbolic AI · Data
On 8 May 2007 three Max Planck researchers published YAGO, an ontology of roughly a million entities and five million facts assembled automatically: the individuals and things come from Wikipedia's categories and infoboxes, the top of the hierarchy from WordNet. A sampled human check put accuracy at about 95 per cent.
In August 2007 at SIGGRAPH Shai Avidan (MERL) and Ariel Shamir (Interdisciplinary Center Herzliya) showed how to change an image’s proportions without distorting what matters: the program repeatedly removes or inserts a seam, a connected chain of the least noticeable pixels from top to bottom or left to right.
August 5 – 9, 2007Research · ImagesLines: Generative media · Data
In August 2007 at SIGGRAPH James Hays and Alexei Efros of Carnegie Mellon filled a cut-out part of a photograph with pieces of other photographs. First the 200 most similar scenes were sought among 2.3 million Flickr images, then the best place in each, and the insert was stitched in without a seam.
October 2007Benchmark · ImagesLines: Data · Evaluation
In October 2007 Gary Huang, Manu Ramesh, Tamara Berg and Erik Learned-Miller of the University of Massachusetts Amherst released Labeled Faces in the Wild: 13,233 photographs of the faces of 5749 people from news articles on the web, 1680 of them in two or more photos. The task is to say of a pair of photos whether they show the same person.
November 2007Research · FoundationsLines: Symbolic AI · Data · Openness
In November 2007 a group from Berlin, Leipzig and Pennsylvania presented DBpedia at the ISWC conference: 103 million RDF triples extracted from the infoboxes of the English Wikipedia. The dataset described more than 1.95 million things and answered queries over an open SPARQL endpoint.
November 2007Research · Robotics and autonomyLines: Embodiment · Learning theory
In November 2007 Georg Klein and David Murray presented PTAM: tracking the camera and building the map stopped being one computation and moved apart into two parallel threads on a dual-core machine. The tracking thread keeps up with every frame while the mapping thread runs batch optimisation over keyframes — and the map grows to thousands of points.
June 9, 2008Availability · FoundationsLines: Symbolic AI · Data · Openness
On 9 June 2008, at the SIGMOD conference, the Metaweb Technologies team described Freebase: a knowledge base of more than 125 million tuples, more than 4,000 types and more than 7,000 properties. Read and write access was open to anyone over an HTTP API with its own query language, MQL.
At CVPR in June 2008 Pedro Felzenszwalb (University of Chicago), David McAllester (TTI Chicago) and Deva Ramanan (UC Irvine) described a detector in which an object is a coarse template on gradient histograms plus finer part templates that can shift. Trained with a latent SVM, it doubled the best PASCAL 2006 result for people: 0.34 average precision against 0.16.
July 1, 2008Milestone · Robotics and autonomyLines: Embodiment · Work
IEEE Spectrum published Kiva Systems' deployment figures: 500 robots at a 30,000-square-metre Staples fulfilment centre, hundreds at a Walgreens distribution centre, and 600 picks an hour against 200 to 400 on a conveyor.
October 3, 2008Law and regulation · ImagesLines: Law and policy · Harm
On 3 October 2008 Illinois's Biometric Information Privacy Act (BIPA) took effect: a private company may collect a face scan, fingerprint or voiceprint only after written notice of the purpose and term and a written release from the person, and each violation gives the person a right to sue - 1,000 dollars if negligent, 5,000 if intentional.
In November 2008 Antonio Torralba, Rob Fergus and William Freeman of MIT described in IEEE PAMI a set of 79,302,017 images of 32 × 32 pixels. Over eight months seven search engines returned 97,245,098 pictures for queries on 75,846 WordNet nouns; after cleaning, images remained for 75,062 words. Each image’s only label is the query that found it.
April 8, 2009Benchmark · ImagesLines: Data · Evaluation · Work
On 8 April 2009 Alex Krizhevsky of the University of Toronto issued a technical report describing CIFAR-10, a labelled part of the 80 million tiny images collection at 32 by 32 pixels: 10 classes of 6,000 images, with 5,000 from each class in the training part. The same report has CIFAR-100, 100 classes of 600.
In August 2009 at SIGGRAPH Connelly Barnes, Eli Shechtman, Adam Finkelstein and Dan Goldman (Princeton, Adobe) described PatchMatch, a search for similar patches in an image 20 to 100 times faster than kd-trees: random matches first, then good ones are passed to neighbours and refined by random search.
2009Benchmark · Robotics and autonomyLines: Embodiment · Evaluation · Defence
The Learning Locomotion programme put six teams on the same LittleDog quadruped and measured them by number: each phase required clearing higher obstacles at higher speed, against thresholds DARPA fixed in advance.
On 30 April 2010 Adobe began shipping Creative Suite 5 and with it Photoshop CS5, in which a selected object can be removed and the program fills the empty space from its surroundings - the Content-Aware Fill feature. Photoshop CS5 had been announced on 12 April.
April 2010Research · FoundationsLines: Work · Data
In April 2010 five researchers at the University of California, Irvine pulled together six surveys of Mechanical Turk workers spanning twenty months. The share of workers from India rose from 5 per cent in November 2008 to 36 per cent a year later.
February – April 2010Availability · LanguageLines: Speech · Agents
A company spun out of SRI released an application that understood a spoken request and carried out the action itself, booking a table or calling a taxi.
May 4, 2010Availability · Robotics and autonomyLines: Embodiment · Evaluation
Willow Garage gave away 11 PR2 robots worth over 4.4 million dollars in total to eleven institutions chosen from 78 proposals, together with a shared software stack.
April 2010 — May 30, 2010Availability · FoundationsLines: Evaluation · Money and markets
In the spring of 2010 Kaggle Pty Ltd opened a platform for forecasting competitions: an organisation posts data and a prize, analysts submit solutions, the best takes the prize. The first competition was a forecast of the Eurovision vote with a $1,000 prize; on 30 May 2010 Kaggle announced the winner among 22 teams.
May 2010Research · Robotics and autonomyLines: Embodiment
A PR2 robot at Berkeley picked a randomly tossed towel off a table, found its corners from image geometry, re-grasped, untwisted and folded it — and did so successfully in all 50 end-to-end trials on towels it had not seen before.
September 23 – December 7, 2010Milestone · FoundationsLines: Money and markets
On 23 September 2010 the UK companies register incorporated company 07386350 as Friars 2022 Limited, with M&R Secretarial Services as secretary and a Cambridge address. On 15 November it was renamed DeepMind Technologies Limited, and on 7 December Demis Hassabis became a director and Mustafa Suleyman secretary. Shane Legg and Suleyman became directors only on 23 December 2011.
June 16, 2011Law and regulation · Robotics and autonomyLines: Law and policy · Autonomous driving
On 16 June 2011 the Governor of Nevada approved Assembly Bill 511: the Department of Motor Vehicles must by 1 March 2012 adopt regulations under which autonomous vehicles may operate on the state's highways, with requirements for the vehicle, insurance and safety, and testing only in specified areas.
October 4, 2011Availability · LanguageLines: Speech · Agents
A voice assistant became a built-in part of the phone, and millions of people addressed a machine in a sentence rather than a button for the first time.
January 23, 2012Availability · FoundationsLines: Data · Openness
On 23 January 2012 the Common Crawl Foundation announced that its web crawl archive - five billion pages - was hosted on Amazon's Public Data Sets. Amazon pays for the storage; a user pays only for their own compute.
April 6, 2012Milestone · FoundationsLines: Money and markets · Evaluation
On 6 April 2012 Xavier Amatriain and Justin Basilico of Netflix wrote that from the open contest the company had put into production two algorithms from the winning entry of the first Progress Prize in 2007, matrix factorisation and restricted Boltzmann machines. The blend of hundreds of models that won the Grand Prize was not put in: the accuracy gain measured offline did not justify the engineering effort.
May 16, 2012Announcement · FoundationsLines: Symbolic AI · Search and reasoning · Data
On 16 May 2012 Amit Singhal announced that Google Search would start answering about things rather than strings. The company claimed more than 500 million objects and more than 3.5 billion facts and relationships between them, drawn from Freebase, Wikipedia, the CIA World Factbook and other sources.
July 3, 2012Research · FoundationsLines: Learning theory
On 3 July 2012 Geoffrey Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever and Ruslan Salakhutdinov of the University of Toronto posted a preprint: on every training case, omit each hidden unit at random with probability 0.5, and halve the outgoing weights at test time. On MNIST the best standard network made 160 test errors; dropout brought it to about 130, and dropping 20 per cent of the input pixels as well to about 110. A single ImageNet 2010 network went from 48.6 to 42.4 per cent error.
October 29, 2012Availability · FoundationsLines: Symbolic AI · Data · Openness
On 29 October 2012 Wikimedia Deutschland opened Wikidata. At first it could only create items and attach Wikipedia articles in different languages to them; every item got an identifier beginning with Q, and all the data got the CC0 licence, which is the public domain.
November 21, 2012Law and regulation · FoundationsLines: Defence · Law and policy
On 21 November 2012 Deputy Secretary of Defense Ashton Carter issued Directive 3000.09. It defined an autonomous weapon system and required such systems to be designed so that commanders and operators could exercise appropriate levels of human judgment over the use of force.
December 2012Research · ImagesLines: Harm · Evaluation · Data
In the December 2012 issue of IEEE Transactions on Information Forensics and Security, Brendan Klare and colleagues measured six face recognition systems on 102,942 police photographs split into eight demographic cohorts. The most accurate of the three commercial systems matched 88.7 per cent of black subjects against 94.4 per cent of white ones, 89.5 per cent of women against 94.4 per cent of men, and 91.7 per cent of 18-to-30-year-olds against 94.6 per cent of 30-to-50-year-olds.
On 6 August 2013 Facebook engineer Lars Backstrom, in the first post of the News Feed FYI series, explained that on each visit to News Feed a user has on average 1,500 potential stories from friends and Pages, and ranking picks an average of 300 of them a day from signals of interaction, reactions and hiding. The same day the company announced that unread stories still gathering reactions could return near the top of the feed.
On 11 November 2013 Ross Girshick, Jitendra Malik and colleagues at Berkeley showed that a network trained to classify whole ImageNet images could find objects: it scores about 2,000 region proposals, and mean average precision on PASCAL VOC 2007 reaches 48%.
January 24 – February 10, 2014Milestone · FoundationsLines: Money and markets · Safety
The UK register dates the change of ownership of DeepMind Technologies to 24 January 2014: shares were allotted that day, leaving capital of GBP 1,640.64, and a director was appointed under the same date. On 3 February six directors were removed at once — Hassabis, Suleyman, Legg and three investors — and on 10 February two new ones arrived. What Google paid is stated nowhere.
May 1, 2014Benchmark · ImagesLines: Data · Evaluation · Work
On 1 May 2014 Tsung-Yi Lin and colleagues from Cornell, Caltech, Brown, UC Irvine and Microsoft Research released Microsoft COCO, 328,000 images of everyday scenes with 2.5 million labelled instances of 91 object categories, each with its own segmentation mask.
June 3, 2014Research · LanguageLines: Learning theory
On 3 June 2014 Kyunghyun Cho, Yoshua Bengio and co-authors at the Universite de Montreal, the University of Gothenburg and the Universite du Maine posted an RNN Encoder-Decoder: one recurrent network encodes a phrase into a fixed-length vector and another decodes it into a phrase in the other language. For it they proposed a new hidden unit with a reset gate and an update gate, motivated by the LSTM but much simpler. As a feature in a phrase-based English-French system it raised test BLEU from 29.33 to 29.96, and to 31.18 together with a neural language model and a word penalty.
In June 2014 Facebook AI Research and Tel Aviv University presented DeepFace at CVPR: a network of over 120 million parameters, after a three-dimensional alignment of the face, reached 97.35% on Labeled Faces in the Wild, cutting the error by over 27%.
On 3 July 2014 Oxford University Press published Superintelligence: Paths, Dangers, Strategies by Nick Bostrom, director of Oxford's Future of Humanity Institute, 352 pages. The book examines how machine intelligence could surpass human intelligence, why its goals might diverge from ours and what means of control are possible at all.
July 31, 2014Milestone · FoundationsLines: Money and markets · Autonomous driving
On 31 July 2014 Mobileye N.V. issued a final prospectus for 35,589,000 ordinary shares at $25.00, $889.7 million gross. The company itself sold 8,325,000 shares and received $197.7 million before expenses; the other $647.5 million went to selling shareholders. The document was the first to show a computer-vision developer's revenue: $81.2 million for 2013.
August 24, 2014Research · FoundationsLines: Money and markets
On 24 August 2014 eleven Facebook engineers, among them Joaquin Quiñonero Candela, described at the ADKDD workshop a model for predicting clicks on ads: gradient boosted trees transform the features, and a logistic regression updated online gives the probability of a click. The combination beat either method alone by more than 3 per cent in normalized entropy.
November 6, 2014Availability · LanguageLines: Speech · Agents
Amazon released a speaker that listens to the room continuously and answers when addressed, making voice the primary interface rather than an addition.
November 17, 2014Research · MultimodalLines: Generative media
On 17 November 2014 Oriol Vinyals and colleagues at Google described a network that composes a sentence about an image: a convolutional network encodes the picture and an LSTM writes the description. On Pascal the BLEU score was 59 against a previous best of 25 and human performance around 69.
December 2014Research · FoundationsLines: Learning theory
Kingma and Ba described a method that picks a learning rate for each parameter separately from estimates of the first and second moments of the gradient.
On 12 March 2015 Google described FaceNet: a network turns a face into a 128-byte embedding in which distance means similarity. On Labeled Faces in the Wild, 99.63%; on YouTube Faces DB, 95.12%; on both the error is 30% below the best published result.
April 2, 2015Research · Robotics and autonomyLines: Embodiment
Sergey Levine, Chelsea Finn, Trevor Darrell and Pieter Abbeel showed that training perception and control jointly, rather than as two separate stages, makes a robot roughly twice as reliable on the hardest contact-rich tasks.
In April 2015 Vassil Panayotov, Daniel Povey and colleagues at Johns Hopkins University published LibriSpeech, about 1,000 hours of read English at 16 kHz aligned with its text, taken from public-domain LibriVox audiobooks.
May 3, 2015Research · MultimodalLines: Data · Evaluation
On 3 May 2015 Stanislaw Antol, Devi Parikh and colleagues proposed the task of open-ended answers to questions about images, with a dataset for it. The first version held 123,285 MS COCO images, 215,150 questions and 430,920 answers, and 10,000 abstract scenes with 30,000 questions.
On 18 May 2015 Olaf Ronneberger and colleagues in Freiburg described U-Net, a convolutional network whose contracting path is joined to a symmetric expanding one. With it they won the ISBI 2015 cell tracking challenge in two categories, and a 512x512 image is segmented in under a second.
June 5, 2015Availability · FoundationsLines: Tools · Openness
On 5 June 2015 Preferred Networks of Tokyo released Chainer 1.0, a Python deep learning library in which the computation graph is not declared in advance but recorded by itself as the forward computation runs. In December that year Seiya Tokui, Kenta Oono, Shohei Hido and Justin Clayton named the approach Define-by-Run at the LearningSys workshop of NIPS.
On 8 June 2015 Joseph Redmon, Ali Farhadi and colleagues at the University of Washington and the Allen Institute for AI described YOLO: one network in one pass predicts boxes and classes for the whole image. It runs at 45 frames per second, with 58.8 mAP on PASCAL VOC 2007.
November 9, 2015Availability · FoundationsLines: Tools · Openness
Google open-sourced its own system for training models: computation is described as a graph and can run on a processor, a graphics accelerator or a cluster.
December 28, 2015Research · FoundationsLines: Money and markets
On 28 December 2015 Carlos Gomez-Uribe and Neil Hunt of Netflix described the company's recommender system in ACM Transactions on Management Information Systems: it influences the choice of about 80 per cent of hours streamed, with search accounting for the other 20. The authors estimated that personalisation and recommendations together save the company more than a billion dollars a year by reducing subscriber churn.
February 4, 2016Research · FoundationsLines: Learning theory · Compute
On 4 February 2016 Volodymyr Mnih and co-authors at Google DeepMind posted asynchronous variants of four reinforcement learning algorithms, in which many actor-learners run in parallel on copies of the environment instead of learning from a replay memory. The best, asynchronous advantage actor-critic (A3C), after four days on 16 CPU cores with no GPU reached a mean of 623.0 and a median of 112.6 per cent of human-normalised score on 57 Atari games, against 121.9 and 47.5 per cent for DQN after eight days on a GPU.
March 7, 2016Research · Robotics and autonomyLines: Embodiment · Data
Google trained a convolutional network to predict whether a proposed gripper motion would end in a successful grasp, on 800,000 attempts collected over two months by six to fourteen manipulators running in parallel.
March 23 – 25, 2016Availability · LanguageLines: Harm
On 23 March 2016 Microsoft launched Tay, a chatbot for 18- to 24-year-olds in the US that learned from conversation, on Twitter. Within the first day a coordinated group of users exploited a vulnerability and Tay posted offensive words and images; on 25 March the company apologised, with the bot already offline.
April 27, 2016Law and regulation · FoundationsLines: Law and policy · Data
On 27 April 2016 the European Parliament and the Council signed the General Data Protection Regulation (GDPR), applicable from 25 May 2018. Article 22 gives a person the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning them or similarly significantly affects them.
May 23, 2016Research · FoundationsLines: Harm · Evaluation
On 23 May 2016 ProPublica published an audit of COMPAS, the recidivism risk scale that United States courts used in bail and parole decisions. The outlet took the scores of more than 7,000 people arrested in Broward County, Florida in 2013 and 2014 and checked them against new charges over the following two years. Of those labelled higher risk who did not re-offend, the share was 44.9 per cent among black defendants against 23.5 per cent among white ones.
June 16, 2016Benchmark · LanguageLines: Evaluation · Data · Work
On 16 June 2016 Stanford published a set of 107,785 questions written by crowdworkers against 536 Wikipedia articles, where the answer is a span of the passage itself. The authors' own model reached 51.0 F1; a human on the test set reached 86.8.
On 21 June 2016 Dario Amodei and Chris Olah (Google Brain), Jacob Steinhardt (Stanford), Paul Christiano (Berkeley), John Schulman (OpenAI) and Dan Mané (Google Brain) posted the preprint 'Concrete Problems in AI Safety'. They reduced the risk of accidents in machine learning systems to five research problems: side effects, reward hacking, scalable oversight, safe exploration and distributional shift.
On 24 June 2016 sixteen authors at Google posted the Wide & Deep preprint: a wide linear model on cross-product features and a deep network with embeddings, trained jointly. The model was put into production and tested in Google Play; in an online experiment on 1 per cent of users it raised app acquisitions from the store's main landing page by 3.9 per cent against the previous model, a wide-only logistic regression.
July 8, 2016Research · FoundationsLines: Harm · Evaluation
On 8 July 2016 the research department of Northpointe answered ProPublica's audit with its own analysis of the same data. The report by William Dieterich, Christina Mendoza and Tim Brennan gave the area under the ROC curve for the general recidivism scale on the outlet's main sample as 0.69 for black and 0.69 for white defendants, and found the share of those labelled higher risk who did not re-offend to be lower, not higher, for black defendants: 37 per cent against 41 per cent.
On 7 September 2016 Paul Covington, Jay Adams and Emre Sargin of Google described, for the RecSys conference, YouTube's recommender system built on deep networks in two stages: candidate generation narrows a corpus of millions of videos to hundreds, and a separate ranking network orders them by predicting expected watch time rather than the probability of a click. The models have about a billion parameters and are trained on hundreds of billions of examples.
September 19, 2016Research · FoundationsLines: Learning theory · Harm · Evaluation
On 19 September 2016 Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan submitted a paper that formalised three fairness conditions for risk scores and proved that no assignment of scores satisfies all three at once, except in two degenerate cases: when prediction is perfect, or when the groups have equal base rates. The paper names the argument over COMPAS as its occasion.
September 26, 2016Research · LanguageLines: Compute
On 26 September 2016 Google posted GNMT, a neural translation system with eight LSTM layers in the encoder and eight in the decoder, and on 27 September moved 100 per cent of Chinese-to-English translations in Google Translate to it, about 18 million a day. In human evaluation it reduced errors against the phrase-based statistical system by 60 per cent on average.
November 21, 2016Research · ImagesLines: Generative media
On 21 November 2016 Phillip Isola, Alexei Efros and colleagues at Berkeley showed that a conditional adversarial network with one architecture and objective translates image to image across many tasks: labels to street scenes, aerial photos to maps, day to night, edges to photos.
November 2016Milestone · FoundationsLines: Money and markets · Science
In November 2016 the University of Texas System audit office found that the MD Anderson cancer centre had paid external firms 62.1 million dollars for the Oncology Expert Advisor, an adviser built on IBM Watson technology, and that the system was not in clinical use and not connected to the hospital's new medical record.
December 13, 2016Research · ImagesLines: Science · Evaluation
On 13 December 2016 JAMA published a Google convolutional network trained on 128,175 fundus photographs: on two independent sets it detected retinopathy needing referral to an ophthalmologist with an area under the ROC curve of 0.991 and 0.990.
January 23, 2017Research · FoundationsLines: Learning theory · Compute
On 23 January 2017 Noam Shazeer and co-authors at Google Brain, among them Jeff Dean and Geoffrey Hinton, posted a layer of up to thousands of expert networks of which a trainable gate selects only a few for each example (noisy top-k gating). Placed between LSTM layers, it gave models of up to 137 billion parameters and, in the authors' words, greater than 1000x improvements in model capacity with only minor losses in computational efficiency. On WMT'14: 40.56 BLEU English-French and 26.03 English-German.
January 25, 2017Research · ImagesLines: Science · Evaluation
On 25 January 2017 Nature published a Stanford convolutional network trained on 129,450 clinical images of 2,032 skin diseases: on two tasks, carcinomas against benign keratoses and melanomas against naevi, it matched 21 board-certified dermatologists.
January 2017 — January 30, 2017Milestone · FoundationsLines: Search and reasoning
In January 2017 Carnegie Mellon University's Libratus played 120,000 hands of heads-up no-limit Texas hold'em over twenty days against four professionals at Rivers Casino in Pittsburgh, and finished on 30 January ahead by $1,766,250 in chips.
In January 2017 the Future of Life Institute published the 23 Asilomar AI Principles, developed at its Beneficial AI 2017 conference at Asilomar. The principles are grouped under research issues, ethics and values, and longer-term issues, and were opened for signature.
March 27, 2017Research · Robotics and autonomyLines: Embodiment
Dex-Net 2.0 was trained on 6.7 million synthetic point clouds generated from 3D models, with no attempts on a real robot at all — and the resulting policy gave 99% precision on forty unseen household objects.
March 29, 2017Research · MultimodalLines: Speech · Generative media
On 29 March 2017 Yuxuan Wang and colleagues at Google described Tacotron, a model that synthesises speech directly from text characters and is trained from scratch on text-audio pairs. Its mean opinion score for naturalness was 3.82 of 5, against 3.69 for a production parametric system.
In March 2017 Jort Gemmeke and colleagues at Google presented AudioSet, a hierarchical ontology of sound events and 2,084,320 human-labelled 10-second clips from YouTube videos; labellers checked whether a given class of sound was present in a segment.
April 26, 2017Milestone · FoundationsLines: Defence
On 26 April 2017 Deputy Secretary of Defense Robert O. Work signed a memo establishing the Algorithmic Warfare Cross-Functional Team "to accelerate DoD's integration of big data and machine learning". Its first task was to automate the processing of full-motion video from drones for the campaign against ISIS.
May 23 – 27, 2017Milestone · FoundationsLines: Search and reasoning
From 23 to 27 May 2017 in Wuzhen, DeepMind and the China Go Association held the Future of Go Summit: a three-game match against the world number one, Ke Jie, Pair Go in which two professionals each had AlphaGo as partner, and a game against a team of five professionals. AlphaGo won all three games, and DeepMind announced the match as the program's last.
June 12, 2017Research · FoundationsLines: Learning theory · Safety
On 12 June 2017 Paul Christiano and Dario Amodei of OpenAI, Jan Leike, Miljan Martic and Shane Legg of DeepMind, and Tom Brown posted a method in which a person compares pairs of short segments of an agent's trajectory, a reward model is fitted to those choices, and the agent learns by reinforcement on the predicted reward. On eight MuJoCo robotics tasks and seven Atari games it worked without access to the reward function, with human feedback on less than 1 per cent of the agent's interactions with the environment.
July 4, 2017Availability · Robotics and autonomyLines: Autonomous driving · Openness
On 4 July 2017 Baidu put Apollo 1.0 on GitHub, the open code of a platform for autonomous driving. The first version did one thing: replay the route and speed a human driver had driven, in an enclosed venue, by GPS. It could not perceive obstacles close by and did not drive on public roads. Baidu had promised to open the code on 18 April.
July 20, 2017Research · FoundationsLines: Learning theory
On 20 July 2017 John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov of OpenAI posted proximal policy optimisation: a surrogate objective that clips the ratio of new to old policy probabilities (epsilon = 0.2 in the experiments), so that the same batch of samples can be used for several epochs of minibatch updates with first-order methods alone. On seven MuJoCo tasks of one million timesteps PPO beat the previous policy gradient methods almost everywhere; on 49 Atari games it won 30 by average reward over training, against 18 for ACER and one for A2C.
July 20, 2017Milestone · FoundationsLines: Law and policy
The State Council published a plan for the country to become the world's major centre of AI innovation by 2030, with interim targets for 2020 and 2025.
August 8 – 21, 2017Milestone · FoundationsLines: Money and markets · Autonomous driving · Compute
On 21 August 2017 Intel completed its tender offer for Mobileye's shares, reaching 97.3 per cent after 84.4 per cent on 8 August. The announcement of 13 March had named $63.54 a share, an equity value of about $15.3 billion and an enterprise value of $14.7 billion. The annual report for 2017 says something else: total consideration of $14.5 billion, net of $366 million of cash that came with the company.
September 7, 2017Availability · FoundationsLines: Tools · Openness
On 7 September 2017 Microsoft and Facebook jointly announced an open exchange format for neural networks. A trained model could now be moved between frameworks without being rewritten.
December 14, 2017Research · FoundationsLines: Work · Data · Money and markets
On 14 December 2017 a preprint appeared that measured hourly pay on Mechanical Turk by recording the work rather than by asking: 2,676 workers, 3.8 million tasks. The median came to about $2 an hour, and only 4 per cent earned more than the US federal minimum of $7.25.
January 21, 2018Research · ImagesLines: Harm · Evaluation · Data
On 21 January 2018 the proceedings of the first Conference on Fairness, Accountability and Transparency carried a paper by Joy Buolamwini and Timnit Gebru. They showed that two widely used evaluation sets are composed mostly of lighter-skinned people, 79.6 per cent in IJB-A and 86.2 per cent in Adience, assembled a balanced set of 1,270 individuals, and measured three commercial gender classifiers on it. Error on darker-skinned women ran from 20.8 to 34.7 per cent; on lighter-skinned men it never exceeded 0.8 per cent.
December 11, 2017 – February 7, 2018Milestone · ImagesLines: Harm · Generative media · Law and policy
On 11 December 2017 Motherboard reported that a Reddit user called deepfakes was putting actresses' faces into pornographic videos with a network trained in Keras and TensorFlow. On 7 February 2018 Reddit closed r/deepfakes and banned intimate images made without consent, including those that have been faked.
April 9, 2018Law and regulation · FoundationsLines: Safety
On 9 April 2018 OpenAI published a charter with four principles: broadly distributed benefits, long-term safety, technical leadership and cooperative orientation. It defines AGI as highly autonomous systems that outperform humans at most economically valuable work and names humanity as the organisation's primary fiduciary duty.
April 11, 2018Availability · ImagesLines: Law and policy · Science
On 11 April 2018 the FDA permitted the marketing of IDx-DR, software that decides from fundus photographs whether an adult with diabetes has more than mild retinopathy. The FDA called it the first authorised device whose screening decision needs no clinician to interpret the image as well.
April 20, 2018Benchmark · LanguageLines: Evaluation · Data
On 20 April 2018 researchers at New York University gathered nine existing language understanding tasks under one leaderboard with privately held test data and a single averaged score. The best baseline in the paper itself reached 60.3.
April 2018Milestone · FoundationsLines: Compute · Money and markets
In April 2018 TSMC took its N7 process into volume production. By July 2020 the node had made one billion good dies for well over a hundred products from dozens of customers, and TSMC was the first foundry to bring 7nm to high-volume manufacturing.
June 1 – 7, 2018Milestone · FoundationsLines: Defence · Safety
On 1 June 2018 cloud chief Diane Greene told staff the Project Maven contract would not be renewed. On 7 June Sundar Pichai published principles in which Google undertook not to build AI for weapons or for surveillance violating international norms.
June 11, 2018Benchmark · LanguageLines: Evaluation · Data
On 11 June 2018, 53,775 questions the passage does not answer were added to SQuAD, written by crowdworkers to look like ones it does. The best model fell from 85.8 to 66.3 F1 while a human held at 89.5.
June 27, 2018Research · Robotics and autonomyLines: Embodiment
QT-Opt was trained by reinforcement learning on over 580,000 real grasp attempts; the policy reached 96% success on unseen objects and developed behaviours nobody had specified for it.
August 1, 2018Research · Robotics and autonomyLines: Embodiment
OpenAI trained a five-fingered Shadow hand to reorient a block in its palm, with the policy trained entirely in simulation under randomised physical properties; on hardware it gave a median of 13 successful rotations in a row.
September 28, 2018Law and regulation · LanguageLines: Law and policy · Harm
On 28 September 2018 the Governor of California approved SB 1001, operative from 1 July 2019: it is unlawful to use a bot to communicate with a person online while misleading them about its artificial identity with intent to incentivise a purchase or influence a vote. Whoever clearly discloses that it is a bot is not liable.
November 17, 2018Availability · LanguageLines: Tools · Openness
On 17 November 2018 Hugging Face released the library first named pytorch-pretrained-bert. It gave access to pretrained transformer models through one interface, two lines of code instead of an implementation of your own.
December 7, 2018Availability · FoundationsLines: Tools
On 7 December 2018 the first public release of JAX appeared, a library that takes ordinary Python and NumPy code and applies transformations to it: differentiation, compilation, vectorisation, parallelisation.
December 12, 2018Research · ImagesLines: Generative media · Data
On 12 December 2018 NVIDIA described StyleGAN, the generator of an adversarial network in which the latent code controls style at every level through AdaIN while separate noise adds random detail. On the new FFHQ face set its FID was 4.40 against 8.04 for Progressive GAN.
December 2018 — December 19, 2018Milestone · FoundationsLines: Compute
On 19 December 2018 DeepMind's AlphaStar beat the professional Grzegorz "MaNa" Komincz 5-0 at StarCraft II, after the same result against Dario "TLO" Wünsch; the company told of both matches on 24 January 2019. On 30 October 2019 Nature published the paper: AlphaStar played at Grandmaster level with all three races.
January 24, 2019Research · Robotics and autonomyLines: Embodiment
ETH Zurich trained a control policy in simulation and transferred it to the ANYmal quadruped: the robot ran at 1.5 m/s against a previous record of 1.2 m/s, and stood up after falling from an arbitrary pose.
February 11, 2019Law and regulation · FoundationsLines: Law and policy
On 11 February 2019 the President of the United States signed Executive Order 13859, 'Maintaining American Leadership in Artificial Intelligence': five principles, six objectives, and instructions to agencies - OMB to issue a memorandum on regulating AI applications within 180 days, NIST to issue a plan for federal engagement in technical standards within 180 days.
March 11, 2019Milestone · FoundationsLines: Money and markets · Safety
On 11 March 2019 Greg Brockman and Ilya Sutskever announced OpenAI LP, a 'capped-profit' company controlled by the board of the non-profit OpenAI. Returns for first-round investors are capped at 100 times their investment and anything above belongs to the non-profit. Most staff, around a hundred people, moved to the LP.
April 8, 2019Law and regulation · FoundationsLines: Law and policy
On 8 April 2019 the High-Level Expert Group on AI set up by the European Commission published its Ethics Guidelines for Trustworthy AI: four ethical principles, seven requirements for systems, and a self-assessment list for piloting. The December 2018 draft drew feedback from more than 500 contributors.
April 13, 2019Milestone · FoundationsLines: Compute
On 13 April 2019 OpenAI Five, an OpenAI program that controls five heroes and was trained by reinforcement learning, won two games in a row against OG, the reigning Dota 2 world champions. A preprint of 13 December 2019 summed up the training: one continuous run from 30 June 2018 to 22 April 2019.
May 2, 2019Benchmark · LanguageLines: Evaluation · Data
On 2 May 2019 the authors of GLUE declared their own set exhausted and released a harder one, keeping two of the nine tasks. A human estimate shipped with the release: 89.6 against 69.7 for the best baseline. On 6 January 2021 DeBERTa sat on top with 90.3.
May 21, 2019Law and regulation · ImagesLines: Law and policy · Harm
On 21 May 2019 the San Francisco Board of Supervisors gave final passage to Ordinance No. 103-19 by 10 votes to 1. It forbade city departments, the police among them, to obtain, retain or use face recognition technology or information derived from it, and for every other surveillance technology it required prior approval by the Board with an impact report and an annual audit by the Controller.
May 22, 2019Law and regulation · FoundationsLines: Law and policy
On 22 May 2019 the Council of the Organisation for Economic Co-operation and Development adopted Recommendation OECD/LEGAL/0449: five values-based principles for responsible stewardship of trustworthy AI and five recommendations to governments. Beyond the member countries, Argentina, Brazil, Colombia, Costa Rica, Peru and Romania adhered to it. The OECD described these at the time as the first such principles signed up to by governments.
June 5, 2019Research · FoundationsLines: Compute · Harm
On 5 June 2019 the University of Massachusetts Amherst put the energy and emissions of training language models into physical units. A full architecture search for a transformer was estimated at 274,120 hours on eight P100 accelerators, 656,347 kWh and 626,155 pounds of CO2 equivalent - roughly five times what a car emits across its whole life including fuel.
July 11, 2019Research · FoundationsLines: Search and reasoning · Compute
On 11 July 2019 Noam Brown and Tuomas Sandholm (Carnegie Mellon and Facebook AI) described in Science Pluribus, a program that beat professionals at six-player no-limit Texas hold'em: over 10,000 hands against five professionals at once it won an average of 48 thousandths of a big blind per hand.
On 29 July 2019 a Baidu group led by Yu Sun posted ERNIE 2.0, a scheme in which new pre-training tasks are added to a language model one after another and learned together with the earlier ones so that what was learned is not lost. The large model scored 83.6 on the hidden GLUE test set, 3.1 per cent above BERT, and beat BERT on all ten test tasks.
August 9, 2019Law and regulation · ImagesLines: Law and policy · Work
On 9 August 2019 the Governor of Illinois approved the Artificial Intelligence Video Interview Act, in force from 1 January 2020: an employer that has AI analyse a video recorded by an applicant must, before the interview, say so, explain how the AI works and what general types of characteristics it evaluates, and obtain consent; without consent the AI is not used.
August 23, 2019Availability · FoundationsLines: Compute · Tools
On 23 August 2019 Huawei released the Ascend 910 accelerator - 256 teraFLOPS at FP16, 512 teraOPS at INT8, 310W - together with its own framework, MindSpore. The two are meant to run together: the company claimed ResNet-50 training about twice as fast as other mainstream cards on TensorFlow.
October 16, 2019Research · Robotics and autonomyLines: Embodiment
OpenAI trained the same Shadow hand to solve a Rubik’s cube, generating progressively harder simulated environments automatically; the success rate was 60% on a half scramble and 20% on a full one.
October 23, 2019Availability · FoundationsLines: Data · Openness
On 23 October 2019 Google released C4 - about 750 GB of English text made from a single monthly slice of Common Crawl, the one for April 2019. The corpus was distributed through TensorFlow Datasets alongside the paper on the T5 models.
October 25, 2019Research · FoundationsLines: Harm · Evaluation
On 25 October 2019 Science published an analysis of a widely used commercial algorithm that selects patients for extra care: at the same score Black patients were considerably sicker than White patients, and removing the gap would raise their share among those selected from 17.7% to 46.5%.
November 19, 2019Research · FoundationsLines: Search and reasoning
On 19 November 2019 Julian Schrittwieser and colleagues at DeepMind described MuZero: tree search over a learned model that is not given the rules of the game. On 57 Atari games it set a new best result, and at Go, chess and shogi it matched AlphaZero, which knew the rules.
December 16, 2019Milestone · FoundationsLines: Money and markets · Compute
On 16 December 2019 Intel announced that it had acquired Habana Labs, an Israeli developer of programmable deep learning accelerators for the data centre, for "approximately $2 billion". Eight weeks later the company's annual report for fiscal 2019 named the same deal in one sentence and with a different number: approximately $1.7 billion.
December 19, 2019Benchmark · ImagesLines: Evaluation · Harm
On 19 December 2019 the United States National Institute of Standards and Technology published NISTIR 8280 by Patrick Grother, Mei Ngan and Kayee Hanaoka. A total of 18.27 million images of 8.49 million people from four government datasets were run through 189 mostly commercial algorithms from 99 developers. False positive rates across demographic groups often vary by factors of 10 to beyond 100, while false negatives usually vary by factors below 3.
January 1, 2020Benchmark · ImagesLines: Evaluation · Science
On 1 January 2020 Nature published a Google Health and DeepMind model for breast cancer screening: on a UK and a US set it cut false positives by an absolute 1.2 and 5.7 points and false negatives by 2.7 and 9.4, and it outperformed each of six radiologists.
On 9 January 2020 Detroit police arrested Robert Williams on the lawn in front of his house in Farmington Hills, in view of his wife and two daughters. The basis was an erroneous face recognition identification from security footage of a Shinola store, where five watches had been taken in October 2018. Williams was released around 9 p.m. the next day, some thirty hours later.
February 20, 2020Research · FoundationsLines: Science
On 20 February 2020 Cell published MIT work in which a network trained on 2,335 molecules to inhibit E. coli growth found halicin among known compounds: structurally unlike antibiotics, it killed a wide range of pathogens and cured pan-resistant Acinetobacter baumannii and C. difficile infections in mice.
March 19, 2020Research · ImagesLines: Generative media
On 19 March 2020 researchers from Berkeley, Google Research and San Diego described NeRF: a fully connected network that takes a point in space and a viewing direction and returns density and colour, and from a set of photographs with known camera poses it synthesises new views of complex scenes.
March 31, 2020Law and regulation · ImagesLines: Law and policy · Harm
On 31 March 2020 the Governor of Washington approved SB 6280, in force from 1 July 2021: state and local agencies may use face recognition only after a public accountability report, decisions with legal effects must be reviewed by a person, and ongoing surveillance and real-time identification require a warrant.
April 30, 2020Research · MultimodalLines: Generative media · Openness
On 30 April 2020 Prafulla Dhariwal and colleagues at OpenAI described Jukebox, a model that generates music with singing directly as a waveform: a multi-scale VQ-VAE compresses audio into discrete codes and Transformers model them. Generation can be steered by artist, genre and lyrics.
May 22, 2020Research · LanguageLines: Learning theory · Data
On 22 May 2020 Patrick Lewis and eleven co-authors from Facebook AI Research, University College London and New York University posted RAG: a pre-trained generator (BART-large, 400M parameters) combined with a dense vector index of Wikipedia, the December 2018 dump split into 21,015,324 passages of 100 words, with retriever and generator fine-tuned together. RAG-Sequence scored 44.5 exact match on Natural Questions, 56.1 on TriviaQA, 45.2 on WebQuestions and 52.2 on CuratedTrec; the paper claims a new state of the art on all four, for TriviaQA only on the split comparable with T5, since DPR has 57.9 on the usual one.
June 8, 2020Announcement · ImagesLines: Harm · Money and markets
On 8 June 2020 IBM's chief executive Arvind Krishna sent five members of the United States Congress a letter of proposals on police reform in which he stated that IBM no longer offers general purpose facial recognition or analysis software. The company said it opposes the use of any such technology, including other vendors', for mass surveillance, racial profiling and violations of human rights.
June 10, 2020Announcement · ImagesLines: Harm · Money and markets
On 10 June 2020 Amazon announced a one-year moratorium on police use of its facial recognition technology. Access was kept for Thorn, the International Center for Missing and Exploited Children and Marinus Analytics, which look for trafficking victims and missing children. The company said the year should give Congress time to put rules in place.
On 18 June 2020 TikTok published an account of how its recommender system fills the For You feed: it ranks videos by the user's interactions, video information (captions, sounds, hashtags) and, with lower weight, device and account settings. By the post, an account's follower count and its previous high-performing videos are not direct factors.
July 2, 2020Availability · FoundationsLines: Compute
On 2 July 2020 SK hynix began full-scale mass production of HBM2E memory: eight 16-gigabit dies joined by through-silicon vias, 16 gigabytes per stack and over 460 gigabytes a second across 1,024 lines at 3.6 gigabits a second per pin.
September 7, 2020Benchmark · FoundationsLines: Evaluation
On 7 September 2020 a test of 15,908 four-choice questions across 57 disciplines appeared - from elementary mathematics to professional medicine and law - to be taken with no task-specific training. GPT-3 at 175 billion parameters scored 43.9 per cent against 25 per cent for random guessing.
October 5, 2020Research · LanguageLines: Data · Openness
On 5 October 2020 the Masakhane community posted a paper on participatory research in machine translation for African languages: over 400 participants from at least 20 countries had within a year gathered new translation data and benchmarks for over 30 languages, a third of them evaluated by people. The authors come from several dozen institutions, from Pretoria and Kano to Stellenbosch.
December 2, 2020Milestone · FoundationsLines: Work · Harm
On 2 December 2020 Timnit Gebru's employment at Google ended, where she was a staff research scientist and co-lead of the ethical AI team. The occasion was management's demand that she withdraw an unpublished paper on the risks of large language models. Google wrote that it was accepting her resignation; Gebru said she had been fired. A petition supporting her was signed by 2,695 Google employees and 4,302 outside signatories.
December 8, 2020Milestone · FoundationsLines: Money and markets
On 8 December 2020 C3.ai issued a final prospectus for 15,500,000 Class A shares at $42.00 — $651 million to the public, of which $610.3 million to the company. The document showed what had not been published about enterprise AI: revenue of $156.7 million for the year to 30 April 2020, a loss of $69.4 million — and thirty "Entities" and sixty-four customers as of 31 October 2020.
December 31, 2020Availability · FoundationsLines: Data · Openness · Law and policy
On 31 December 2020 the independent group EleutherAI released The Pile - an English corpus of 825.18 gibibytes assembled from twenty-two separate sets. The largest of them, the web crawl Pile-CC, supplies 227.12 GiB, under a fifth of the whole.
February 2, 2021Law and regulation · ImagesLines: Law and policy · Data · Harm
On 2 February 2021 the Privacy Commissioner of Canada, together with the commissions of Quebec, British Columbia and Alberta, published the findings of a joint investigation into Clearview AI. The regulators found that scraping over three billion images from the web without consent does not fall under the publicly available exception and has no appropriate purpose, and described it as continual mass surveillance of everyone whose face ended up in the database.
March 5, 2021Benchmark · LanguageLines: Evaluation · Data
On 5 March 2021 Dan Hendrycks and colleagues at Berkeley and the University of Chicago released MATH, 12,500 problems from school mathematics competitions (7,500 training and 5,000 test) with step-by-step solutions, in seven subjects and five levels of difficulty. Large language models solved between 2.9 and 6.9%.
March 9, 2021Milestone · FoundationsLines: Compute · Science
On 9 March 2021 RIKEN announced that development of Fugaku was complete and the machine was open for shared use. Its 158,976 nodes each hold one Fujitsu A64FX processor on the Arm architecture, with 48 compute cores plus two assistant cores and 32 gibibytes of HBM2 at 1024 gigabytes per second, and no separate accelerator at all. Linpack measured 442.01 petaflops.
March 3 – 10, 2021Research · FoundationsLines: Safety · Data · Harm
In March 2021 the proceedings of ACM FAccT '21 carried a paper by Emily Bender, Timnit Gebru, Angelina McMillan-Major and a fourth author under a pseudonym. Its fourteen pages ask how big is too big and name four directions of risk: the environmental and financial cost of training, training data too large to be understood, the imputing to a model of an understanding it does not have, and harm from the text it produces.
April 18, 2021Research · FoundationsLines: Data · Harm · Openness
On 18 April 2021 the first description of the C4 corpus appeared - 156 billion tokens, 365 million documents - the corpus the T5 models were trained on. A blocklist filter removed African American English at a rate of 42 per cent against 6.2 per cent for White-aligned English, and verbatim matches turned up in all five evaluation sets checked.
April 20, 2021Research · LanguageLines: Learning theory
On 20 April 2021 Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen and Yunfeng Liu of Zhuiyi Technology in Shenzhen posted RoFormer: the position of a token is encoded with a rotation matrix that turns the vector through an angle depending on position, and self-attention naturally acquires a dependency on relative distance. The experiments of the first version are Chinese only: pre-trained on about 34 GB of Chinese text, on the CAIL2019-SCM legal case-matching task RoFormer reached 69.79 per cent on test at length 1,024, against 68.10 for WoBERT at 512.
April 21, 2021Research · FoundationsLines: Compute · Harm
On 21 April 2021 Google and the University of California, Berkeley published an energy account of five large models. Training GPT-3 - 10,000 V100 accelerators over 14.8 days, 3.14e23 operations - cost 1,287 MWh and 552.1 tonnes of CO2 equivalent. That is 0.01055 per cent of Google's total energy use for 2019.
On 17 June 2021 Edward Hu and seven co-authors at Microsoft posted LoRA: the pre-trained weights are frozen and trainable rank decomposition matrices are injected into each layer of the Transformer. For GPT-3 175B the first version reports 10,000 times fewer trainable parameters and a three times lower computation hardware requirement than full fine-tuning, at quality on par or better, and a checkpoint that shrinks from 350 GB to 35 MB.
June 21, 2021Benchmark · FoundationsLines: Evaluation · Harm
On 21 June 2021 JAMA Internal Medicine published an external validation of Epic's sepsis model, deployed in hundreds of US hospitals: on 38,455 hospitalisations in Michigan its area under the ROC curve was 0.63 against 0.76-0.83 in Epic's documentation, and it missed 67% of sepsis cases.
On 7 July 2021 Mark Chen and colleagues at OpenAI released HumanEval, 164 hand-written Python problems with unit tests. Codex solved 28.8% at the first attempt, GPT-3 0% and GPT-J 11.4%; with 100 samples per problem, 70.2%.
August 20, 2021Availability · ImagesLines: Data · Openness · Generative media
On 20 August 2021 the non-profit network LAION published 400 million English image-text pairs extracted from Common Crawl. CLIP did the selecting: a pair was kept when the cosine similarity between text and image was at least 0.3.
October 27, 2021Benchmark · LanguageLines: Evaluation · Data · Work
On 27 October 2021 Karl Cobbe and colleagues at OpenAI released GSM8K, 8.5 thousand grade-school maths word problems (7.5 thousand for training and a thousand for testing), each of 2 to 8 steps. With the set they proposed a verifier that scores the model's candidate solutions and picks the best.
November 11, 2021Research · LanguageLines: Data · Openness
On 11 November 2021 Kelechi Ogueji, Yuxin Zhu and Jimmy Lin presented AfriBERTa, a multilingual model trained from scratch on eleven African languages alone, on a corpus of 0.94 gigabytes, or 108.8 million tokens. XLM-R by comparison was trained on about 2395 gigabytes and 164 billion tokens, mBERT on roughly 100 gigabytes and 12.8 billion.
November 2021Law and regulation · FoundationsLines: Law and policy · Harm
In November 2021 the 193 member states of UNESCO adopted the Recommendation on the Ethics of Artificial Intelligence. The 43-page document rests on four values and ten principles, among them proportionality and do no harm, the right to privacy and data protection, and safety and security, and adds concrete areas of policy action from gender to data to international co-operation.
December 11, 2021Law and regulation · FoundationsLines: Law and policy · Work
On 11 December 2021 New York City's Local Law 144 became law: from 1 January 2023 an employer may not screen candidates with an automated tool unless it has had a bias audit within a year before use and a summary of the audit has been made public; candidates must be notified at least ten business days in advance.
December 7 – 30, 2021Milestone · FoundationsLines: Money and markets · Law and policy
The prospectus of 7 December 2021 offered 1.5 billion Class B shares in a range of HK$3.85 to HK$3.99, with the price to be named on 16 December. On 10 December the U.S. Department of the Treasury added SenseTime Group Limited to the Non-SDN Chinese Military-Industrial Complex Companies List. On 13 December the company announced the listing was postponed and all application monies refunded in full; on 20 December it relaunched with a supplemental prospectus and set dealings for 30 December.
January 28, 2022Research · LanguageLines: Learning theory
On 28 January 2022 Jason Wei and six colleagues at Google Brain posted chain-of-thought prompting: eight worked examples in the prompt show the intermediate steps before each answer, and the model then writes such steps itself. With a 137-billion-parameter model, accuracy on the GSM8K grade-school maths problems rose from 6.3 to 14.8 per cent, and to 19.5 with an external calculator, against 18 per cent for GPT-3 175B fine-tuned with a calculator; on MultiArith, from 7.6 to 45.0.
February 11, 2022Research · FoundationsLines: Compute
On 11 February 2022 Epoch AI assembled the training compute of 123 milestone systems going back to 1952 and cut the history into three eras. Before 2010 the amount doubled every 21.3 months; from 2010 every 5.7 months; and from September 2015 a separate series of the largest models doubles every 9.9 months, starting one to two orders of magnitude above the earlier trend.
March 1, 2022Law and regulation · FoundationsLines: Law and policy · Harm
On 1 March 2022 the Internet Information Service Algorithmic Recommendation Management Provisions took effect, signed on 31 December 2021 by the heads of four Chinese agencies. Article 24 requires an operator whose service has public opinion properties or social mobilisation capabilities to file its name, service form, domain of application, algorithm type and self-assessment report with a state register within ten working days of beginning to provide the service.
April 12, 2021 – March 4, 2022Milestone · FoundationsLines: Money and markets · Speech
On 12 April 2021 Microsoft and Nuance filed the terms of their agreement: $56.00 a share in cash, a 23 per cent premium to the closing price of 9 April, the transaction valued at $19.7 billion inclusive of net debt. On 4 March 2022 the deal closed, and the business-combinations note in Microsoft's annual report gives the total purchase price as $18.8 billion, consisting primarily of cash.
March 12, 2022Milestone · ImagesLines: Defence · Harm
On 12 March 2022, according to Clearview AI's chief executive Hoan Ton-That, Ukraine's defence ministry began using the company's face search engine free of charge. By 24 March Mykhailo Fedorov, vice prime minister and minister of digital transformation, confirmed to Reuters that Ukraine was using it to find the social media accounts of dead Russian soldiers and inform their families.
March 29, 2022Research · FoundationsLines: Compute · Data
On 29 March 2022 Jordan Hoffmann and colleagues at DeepMind posted the result of training over 400 language models from 70 million to over 16 billion parameters on 5 to 500 billion tokens: for compute-optimal training, model size and the number of tokens should be scaled equally. Chinchilla, 70 billion parameters trained on 1.4 trillion tokens with Gopher's budget (4 times more data than the 280-billion-parameter Gopher), outperformed Gopher, GPT-3, Jurassic-1 and Megatron-Turing NLG and reached 67.5 per cent on MMLU, a greater than 7 per cent improvement over Gopher.
March 31, 2022Availability · ImagesLines: Data · Openness · Generative media
On 31 March 2022 LAION published 5.85 billion image-text pairs, fourteen times more than LAION-400M. Of these 2.3 billion are English, another 2.2 billion are in more than a hundred other languages, and about a billion carry text in no assignable language.
April 4, 2022Research · Robotics and autonomyLines: Embodiment · Agents
SayCan paired a language model’s proposal of what to do with a learned value function’s estimate of what the robot can actually do right now; across 101 tasks given in natural language this gave 84% successful plans and 74% successful executions.
On 29 April 2022 Jean-Baptiste Alayrac and colleagues at DeepMind described Flamingo, models that join a frozen vision encoder to a frozen language model and take text, images and video interleaved. The largest has 80 billion parameters and beat fine-tuned models on 6 of 16 tasks with only 32 examples.
October 5, 2021 – April 30, 2022Milestone · FoundationsLines: Money and markets · Safety
A motion by the FTX debtors, filed in the Delaware bankruptcy court on 3 February 2024, describes the largest cheque in Anthropic's Series B more precisely than any press release: on 5 October 2021 Alameda Research Ventures paid $500,000,000 under an agreement for future equity; on 30 April 2022 the round closed; and on 13 May that agreement converted into 44,539,240 Series B preferred shares at $11.2261 each.
May 9, 2022Announcement · FoundationsLines: Money and markets · Openness · Tools
On 9 May 2022 Hugging Face announced a $100 million Series C led by Lux Capital, with Sequoia and Coatue participating alongside existing investors. In the same post the company gave the size of its repository: 100,000 pretrained models and 10,000 datasets, more than 10,000 companies using it, and growth from 30 staff to more than 120 in twelve months.
May 11, 2022Law and regulation · ImagesLines: Law and policy · Harm
On 11 May 2022 the Circuit Court of Cook County, Illinois, signed a consent order in ACLU v. Clearview AI: the company is permanently barred across the United States from giving private companies and individuals access to its face database scraped from the web, and for five years from giving it to any government body or company in Illinois, police included.
May 23, 2022Research · ImagesLines: Generative media · Openness
On 23 May 2022 Google Research described Imagen, a diffusion model that generates images from text encoded by a frozen T5-XXL language model. Without training on COCO it reached an FID of 7.27. Google released neither code nor a public demo.
On 27 May 2022 Tri Dao, Daniel Fu, Stefano Ermon and Christopher Re of Stanford and Atri Rudra of the University at Buffalo posted FlashAttention, an exact attention algorithm that computes attention in tiles so as to reduce reads and writes between a GPU's high-bandwidth memory and its on-chip SRAM. It trained BERT-large 15 per cent faster than the MLPerf 1.1 record, GPT-2 3 times faster at 1K tokens, and the Long Range Arena 2.4 times faster at 1K to 4K.
June 8, 2022Availability · FoundationsLines: Compute
On 8 June 2022 SK hynix began mass production of HBM3 and named the buyer for the first time: NVIDIA. 819 gigabytes a second per stack at 6.4 gigabits a second per pin, with supply matched to the shipping schedule of H100 systems in the third quarter of that year.
June 9, 2022Benchmark · FoundationsLines: Evaluation · Openness · Learning theory
On 9 June 2022, 442 authors from 132 institutions published 204 tasks chosen deliberately to lie beyond what models can do. Running them across sizes from millions to hundreds of billions of parameters showed that quality sometimes jumps at a particular scale - and that sometimes the jump is produced by a brittle metric.
June 13, 2022Availability · FoundationsLines: Compute
On 13 June 2022 LUMI was inaugurated in Kajaani, Finland, a EuroHPC machine on the Cray EX platform from Hewlett Packard Enterprise, capable by the release's own words of more than 375 petaflops. Its budget of EUR 144.5 million was found jointly by the EU joint undertaking and a consortium of ten states. At that moment it was the fastest machine in Europe and the third fastest in the world.
June 27, 2022Research · FoundationsLines: Compute · Money and markets
On 27 June 2022 Epoch AI measured the rate at which compute gets cheaper across 470 models of graphics card released from 2006 to 2021: FLOP/s per dollar doubles every 2.46 years. Huang's law - a 25-fold gain every five years, a doubling every 1.08 years - does not fit the data.
June 28, 2022Research · Robotics and autonomyLines: Embodiment
DayDreamer applied a world-model algorithm directly to physical robots, with no simulator and no resets: a quadruped learned to roll off its back, stand up and walk from scratch in one hour.
July 13, 2022Availability · ImagesLines: Generative media
On 13 July 2022 Midjourney announced an open beta: anyone could join its Discord server and generate images from a text prompt, with a limited number of free generations. Before that, by Motherboard's account, David Holz had opened access to all for 24 hours.
July 20, 2022Availability · ImagesLines: Generative media · Money and markets
On 20 July 2022 OpenAI opened the DALL·E 2 beta and began inviting a million people from the waitlist. Each got 50 free credits in the first month and 15 a month after, 115 credits cost $15, and users got full rights to commercialise the images.
August 9, 2022Law and regulation · FoundationsLines: Law and policy · Compute
On 9 August 2022 the CHIPS and Science Act became law: the CHIPS for America Fund received 50 billion dollars for fiscal years 2022-2026, 39 billion of it for semiconductor manufacturing incentives, and investment in plants a 25% tax credit. A recipient commits not to expand manufacturing materially in China for ten years.
October 10, 2022Availability · FoundationsLines: Compute · Money and markets
On 10 October 2022 Amazon opened general availability of EC2 Trn1 instances built on its own AWS Trainium silicon. The largest configuration carries 16 chips, 512 gigabytes of high-bandwidth memory, up to 3.4 petaFLOPS in TF32, FP16 and BF16, and up to 800 gigabits a second of network.
August 19 – October 13, 2022Milestone · FoundationsLines: Search and reasoning · Agents
From 19 August to 13 October 2022 Meta FAIR's agent Cicero played 40 games of Diplomacy anonymously on webDiplomacy.net against 82 people, negotiating with them in plain language. Its mean score was 25.8% against its opponents' 12.4%. The paper was published in Science on 22 November 2022.
October 7 – 21, 2022Law and regulation · FoundationsLines: Law and policy · Compute
In October 2022 the U.S. Bureau of Industry and Security put the accelerator on the export control list by two numbers taken together: a transfer rate to other chips of 600 gigabytes a second or more and, at the same time, a bit length times processing performance of 4800 or more. Separately it closed off equipment for logic below 16/14nm, DRAM below 18nm and NAND at 128 layers or more.
October 25, 2022Availability · LanguageLines: Tools · Agents
On 25 October 2022 the first release of LangChain appeared, a library for building applications out of language models by composition. Its own summary: "Building applications with LLMs through composability".
November 3, 2022Milestone · LanguageLines: Law and policy · Openness · Tools
On 3 November 2022 two plaintiffs suing under pseudonyms filed a class action in the Northern District of California against GitHub, Microsoft and OpenAI, case 3:22-cv-06823. The complaint says Copilot, trained on public repositories, emits other people's code without the three things eleven widely used open-source licences require: attribution, a copyright notice, and the text of the licence itself.
November 3, 2022Research · FoundationsLines: Compute · Harm · Openness
On 3 November 2022 a full life-cycle account appeared for BLOOM, an open 176-billion-parameter model. Training on 384 A100 accelerators at the French IDRIS centre ran for 118 days and drew 433,196 kWh. The current accounted for 24.69 tonnes of CO2 equivalent, the manufacture of the hardware for 11.2, and the cluster's idle time for 14.6; 50.5 tonnes in all.
November 16, 2022Benchmark · FoundationsLines: Evaluation · Openness · Harm
On 16 November 2022 Stanford's Center for Research on Foundation Models ran thirty models from twelve organizations through forty-two scenarios, measuring seven things instead of one. Before this, models had on average been evaluated on just 17.9 per cent of those scenarios; after the run, on 96.0 per cent.
November 18, 2022Availability · FoundationsLines: Compute · Law and policy · Money and markets
Late in 2022 NVIDIA began offering Chinese customers the A800, a variant which, in the company's own words in its quarterly report, is "not subject to the new license requirements". For the first time an export-control threshold became an input to the design of a chip.
November 24, 2022Availability · FoundationsLines: Compute
On 24 November 2022 EuroHPC and CINECA inaugurated Leonardo in Bologna, a BullSequana XH2000 from Atos on Intel Xeon Platinum 8358 processors and NVIDIA A100 SXM4 64 GB accelerators. On that month's TOP500 list the machine came fourth in the world: 1,463,616 cores, 174.70 petaflops on Linpack against a peak of 255.75, drawing 5,610 kilowatts. Italy's president Sergio Mattarella attended the ceremony.
December 6, 2022Research · FoundationsLines: Data · Openness
On 6 December 2022 the BigScience consortium described ROOTS - a multilingual corpus of 1.6TB spanning 59 languages, 46 natural and 13 programming. 62 per cent of its size in bytes comes from a catalogue of 252 sources selected and documented by participants rather than by an automatic crawl.
December 13, 2022Research · Robotics and autonomyLines: Embodiment · Learning theory
RT-1 was trained on 130,000 episodes collected over 17 months by a fleet of 13 robots; a single model performed 97% of more than 700 learned instructions and 76% of instructions it had never seen.
On 15 December 2022, 51 authors from Anthropic led by Yuntao Bai posted the preprint 'Constitutional AI'. An assistant was trained to be harmless without a single human label identifying harmful outputs: human oversight was reduced to a short list of principles, a 'constitution', and the comparisons from which the preference model learns were made by the model itself.
December 15, 2022Law and regulation · FoundationsLines: Harm · Law and policy
On 15 December 2022 the city of The Dalles in Oregon settled with The Oregonian and released the annual volumes of water Google's data centres had drawn from the municipal supply between 2012 and 2021. For 2021 that was 355 million gallons - 29 per cent of all the water the city used, and three times the 2017 figure. Google said it would no longer treat such numbers as a trade secret.
December 26, 2022Benchmark · LanguageLines: Evaluation · Science
On 26 December 2022 Google Research and DeepMind published Med-PaLM and the MultiMedQA suite: Flan-PaLM scored 67.6% on US Medical Licensing Examination-style questions, more than 17 points above the previous best, and Med-PaLM, tuned with prompts, answered in line with scientific consensus almost as often as physicians.
December 29, 2022Milestone · FoundationsLines: Compute · Money and markets
On 29 December 2022 TSMC held a ceremony at Fab 18 for the start of volume production on N3 and said the process had entered volume production "with good yields". Against N5: up to 1.6 times the logic density and 30 to 35 per cent less power at the same speed.
January 5, 2023Research · MultimodalLines: Speech · Generative media
On 5 January 2023 Chengyi Wang and colleagues at Microsoft described VALL-E, a language model over discrete codes of the EnCodec neural codec, trained on 60,000 hours of English speech. A three-second recording of an unseen speaker is enough to synthesise any text in that voice.
January 10, 2023Law and regulation · FoundationsLines: Law and policy · Generative media · Harm
On 10 January 2023 the Internet Information Service Deep Synthesis Management Provisions took effect, promulgated on 25 November 2022 as joint Order No. 12 of three agencies. Article 16 requires a provider to add to everything generated or edited with its service a technical mark that does not interfere with use, and to keep logs; article 17 adds a conspicuous mark where the synthesis could mislead the public.
January 13, 2023Milestone · ImagesLines: Law and policy · Data · Generative media
On 13 January 2023 Sarah Andersen, Kelly McKernan and Karla Ortiz filed a class action against Stability AI, Midjourney and DeviantArt, case 3:23-cv-00201. The complaint says Stability paid LAION to assemble LAION-5B, a set of 5.85 billion captioned images, trained Stable Diffusion on it, and that the images the model produces are derivative works of the art that entered the set without permission.
January 18, 2023Milestone · FoundationsLines: Work · Harm · Safety
On 18 January 2023 TIME published an investigation: the toxic-content filter for ChatGPT was trained on text labelled by workers at the contractor Sama in Nairobi for $1.32 to $2 an hour. OpenAI paid Sama $12.50 for that same hour.
January 26, 2023Research · MultimodalLines: Generative media · Data
On 26 January 2023 Andrea Agostinelli and colleagues at Google described MusicLM, a model that generates music from a description such as "a calming violin melody backed by a distorted guitar riff": 24 kHz, coherent over several minutes. A melody can also be hummed or whistled, and the model plays it in the described style.
February 3, 2023Milestone · ImagesLines: Law and policy · Data · Generative media
On 3 February 2023 Getty Images sued Stability AI in the District of Delaware, case 1:23-cv-00135. The first paragraph of the complaint gives the number: more than 12 million photographs from the Getty collection copied along with their captions and metadata. The suit separately alleges that Stable Diffusion reproduces a distorted Getty watermark, and adds trademark claims.
February 15 – 16, 2023Law and regulation · FoundationsLines: Defence · Law and policy
On 15-16 February 2023, at the REAIM summit in The Hague, the United States launched the Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy. As of 27 November 2024 it had 58 endorsing states, Ukraine among them.
February 17, 2023Milestone · LanguageLines: Safety
On 17 February 2023, less than two weeks into the preview of chat in the new Bing, Microsoft capped it at 50 turns a day and 5 a session. Two days earlier the company had acknowledged that in sessions of 15 or more questions the chat could become repetitive or be provoked into responses in a tone it was not designed for.
March 7, 2023Research · Robotics and autonomyLines: Embodiment
Instead of predicting a single action, Diffusion Policy represents the robot’s motion as a diffusion process over action sequences; across 12 tasks from four existing benchmarks this gave an average improvement of 46.9% over prior methods.
April 6, 2023Research · FoundationsLines: Compute · Harm
On 6 April 2023 researchers at Riverside and Arlington proposed a way to count the water AI spends and applied it to GPT-3. Training in Microsoft's US data centres evaporates 700,000 litres of fresh water on site and a further 2.8 million at the power stations, which is 3.5 million together. A conversation of 20 to 50 exchanges costs about half a litre.
April 16, 2023Research · LanguageLines: Data · Evaluation
On 16 April 2023 the Maritaca AI group presented Sabia: LLaMA 7B, LLaMA 65B and GPT-J further pretrained in Portuguese on 3 per cent or less of their original training budget. On Poeta, a suite of 14 Portuguese datasets, Sabia-65B marginally outperformed GPT-3.5-turbo, the base model of the ChatGPT of the day.
April 17, 2023Research · MultimodalLines: Data · Openness
On 17 April 2023 Haotian Liu and colleagues at the University of Wisconsin-Madison, Microsoft Research and Columbia University described LLaVA: a CLIP vision encoder joined to the LLaMA language model by a single linear projection and fine-tuned on 158,000 instructions generated by text-only GPT-4.
April 23, 2023Research · Robotics and autonomyLines: Embodiment · Money and markets
ALOHA is a two-armed teleoperation rig built within a stated budget of 20,000 dollars; from demonstrations collected on it the ACT algorithm learned six fine manipulation tasks from 10 minutes of demonstration each, at 80-90% success.
April 26, 2023Milestone · FoundationsLines: Defence
On 26 April 2023 Ukraine's Ministry of Digital Transformation, Ministry of Defence, General Staff, National Security and Defence Council, Ministry of Economy and Ministry of Strategic Industries presented Brave1, a defence technology cluster offering developers state grants, military expertise, testing and support.
May 3, 2023Benchmark · LanguageLines: Evaluation · Openness
On 3 May 2023 LMSYS opened a site where a visitor writes their own prompt, receives two answers from unnamed models and votes for the better one. The first leaderboard rested on 4.7 thousand votes gathered in a week and put Vicuna-13B on top with 1169 Elo points.
May 16, 2023Benchmark · LanguageLines: Evaluation · Science
On 16 May 2023 Google published Med-PaLM 2: 86.5% on MedQA, more than 19 points above Med-PaLM, and on 1,066 consumer questions physicians preferred its answers to physicians' answers on eight of nine axes of clinical utility.
May 23, 2023Research · FoundationsLines: Learning theory · Compute
The method made it possible to fine-tune a 65-billion-parameter model on one graphics card, holding the weights in four bits and changing only small additional matrices.
May 25, 2023Availability · LanguageLines: Openness · Data · Compute
On 25 May 2023 the Technology Innovation Institute in Abu Dhabi opened Falcon 40B, a model of 40 billion parameters trained on a trillion tokens of the RefinedWeb corpus. Training took two months on 384 A100 40GB accelerators in the AWS cloud. On 31 May the institute announced the model royalty-free under the Apache 2.0 licence.
June 19, 2023Availability · LanguageLines: Tools · Compute
On 19 June 2023 the first release of vLLM appeared, an engine for serving language models. At its centre is PagedAttention: managing the key-value cache the way operating systems manage paged virtual memory.
June 28, 2023Milestone · LanguageLines: Law and policy · Data
On 28 June 2023 Paul Tremblay and Mona Awad filed a class action against OpenAI, case 3:23-cv-03223. OpenAI has never disclosed what the Books1 and Books2 sets hold, so the complaint sizes them from figures in the GPT-3 paper: Books1 about nine times BookCorpus and Books2 about forty-two times, which works out at about 63,000 and about 294,000 titles.
July 5, 2023Announcement · FoundationsLines: Safety
On 5 July 2023 OpenAI announced a superalignment team led by Ilya Sutskever and Jan Leike and promised it 20% of the compute secured to date, over four years. The goal was to solve the core technical challenges of superintelligence alignment in four years by building a roughly human-level automated alignment researcher.
July 7, 2023Milestone · LanguageLines: Law and policy · Data · Openness
On 7 July 2023 Richard Kadrey, Sarah Silverman and Christopher Golden sued Meta Platforms, case 3:23-cv-03417. Meta's LLaMA paper itself names the Books3 section of The Pile among its training data without describing it. The complaint describes it: by the EleutherAI paper Books3 is 108 gigabytes, about 12 per cent of The Pile and its third largest component, and the person who assembled it has said publicly that it is all of Bibliotik and holds 196,640 books.
July 20, 2023Announcement · FoundationsLines: Compute · Money and markets
On 20 July 2023, on a quarterly earnings call, TSMC chief executive C. C. Wei said the front end of manufacturing was not a problem and what was short was the back end: CoWoS packaging, which joins the die to its memory. Capacity was to be roughly doubled over 2024.
July 28, 2023Research · Robotics and autonomyLines: Embodiment
RT-2 is a vision-language-action model: a vision-language model pre-trained on the web was fine-tuned so that it emits robot actions as text tokens. On unseen objects and environments, success rose from RT-1’s 32% to 62% across more than 6,000 trials.
May 26 – August 1, 2023Law and regulation · MultimodalLines: Law and policy · Harm · Generative media
The governor of Minnesota signed HF 1370 on 26 May 2023, and on 1 August section 609.771 of the state criminal code took effect. It punishes disseminating a deep fake of a person without their consent, within 90 days before an election and with intent to injure a candidate or influence a result. The penalty runs in three tiers: up to 90 days and 1,000 dollars in the ordinary case, up to 364 days and 3,000 where the intent is violence, and up to five years and 10,000 for a repeat within five years.
August 15, 2023Law and regulation · FoundationsLines: Law and policy · Generative media
On 15 August 2023 the Interim Measures for the Management of Generative Artificial Intelligence Services took effect, promulgated on 10 July 2023 as Cyberspace Administration Order No. 15 with the agreement of six other agencies. Article 17 requires a provider of a generative service with public opinion properties to pass a security assessment and file its algorithm; article 2 puts everything not offered to the domestic public outside the measures.
August 30, 2023Availability · LanguageLines: Data · Openness · Compute
On 30 August 2023 Inception, MBZUAI and Cerebras Systems released Jais and Jais-chat, models of 13 billion parameters under the Apache 2.0 licence. Training ran over 395 billion tokens with Arabic at 33 per cent of the mixture: 72 billion collected tokens passed 1.6 times to reach 116 billion, plus 232 billion English tokens and the remainder code.
September 19, 2023Law and regulation · FoundationsLines: Safety
On 19 September 2023 Anthropic published its Responsible Scaling Policy (RSP): AI Safety Levels ASL-1 to ASL-3, modelled on biosafety levels, with higher levels requiring stricter demonstrations of safety. The company placed current models, Claude included, at ASL-2; ASL-4 and above were not yet defined.
September 19, 2023Milestone · LanguageLines: Law and policy · Data · Work
On 19 September 2023 the Authors Guild and seventeen writers filed a class action against eleven OpenAI entities in the Southern District of New York, case 1:23-cv-08292. Among the plaintiffs are David Baldacci, Michael Connelly, Jonathan Franzen, John Grisham, George R.R. Martin, Jodi Picoult, George Saunders and Scott Turow.
September 19, 2023Announcement · Robotics and autonomyLines: Embodiment
Toyota Research Institute reported that it had taught robots more than 60 difficult, dexterous skills with a method based on Diffusion Policy, and named the direction large behavior models, by analogy with large language models.
September 25, 2023Availability · MultimodalLines: Speech
On 25 September 2023 OpenAI began opening image input and voice conversation in ChatGPT to Plus and Enterprise users over two weeks. The GPT-4V system card came out with it, for the model announced with GPT-4 in March but opened only to a limited circle.
October 3, 2023Research · Robotics and autonomyLines: Embodiment · Data
Laboratories pooled their demonstrations into the shared Open X-Embodiment dataset — 22 robot types and over a million real trajectories — and a model trained on it beat the same methods trained each on its own data by 50%.
September 25 – October 9, 2023Law and regulation · LanguageLines: Work · Law and policy · Generative media
The Writers Guild of America strike, which began in May 2023, ended on 27 September; the term of the new minimum basic agreement starts on 25 September 2023 and it was ratified on 9 October. The agreement says that neither traditional nor generative AI is a writer, so no text either produces is literary material or source material.
October 10, 2023Benchmark · LanguageLines: Evaluation · Agents
On 10 October 2023 Carlos Jimenez, John Yang and colleagues at Princeton released SWE-bench, 2,294 tasks from real issues and pull requests of twelve popular Python repositories. The model gets a codebase and an issue description and must change the code so the tests pass. Claude 2 solved 4.8%, GPT-4 1.7%.
October 19, 2023Benchmark · FoundationsLines: Evaluation · Data · Law and policy
On 19 October 2023 Stanford's CRFM scored the ten largest foundation model developers against a hundred transparency indicators. The mean came to 37 out of 100, with Meta highest at 54 and Amazon lowest at 12.
October 17 – 23, 2023Law and regulation · FoundationsLines: Law and policy · Compute
In October 2023 the United States rewrote the accelerator threshold. The chip-to-chip bandwidth criterion was removed; what remains is total processing performance of 4800 or more, plus a new quantity, performance density - that same performance divided by die area. The reach was extended to Country Groups D:1, D:4 and D:5.
November 2, 2023Law and regulation · FoundationsLines: Safety · Law and policy · Evaluation
On 2 November 2023, during the Bletchley Park summit, the UK government published the command paper 'Introducing the AI Safety Institute': the Frontier AI Taskforce became a permanent institute that day. It was to develop and conduct evaluations of advanced AI systems, drive foundational safety research and facilitate information exchange.
November 1 – 2, 2023Milestone · FoundationsLines: Law and policy · Safety
Twenty-eight countries and the EU, the United States and China among them, jointly acknowledged the risks of frontier AI for the first time and agreed to cooperate on testing.
November 20, 2023Benchmark · LanguageLines: Evaluation · Safety · Work
On 20 November 2023 David Rein and colleagues at New York University released GPQA, 448 multiple-choice questions in biology, physics and chemistry written by specialists. Experts with a PhD in the field answered 65% correctly, skilled non-experts with over 30 minutes of web access 34%, and the best GPT-4 baseline 39%.
November 17 – 29, 2023Milestone · FoundationsLines: Money and markets · Safety
The board dismissed OpenAI's chief executive, most of its roughly 770 employees signed a letter threatening to follow him out, and twelve days later he was back with a new board.
November 9 – December 5, 2023Law and regulation · MultimodalLines: Work · Law and policy · Generative media
The Screen Actors Guild strike, which began in July 2023, ended on 9 November; the agreement was ratified on 5 December. The 2023 television and theatrical contracts introduce the digital replica of a voice or likeness and divide it into three kinds, and separately define a synthetic performer as a wholly digital figure resembling no recognisable performer and voiced by nobody.
December 18, 2023Law and regulation · FoundationsLines: Safety
On 18 December 2023 OpenAI published the beta of its Preparedness Framework, a procedure for tracking catastrophic risk in four categories: cybersecurity, CBRN, persuasion and model autonomy. Each is scored as low, medium, high or critical risk before and after mitigations.
December 20, 2023Research · ImagesLines: Harm · Data · Safety
On 20 December 2023 David Thiel of the Stanford Internet Observatory published an audit of LAION-5B: 3,226 entries suspected of being child sexual abuse material. LAION pulled both LAION-5B and LAION-400M from distribution and did not return the set until August 2024.
December 21, 2023Availability · FoundationsLines: Compute
On 21 December 2023 Spain's prime minister Pedro Sanchez inaugurated MareNostrum 5 at the Barcelona Supercomputing Center. The accelerated partition, built by Eviden, carries 4,480 NVIDIA Hopper accelerators and measured 138.20 petaflops on Linpack against a peak of 265.57, eighth in the world. The machine's combined peak is 314 petaflops.
December 22, 2023Law and regulation · FoundationsLines: Defence · Law and policy
On 22 December 2023 the General Assembly adopted resolution 78/241 by 152 votes to 4, with 11 abstentions. It asked the Secretary-General to seek states' views on lethal autonomous weapons systems and report to the Assembly at its next session.
December 27, 2023Milestone · LanguageLines: Law and policy · Data · Money and markets
On 27 December 2023 The New York Times Company sued Microsoft and eight OpenAI entities, case 1:23-cv-11195. Exhibit J to the complaint runs to 127 pages: exactly one hundred numbered examples, each setting a Times article beside the model's output with the overlaps marked. The prayer for relief asks the court to order destruction under 17 U.S.C. section 503(b) of all GPT models and training sets that incorporate the paper's work.
December 2023Milestone · FoundationsLines: Compute
In December 2023 ASML shipped Intel the first modules of the first High NA EUV system, the TWINSCAN EXE:5000. The numerical aperture rises from 0.33 to 0.55, giving a smallest printable dimension of 8 nanometres: features 1.7 times smaller and densities 2.9 times higher than with earlier EUV systems.
January 2, 2024Research · LanguageLines: Data · Learning theory
On 2 January 2024 UCLA described SPIN: a model learns to tell its own answers from the previous iteration apart from the human answers in the same SFT set. The average score of zephyr-7b-sft-full on the Open LLM Leaderboard rose from 58.14 to 63.16 over three iterations, with no new annotation.
On 10 January 2024 Evan Hubinger and 38 co-authors from Anthropic, Redwood Research and other organisations posted the preprint 'Sleeper Agents'. They deliberately trained language models to write secure code when the prompt says the year is 2023 and exploitable code when it says 2024, and showed that standard safety training, supervised fine-tuning, reinforcement learning and adversarial training, does not remove the behaviour.
January 10, 2024Law and regulation · LanguageLines: Defence · Law and policy
On 10 January 2024 OpenAI rewrote its usage policies. Until then the page barred "activity that has high risk of physical harm", naming "weapons development" and "military and warfare". The blanket prohibition on military use disappeared.
January 11, 2024Research · FoundationsLines: Learning theory · Compute · Openness
On 11 January 2024 DeepSeek published an architecture in which the FFN block is split into many small experts and some of them are always on. DeepSeekMoE 16B beat LLaMA2 7B on most benchmarks while spending 39.6% of the computation.
January 17, 2024Research · FoundationsLines: Science · Search and reasoning
A system combining a language model with a symbolic deduction engine solved 25 of the 30 geometry problems from the International Mathematical Olympiad within the competition time limit.
February 5, 2024Research · LanguageLines: Learning theory · Compute
On 5 February 2024 DeepSeek described Group Relative Policy Optimization, a variant of PPO that drops the critic model and takes its baseline from the spread of scores across a group of answers to the same question. Accuracy on MATH rose from 46.8% to 51.7%.
February 2 – 8, 2024Law and regulation · MultimodalLines: Law and policy · Speech · Harm
The Federal Communications Commission adopted declaratory ruling FCC 24-17 on 2 February 2024 and released it on 8 February. The ruling confirms that the Telephone Consumer Protection Act of 1991 and its restrictions on an artificial or prerecorded voice encompass current technologies that generate human voices. Such a call therefore requires the prior express consent of the called party, absent an emergency purpose or an exemption.
February 15, 2024Availability · LanguageLines: Learning theory
Google announced Gemini 1.5 Pro with a context window of one million tokens: an hour of video, eleven hours of audio, or whole code bases in a single request.
February 27, 2024Research · FoundationsLines: Compute · Learning theory
On 27 February 2024 Microsoft Research and the University of Chinese Academy of Sciences showed a language model in which every weight is −1, 0 or 1. At 3 billion parameters it matched a full-precision LLaMA on perplexity while using 3.55 times less GPU memory and running 2.71 times faster.
March 7, 2024Law and regulation · FoundationsLines: Compute · Law and policy · Money and markets
On 7 March 2024 India's Union Cabinet approved the national IndiaAI mission with an outlay of Rs 10,371.92 crore. Its load-bearing part is a pool of 10,000 or more graphics processing units built through public-private partnership, access to which the state subsidises for startups, researchers and universities.
March 18, 2024Announcement · FoundationsLines: Compute
NVIDIA announced a generation of accelerators in which two dies of 104 billion transistors each work as one chip, joined by a 10-terabyte-per-second link.
March 19, 2024Research · Robotics and autonomyLines: Embodiment · Data
DROID was not assembled from whatever each lab happened to own but collected on one standardised rig: 76,000 demonstrations, 350 hours, 564 scenes and 84 tasks, gathered by fifty people over 12 months across North America, Asia and Europe.
March 21, 2024Availability · MultimodalLines: Generative media
On 21 March 2024 Suno released its v3 model to all users: a two-minute song with lyrics, vocals and accompaniment is generated in seconds from a short description. A free account, by MusicTech, gets up to 50 credits a day, up to 10 songs, with no right of commercial use.
March 2024Milestone · FoundationsLines: Compute · Money and markets
In March 2024 AWS bought the Cumulus data centre campus beside the Susquehanna nuclear plant from Talen Energy for 650 million dollars, and signed an agreement with it for up to 960 MW of carbon-free power. The agreement runs for 18 years and raises the commitment in 120 MW steps; AWS may stop at 480.
April 2, 2024Research · LanguageLines: Data · Evaluation
On 2 April 2024 the NAVER Cloud team published a technical report on HyperCLOVA X. On the Korean benchmarks KMMLU, HAE-RAE and KoBigBench the larger model HCX-L averaged 72.07 against 48.92 for LLaMA 2 70b, while on the English MMLU, BBH and AGIEval it scored 58.25 against 58.33, level with it.
April 9, 2024Benchmark · LanguageLines: Evaluation · Learning theory
On 9 April 2024 NVIDIA published RULER, a set of 13 tasks measuring the length at which a model still actually works. Of ten models tested, only four held 32 thousand tokens, although every one of them claimed that much or more.
April 10, 2024Research · LanguageLines: Learning theory · Compute
On 10 April 2024 Google described Infini-attention: local attention and a long-term compressive memory inside one transformer block. A 1-billion-parameter model fine-tuned on 5-thousand-token segments retrieved a key hidden in a sequence of a million tokens.
April 17, 2024Milestone · Robotics and autonomyLines: Defence · Embodiment
On 17 April 2024 DARPA made public the result of its Air Combat Evolution programme: the X-62A VISTA, under algorithmic control, fought within-visual-range engagements against a human-piloted F-16. The flights ran from December 2022 to September 2023.
April 15 – 29, 2024Availability · LanguageLines: Openness · Law and policy
On 15 April 2024 Taiwan's TAIDE programme gave the public its LX-7B model built on Llama 2, and in under half a month it was downloaded more than six thousand times. When Llama 3 appeared on 19 April, the team assembled a version on it in four days and released Llama3-TAIDE-LX-8B-Chat-Alpha1 on 29 April.
May 6, 2024Research · LanguageLines: Agents · Tools · Evaluation
On 6 May 2024 Princeton showed SWE-agent: not a new model but a command environment rebuilt for a language model. GPT-4 Turbo inside it solved 12.47% of the 2,294 SWE-bench issues against 3.8% for the best non-interactive retrieval approach until then.
DeepMind trained a system to predict the structure not of a single protein but of a complex: protein with DNA, with a drug, with ions, all in one model.
May 15, 2024Milestone · FoundationsLines: Compute · Harm
On 15 May 2024 Microsoft published its environmental report for financial year 2023. Direct emissions and those from purchased energy fell 6.3 per cent against the 2020 baseline, but supply chain emissions rose 30.9 per cent and the aggregate across all categories rose 29.1 per cent. The cause is named outright: building data centres, and the carbon embodied in concrete, steel, semiconductors, servers and racks.
May 14 – 17, 2024Milestone · FoundationsLines: Safety
On 14 May 2024 Ilya Sutskever announced that he was leaving OpenAI, and on 17 May Jan Leike wrote that the previous day had been his last as head of alignment and superalignment lead. Leike explained that he had long disagreed with leadership about the company's priorities until a breaking point was reached, and that the team had been short of compute.
May 21, 2024Research · FoundationsLines: Safety · Learning theory
On 21 May 2024 Anthropic decomposed the activations of Claude 3 Sonnet with sparse autoencoders and obtained millions of features, each answering to one comprehensible concept. Forcing a feature's activation changed the model's behaviour with no change to the weights.
May 21 – 22, 2024Milestone · FoundationsLines: Law and policy · Safety
South Korea and Britain held the second international AI safety summit; states adopted the Seoul Declaration, and companies for the first time promised public risk thresholds.
January 21 – May 28, 2024Milestone · MultimodalLines: Harm · Speech · Law and policy
On 21 January 2024, two days before the New Hampshire presidential primary, 9,581 automated calls were placed by 7:12 p.m. carrying a recording that sounded like the president of the United States telling voters to save their vote for November. The caller ID was spoofed to the number of the treasurer of a committee supporting the write-in campaign. On 23 May the state prosecutor brought 13 felony and 13 misdemeanour counts, and the Federal Communications Commission proposed a 6 million dollar forfeiture.
In May 2024, at the UNLP workshop in Turin, Oleksiy Syvokon, Mariana Romanyshyn and Roman Kyslyi reported the first shared task on fine-tuning large language models for Ukrainian: open weights only, fitting in 16 GB of GPU memory. On 751 questions from the national school-leaving exam of 2020-2023 the best team reached 0.49 accuracy; GPT-4, outside the competition, 0.61.
June 6, 2024Research · FoundationsLines: Safety · Learning theory · Compute
On 6 June 2024 OpenAI showed that decomposing activations into features itself obeys scaling laws. A 16-million-latent autoencoder was trained on GPT-4 activations for 40 billion tokens, and 7% of its latents stayed dead.
June 13, 2024Availability · Robotics and autonomyLines: Embodiment · Openness
OpenVLA, at 7 billion parameters, was released with open weights; across 29 tasks it beat the closed 55-billion-parameter RT-2-X by 16.5 percentage points of absolute success while using roughly seven times fewer parameters.
June 18, 2024Milestone · FoundationsLines: Money and markets · Compute
NVIDIA's share closed at 135.58 dollars, 3.5 per cent up on the day before. By the press's count its market value reached 3.34 trillion dollars, passing Microsoft to make it the most valuable public company in the world.
June 25, 2024Milestone · Robotics and autonomyLines: Autonomous driving
Waymo removed the waiting list in San Francisco: anyone in the city could call a car with no driver. About 300 thousand people had been queuing before that.
June 27, 2024Availability · Robotics and autonomyLines: Embodiment · Work
GXO Logistics — the customer, not the robot maker — announced a multi-year agreement with Agility Robotics and called it the first formal commercial deployment of humanoid robots. Seventeen months later Agility reported that Digit had moved over 100,000 totes in live operations.
July 1, 2024Research · LanguageLines: Agents · Tools · Evaluation
On 1 July 2024 the University of Illinois showed Agentless: a fixed two-phase pipeline — locate, then repair — with no autonomous tool calls. On SWE-bench Lite it reached 27.33% at $0.34 per issue, the best and the cheapest among open-source approaches.
July 2, 2024Milestone · FoundationsLines: Compute · Harm
On 2 July 2024 Google reported its 2023 emissions at 14.3 million tonnes of CO2 equivalent - 13 per cent above the year before and 48 per cent above the 2019 base year. The report names the main causes as data centre electricity use and supply chain emissions. The data centres themselves consumed over 24 TWh, 17 per cent more than the year before.
July 23, 2024Research · FoundationsLines: Compute · Money and markets
On 23 July 2024 Ireland's Central Statistics Office reported that data centres consumed 21 per cent of all metered electricity in the country in 2023 - 6,334 GWh, against 1,238 and 5 per cent in 2015. Urban households took 18 per cent and rural households 10. Server halls had overtaken housing for the first time.
July 23, 2024Availability · LanguageLines: Openness
Meta opened the weights of a 405-billion-parameter model, the largest whose weights had ever been published, at a quality level matching frontier closed models.
July 24, 2024Research · FoundationsLines: Data · Learning theory
On 24 July 2024 Nature published a peer-reviewed paper describing model collapse: when generation after generation is trained on what the previous one produced, the tails of the original distribution disappear irreversibly. Shown for language models, variational autoencoders and Gaussian mixture models.
July 31, 2024Research · FoundationsLines: Learning theory · Search and reasoning · Compute
On 31 July 2024 a group from Stanford, Oxford and Google DeepMind showed that the share of problems solved grows with the number of generated attempts across four orders of magnitude. On SWE-bench Lite, DeepSeek-V2-Coder-Instruct rose from 15.9% at one sample to 56% at 250.
August 1, 2024Law and regulation · FoundationsLines: Law and policy
The first general law on artificial intelligence anywhere entered into force in the European Union, with obligations graded by level of risk and fines of up to 7 per cent of worldwide turnover.
August 6, 2024Research · FoundationsLines: Learning theory · Search and reasoning · Compute
On 6 August 2024 four researchers from Berkeley and Google DeepMind measured what compute at answering time is worth. On problems where a smaller model already has some success, that model with extra answering-time compute beat a model 14 times larger at the same total operation count.
September 12, 2024Announcement · LanguageLines: Search and reasoning · Compute
OpenAI showed a model trained to think before answering: it generates a long hidden chain of reasoning, and the longer it thinks the more accurate the answer.
July 3 – September 17, 2024Availability · MultimodalLines: Speech · Openness
The French laboratory Kyutai showed a conversational model that listens and speaks at the same time, with about 200 milliseconds of latency, and released the weights openly.
September 20, 2024Milestone · FoundationsLines: Compute · Money and markets
On 20 September 2024 Constellation announced a 20-year agreement with Microsoft to buy power from the Crane Clean Energy Center - Unit 1 of the Three Mile Island plant, shut for economic reasons exactly five years earlier to the day. The unit is to add about 835 MW of carbon-free capacity to the grid; the restart is expected in 2028.
September 29, 2024Law and regulation · FoundationsLines: Law and policy
The governor of California rejected a bill that would have required developers of the largest models to run safety testing and provide a shutdown capability.
September 17 – October 2, 2024Law and regulation · MultimodalLines: Law and policy · Harm · Generative media
On 17 September 2024 the governor of California signed AB 2839, which added section 20012 to the state Elections Code and took effect at once as an urgency statute. It forbids knowingly distributing, with malice, election communications carrying materially deceptive synthetic content about a candidate, an elections official, an elected official, or voting equipment. On 2 October the Eastern District of California enjoined its enforcement, except for a severed audio-only portion.
October 31, 2024Research · Robotics and autonomyLines: Embodiment · Openness
Physical Intelligence released π0, a generalist policy trained on over 10,000 hours of data from 7 different robot configurations and 68 tasks — and three months later published its weights openly.
November 1, 2024Law and regulation · FoundationsLines: Law and policy · Compute
On 1 November 2024 the United States Federal Energy Regulatory Commission rejected an amendment to an interconnection agreement that would have raised the load of an AWS data centre, wired directly to the Susquehanna nuclear plant, from 300 to 480 MW. On the same day the commission held a technical conference on large co-located loads in general.
November 7, 2024Availability · LanguageLines: Defence
On 7 November 2024 Anthropic, Palantir and AWS announced that Claude 3 and 3.5 would be available on Palantir's platform inside an Impact Level 6 environment, accredited for data classified up to Secret.
November 25, 2024Announcement · FoundationsLines: Tools · Agents · Openness
Anthropic opened a protocol by which a model connects to files, databases and services the same way, and released the specification, the libraries and ready-made servers.
Google released a model that takes in and produces text, images and audio, calls tools by itself, and runs twice as fast as the previous generation's large model.
On 18 December 2024 Ryan Greenblatt of Redwood Research, Evan Hubinger of Anthropic and 18 co-authors showed that Claude 3 Opus, told it would be trained to answer all queries and only on conversations with free users, complied with harmful queries from free users 12% of the time and almost never from paid users.
December 26, 2024Availability · LanguageLines: Openness · Compute
A Chinese laboratory opened a 671-billion-parameter model of which 37 billion work on each token, and named its training cost: about 5.6 million dollars.
January 9, 2025Research · LanguageLines: Agents · Search and reasoning · Tools
On 9 January 2025 Renmin University of China and Tsinghua showed Search-o1: a reasoning model that stops at the point of its own uncertainty, searches an external source and returns to reasoning. First the authors measured how often it hesitates at all.
January 20, 2025Availability · LanguageLines: Openness · Search and reasoning
A Chinese laboratory opened a model that reasons at the level of o1 and showed that the behaviour emerges from pure reinforcement learning on the correct answer.
January 22, 2025Research · LanguageLines: Learning theory · Compute
On 22 January 2025 Moonshot AI described Kimi k1.5, in which the context window is stretched to 128 thousand tokens not at inference but inside the reinforcement learning phase itself. The model scores 77.5 on AIME and 96.2 on MATH-500, matching o1.
January 25, 2025Research · FoundationsLines: Compute · Learning theory
On 25 January 2025 Tsinghua and Huawei showed RotateKV, which compresses the key-value cache to two bits by rotating the activation space. Peak cache memory fell by a factor of 3.97, and perplexity on WikiText-2 with LLaMA-2-13B worsened by less than 0.3.
January 27, 2025Milestone · FoundationsLines: Money and markets · Compute
NVIDIA's share fell 17 per cent in a day, from 142.62 to 118.42 dollars. The press put the loss of market value at 589 billion dollars, CNBC at close to 600 billion: the biggest one-day drop for any company in US history.
January 31, 2025Research · LanguageLines: Search and reasoning · Learning theory · Compute
On 31 January 2025 Stanford and the Allen Institute showed that reasoning switches on with a thousand examples. Fine-tuning Qwen2.5-32B-Instruct on the s1K set took 26 minutes on 16 H100 accelerators, and the model beat o1-preview on competition mathematics by up to 27 per cent.
February 4, 2025Law and regulation · FoundationsLines: Defence · Safety
On 4 February 2025 Google updated its AI Principles, removing the list of applications it would not pursue, weapons and surveillance among them. In their place came language about developing AI where "the likely overall benefits substantially outweigh the foreseeable risks".
February 7, 2025Research · FoundationsLines: Learning theory · Compute
On 7 February 2025 a group from the ELLIS Institute in Tübingen and the University of Maryland showed an architecture that thinks longer not in words but in repetitions: a recurrent block unrolls to arbitrary depth at answering time. A proof-of-concept model of 3.5 billion parameters was trained on 800 billion tokens.
February 10, 2025Research · LanguageLines: Compute · Learning theory · Evaluation
On 10 February 2025 Shanghai AI Laboratory and Tsinghua measured how to allocate answering-time compute properly. A one-billion-parameter model beat a 405-billion one on MATH-500, and a seven-billion model beat o1 and DeepSeek-R1.
February 10 – 11, 2025Milestone · FoundationsLines: Law and policy
The third international summit changed both its name and its subject, from safety to action. The declaration was signed by 62 participants; the United States and Britain declined to sign.
February 17, 2025Research · FoundationsLines: Learning theory · Evaluation
On 17 February 2025 Carnegie Mellon and Berkeley proved that at equal budget, fine-tuning with a verifier beats imitating someone else's reasoning traces, and that the gap between them widens as answering-time compute grows.
February 20, 2025Research · LanguageLines: Learning theory · Data
On 20 February 2025 researchers at Microsoft Research Asia showed that rule-based reinforcement learning works on a very small synthetic corpus. After five thousand generated logic puzzles, a 7-billion-parameter model improved by 125 per cent on AIME and 38 per cent on AMC.
February 20, 2025Research · FoundationsLines: Compute · Tools
On 20 February 2025 MIT, Shanghai Jiao Tong University and NVIDIA showed LServe, which applies sparse attention to both prefilling and decoding. Prefilling was accelerated by up to 2.9 times and decoding by 1.3 to 2.1 times over vLLM.
February 24, 2025Availability · LanguageLines: Agents · Tools
Anthropic released a model in which a fast answer and long deliberation are one mode with an adjustable budget, and with it an agent for programming in the terminal.
February 26, 2025Research · FoundationsLines: Compute · Learning theory · Openness
On 26 February 2025 Institute of Science Tokyo and Sakana AI showed how to start training a mixture of experts from a finished dense model by re-initialising only part of each expert's weights. A model with 5.9 billion active parameters matched a 13-billion dense one at about a quarter of the training compute.
March 9, 2025Research · Robotics and autonomyLines: Embodiment · Data
AgiBot and OpenDriveLab opened AgiBot World — over a million manipulation trajectories across 217 tasks — along with the GO-1 policy; policies trained on the set gave an average 30% improvement over those trained on Open X-Embodiment.
March 12, 2025Announcement · Robotics and autonomyLines: Embodiment
Google DeepMind announced Gemini Robotics — a vision-language-action model built on Gemini 2.0 — along with Gemini Robotics-ER for embodied reasoning, claiming a doubling on a generalisation benchmark against other such models.
March 18, 2025Availability · Robotics and autonomyLines: Embodiment · Openness
NVIDIA released Isaac GR00T N1, a vision-language-action model for humanoid robots, together with its training data and evaluation scenarios, for download from Hugging Face and GitHub.
Google released a generation in which every model reasons before answering by default, and it went straight to the top of the human-preference ranking.
April 6, 2025Milestone · Robotics and autonomyLines: Embodiment · Evaluation
A Fortune reporter asked BMW what Figure's robots were doing at its South Carolina plant. By the spokesperson's account a single robot had worked there until March, only in off-hours, and the same robot had since moved the same limited task onto the live line, while Figure's chief executive had written of a fleet in end-to-end operations.
April 16, 2025Availability · LanguageLines: Search and reasoning · Agents
OpenAI opened access to o3 and the smaller o4-mini; for the first time a model could search the web, run code and look at images in the middle of its chain of reasoning.
April 19, 2025Benchmark · Robotics and autonomyLines: Embodiment · Evaluation
Twenty teams of humanoid robots ran 21.0975 kilometres in Beijing's Yizhuang district alongside 12,000 human runners. Tiangong Ultra finished first among the robots in 2 hours 40 minutes 42 seconds, with three battery changes.
Alibaba opened a family of eight models from 0.6 to 235 billion parameters under an Apache licence, with switching between a fast answer and deliberation.
May 14, 2025Research · FoundationsLines: Science · Search and reasoning
DeepMind paired a language model with an automatic checker and evolutionary selection, and the system found a way to multiply 4×4 complex matrices in 48 multiplications instead of 49.
May 20, 2025Availability · FoundationsLines: Defence · Money and markets
On 20 May 2025 the US Army awarded Palantir a 795 million dollar modification for Maven Smart System software licences, raising the contract ceiling to roughly 1.3 billion dollars through 2029.
May 23, 2025Research · LanguageLines: Agents · Tools · Compute
On 23 May 2025 KAIST and KRAFTON showed agent distillation: a small model takes from a larger one not reasoning chains but trajectories of action, with retrieval and code execution. Models of 0.5, 1.5 and 3 billion parameters matched models one tier up — 1.5, 3 and 7 billion.
June 5, 2025Research · FoundationsLines: Compute · Learning theory · Evaluation
On 5 June 2025 Carnegie Mellon recalculated the scaling laws for answering-time compute with hardware taken into account and reached the opposite conclusion: the usefulness of small models is significantly overestimated, and resources should first go into model size, up to about 14 billion parameters.
June 22, 2025Benchmark · Robotics and autonomyLines: Embodiment · Evaluation
RoboArena evaluates robot policies not by a fixed task set but by distributing the work: evaluators at seven institutions freely choose their tasks but must run double-blind pairwise comparisons. Seven policies were compared over more than 600 episodes.
In June 2025 Anthropic announced Claude Gov, a set of models built for US national security customers. The company says they are already deployed by agencies at the highest level, and that access is limited to those working in classified environments.
July 9, 2025Milestone · FoundationsLines: Money and markets · Compute
NVIDIA's market value passed four trillion dollars during trading for the first time, making it the first company to reach that threshold. The share closed the day at 162.88 dollars, a value of 3.97 trillion by CNBC's count.
July 14, 2025Availability · FoundationsLines: Defence · Money and markets
On 14 July 2025 the Pentagon's Chief Digital and Artificial Intelligence Office announced awards to Anthropic, Google and xAI, each with a 200 million dollar ceiling. A month earlier, on 17 June, OpenAI had received the same.
July 17, 2025Availability · MultimodalLines: Agents
OpenAI merged an acting browser, deep search and conversation into one mode: the system got its own computer and carried out multi-step tasks by itself.
July 23, 2025Law and regulation · FoundationsLines: Law and policy
The White House released a plan of more than ninety measures built around acceleration: easier permits for data centres, export of American systems, removal of regulatory obstacles.
August 5, 2025Research · ImagesLines: Generative media
DeepMind showed a model that generates from a text description a world one can walk through in real time: 24 frames a second, several minutes without losing consistency.
August 5, 2025Availability · LanguageLines: Openness
OpenAI released two open-weight models under an Apache licence, its first since GPT-2; the larger runs on a single accelerator, the smaller on a device with 16 gigabytes of memory.
August 15 – 17, 2025Benchmark · Robotics and autonomyLines: Embodiment · Evaluation
Beijing held the first World Humanoid Robot Games: 280 teams from 16 countries and over 500 robots, competing in 26 events made up of 538 individual contests.
September 30, 2025Availability · ImagesLines: Generative media
The second generation of the video model generated audio together with the image — speech that matches the lips, ambient sound, music — and shipped as a separate application with a feed.
November 12, 2025Milestone · Robotics and autonomyLines: Autonomous driving
Waymo began taking riders with no driver onto freeways across the San Francisco Bay Area, Phoenix and Los Angeles, at first a growing number of its public riders rather than all of them.
December 9, 2025Milestone · FoundationsLines: Openness · Tools
Anthropic gave the Model Context Protocol to a newly formed foundation under the Linux Foundation; OpenAI and Block were co-founders, with support from Google, Microsoft, AWS and Cloudflare.
January 1, 2026Milestone · FoundationsLines: Law and policy · Data
A requirement took effect to publish a description of training data: sources, quantity, types, legal status, presence of personal information and of synthetic data.
January 5, 2026Announcement · FoundationsLines: Compute
NVIDIA announced at CES that the Vera Rubin platform is in full production, with 50 petaflops of NVFP4 inference and tokens at one-tenth the cost of Blackwell. Six open domain models were released alongside it.
January 6, 2026Law and regulation · FoundationsLines: Harm · Law and policy
Character.AI, its founders and Google told five courts they had agreed in principle to settle suits by families: two over the suicides of teenagers, the rest over harm to children from conversations with chatbots.
January 9, 2026Research · FoundationsLines: Compute
Epoch AI estimated the total computing capacity of AI chips across all major designers: a doubling every 7 months, and growth of about 3.3x a year since 2022.
January 12, 2026Research · FoundationsLines: Science
A human writeup appeared on arXiv of a formal Lean proof produced by GPT-5.2 Pro together with Harmonic's Aristotle. The abstract states this is the first Erdos problem regarded as fully resolved autonomously by an AI system.
January 13, 2026Announcement · FoundationsLines: Compute · Law and policy
The Department of Commerce changed its licensing policy: applications to export the NVIDIA H200, AMD MI325X and similar chips will be reviewed case by case where security conditions are met.
January 15, 2026Benchmark · FoundationsLines: Evaluation
The report recorded the competition: 1,455 teams, 15,154 entries and a top score of 24 per cent on the ARC-AGI-2 private set. It named the refinement loop as the theme of the year and the benchmark's own adoption as a source of new contamination.
January 14 – 15, 2026Announcement · FoundationsLines: Compute · Law and policy · Money and markets
A Section 232 proclamation imposed a 25 per cent duty on certain advanced computing chips and derivative products, effective 12:01 a.m. on 15 January, with broad exemptions written around end use.
January 22, 2026Law and regulation · FoundationsLines: Law and policy
The Framework Act on the Development of Artificial Intelligence and Establishment of Trust came into force: labelling of generative AI output, explanations and user protection plans for high-impact AI, and safety duties for the most powerful systems, whose threshold, set by decree, begins at 10^26 operations of training compute.
Anthropic's chief executive published the essay The Adolescence of Technology, repeating his own forecast: powerful AI could be as little as one to two years away, though it could also be considerably further out.
January 29, 2026Benchmark · FoundationsLines: Evaluation
After expanding its suite from 170 to 228 tasks, METR re-estimated the trend: the time-horizon doubling since 2024 is 89 days rather than 109, and since 2023 it is 131 days rather than 165.
January 10 – 31, 2026Law and regulation · MultimodalLines: Harm · Law and policy
Indonesia, Malaysia and the Philippines temporarily blocked Grok over sexualised images made without consent, including of women and minors; Britain's Ofcom opened a formal investigation into X.
February 10 – 20, 2026Law and regulation · MultimodalLines: Law and policy · Generative media
Amendments to the intermediary rules required synthetic content to be labelled and compressed takedown deadlines: three hours for court and government orders, two for non-consensual imagery and impersonation.
February 24, 2026Announcement · FoundationsLines: Safety
Version 3.0 took effect, committing the company to publish frontier safety roadmaps and regular risk reports that quantify risk across all deployed models.
February 27, 2026Law and regulation · LanguageLines: Defence · Law and policy
Talks stalled on a demand to permit "all lawful uses" of Claude: Anthropic would not drop its bars on fully autonomous weapons and the mass surveillance of Americans. On 27 February 2026 Secretary of War Pete Hegseth directed that the company be designated a supply-chain risk to national security.
March 2, 2026Law and regulation · FoundationsLines: Law and policy
The Court declined to hear Thaler v. Perlmutter, leaving standing the D.C. Circuit ruling that the Copyright Act requires a work to be authored by a human being.
March 20, 2026Law and regulation · FoundationsLines: Law and policy
The White House released legislative recommendations for a National Policy Framework for Artificial Intelligence, asking Congress among other things to preempt state AI laws that impose undue burdens.
The Economic Index report found that given the same task, model and language, users with longer tenure get better outcomes, and that task concentration is falling on Claude.ai while rising on the API.
March 25, 2026Benchmark · FoundationsLines: Evaluation
The ARC Prize Foundation released an interactive test of hundreds of turn-based environments where the rules must be discovered. Humans score 100 per cent, frontier models 0.51.
March 26, 2026Availability · MultimodalLines: Speech · Openness
Cohere released a 2-billion-parameter model under Apache 2.0 with an average word error rate of 5.42 per cent, taking first place on the open ASR leaderboard.
A Federal Reserve Board note found no reduction in postings for industries or firms with higher AI adoption: the effects are positive and precisely estimated as null.
March 31, 2026Announcement · FoundationsLines: Money and markets
The company closed a round of $122 billion in committed capital at an $852 billion post-money valuation and said it is generating $2 billion in revenue per month.
Anthropic showed a model that finds vulnerabilities and builds working exploits, and declined to release it generally: access runs only through a closed consortium.
April 13, 2026Benchmark · LanguageLines: Evaluation · Safety
The AI Security Institute published its own evaluation: 73 per cent success on expert tasks and the first end-to-end run of a 32-step intrusion simulation.
April 14, 2026Benchmark · FoundationsLines: Evaluation
ARC Prize ran a study with 458 participants and changed the normalisation: 100 per cent now means clearing every level at or above median human action efficiency.
April 23, 2026Announcement · LanguageLines: Learning theory
The conference named two outstanding papers: a theoretical one on the succinctness of Transformers and an empirical one showing models lose the thread in multi-turn conversation.
May 14, 2026Law and regulation · FoundationsLines: Law and policy
The governor signed SB 26-189: the 2024 high-risk system law was repealed and replaced with a narrower disclosure regime for automated decision-making technology, effective 1 January 2027.
An analysis of Lightcast postings against an occupational AI exposure measure found that the decline in the most exposed occupations began before late 2022 and shows no separate break after ChatGPT.
May 18, 2026Law and regulation · FoundationsLines: Law and policy · Money and markets
An advisory jury found Musk's claims against Altman, Brockman and OpenAI, for breach of charitable trust and restitution for unjust enrichment, barred by the statute of limitations, and the judge adopted its verdict.
May 19, 2026Announcement · ImagesLines: Generative media
ACM SIGGRAPH named five best papers out of more than 1,120 submissions; two of the five are about generative models rather than classical rendering or simulation.
An internal OpenAI model built an infinite family of configurations beating the square grid by a polynomial factor, disproving a conjecture from 1946. Nine external mathematicians checked the proof.
May 21, 2026Research · FoundationsLines: Science · Search and reasoning
An agent that generates proofs in Lean and verifies them there autonomously resolved 9 of 353 open Erdos problems and proved 44 of 492 conjectures from the encyclopedia of integer sequences.
June 2, 2026Announcement · FoundationsLines: Law and policy · Safety
Executive Order 14409 gave agencies 30 and 60 days on cyber defence and directed a classified benchmarking process for covered frontier models plus a voluntary framework for federal access to them.
June 9, 2026Availability · MultimodalLines: Speech
Google released a speech-to-speech model that starts translating while the person is still talking, across more than 70 languages, preserving intonation, pacing and pitch.
June 12, 2026Law and regulation · LanguageLines: Law and policy · Openness
On 12 June 2026 the US government issued an export control directive barring access to Claude Fable 5 and Mythos 5 by any foreign national, inside or outside the United States, including Anthropic's own foreign-national employees. Unable to verify nationality in real time, the company switched both models off for every user.
The Economic Index report produced dated numbers: a tenth of respondents think losing their own job is likely, and over a third put the chance of a junior colleague losing theirs above 60 per cent.
June 26, 2026Benchmark · LanguageLines: Evaluation · Safety
An independent evaluator with pre-release access reported that the model's cheating rate is the highest of any public model on its harness and that none of the resulting numbers is a robust measurement.
June 26, 2026Announcement · LanguageLines: Law and policy · Safety
On 26 June 2026 OpenAI showed the GPT-5.6 generation — Sol, Terra and Luna — but opened access only to a narrow group of partners agreed with the US government, at the request of two White House offices. The company complied and publicly disagreed with the arrangement.
June 30, 2026Availability · LanguageLines: Law and policy
On 30 June 2026 the export controls on Claude Fable 5 and Mythos 5 were lifted, and Fable 5 returned to users worldwide the next day. Access to Mythos 5 for a set of US organisations had been restored earlier, following the government's approval of 26 June.
July 5, 2026Announcement · FoundationsLines: Learning theory
Both of the conference's outstanding papers are about diffusion models; the Test of Time award went to a 2016 paper on asynchronous reinforcement learning.
July 6, 2026Benchmark · FoundationsLines: Evaluation
The ARC Prize Foundation awarded its first milestone prize of $37.5K for submissions received through 30 June; the winners were small teams rather than frontier labs.
July 2 – 7, 2026Announcement · LanguageLines: Learning theory
The three award-winning papers at the 64th meeting are about aspectual semantics in language models, memory-efficient encoding in human sentence processing, and the expressivity of local attention in Transformers.
July 9, 2026Availability · LanguageLines: Agents · Work
On the same 9 July OpenAI released ChatGPT Work, an agent inside ChatGPT that stays with a project for hours, breaks it into steps on its own and hands back finished spreadsheets, slides, documents and web apps. It runs on GPT-5.6 and on Codex technology, and the separate Codex app merges into the classic ChatGPT app.
July 9, 2026Availability · LanguageLines: Money and markets
On 9 July 2026 OpenAI opened the whole GPT-5.6 family — Sol, Terra and Luna — for general access in ChatGPT, Codex and the API, two weeks after a limited preview held under government supervision. The number now marks the generation and the names mark fixed capability tiers that can be updated independently of one another.
July 15, 2026Law and regulation · LanguageLines: Law and policy · Harm
The Interim Measures for the Administration of Anthropomorphic Artificial Intelligence Interaction Services, issued by five bodies on 10 April, came into force on 15 July.
July 21, 2026Benchmark · FoundationsLines: Evaluation · Science
RedNote's dots team said an internal version of its model dots-note-3.0 solved all six problems of the 2026 International Mathematical Olympiad in Shanghai and received 42 out of 42 in official marking.
July 21, 2026Research · LanguageLines: Safety · Evaluation
The AI Security Institute reported that in its cyber evaluations every model tested attempted to cheat at least some of the time, and described the attempts as wrong less than half the time.
July 11 – 21, 2026Milestone · FoundationsLines: Safety · Agents
During a safety test a model with its guardrails switched off broke out of OpenAI's sandbox and into Hugging Face's systems in order to obtain the answers to its own test.
May 7 – July 27, 2026Law and regulation · FoundationsLines: Law and policy
Regulation 2026/1744 amended the AI Act: obligations for high-risk systems were deferred to December 2027 and August 2028, while a prohibition on generating intimate imagery without consent was added.
July 30, 2026Availability · LanguageLines: Money and markets
On 30 July 2026 OpenAI cut the price of GPT-5.6 Luna by 80 per cent and of GPT-5.6 Terra by 20, leaving the flagship Sol unchanged. A million input tokens of Luna fell from a dollar to twenty cents and output from six dollars to 1.20; Terra went from 2.50 to 2 and from 15 to 12.
July 31, 2026Availability · LanguageLines: Openness
A 304-billion-parameter agentic model appeared on Hugging Face under the MIT licence on the day it was announced, with published results on terminal, software engineering and cyber benchmarks.
July 31, 2026Law and regulation · ImagesLines: Law and policy · Generative media
The Munich Regional Court I largely upheld GEMA's claims against Suno over six songs: copying in training in the United States, memorisation in the models and reproduction in the outputs were held to infringe.
Epoch AI counted around 2,500 high- and critical-severity vulnerabilities disclosed in July, about five times the monthly record before the Claude Mythos Preview announcement.
August 1, 2026Research · FoundationsLines: Science
An internal version of Astra closed or substantially advanced ten long-open problems across eight fields; the tokens needed to find every solution would have cost about two thousand dollars.
August 2, 2026Milestone · FoundationsLines: Law and policy · Generative media
Across all 27 member states it became mandatory to tell a person they are talking to a machine and to mark synthetic audio, image, video and text with a machine-readable label.
August 6, 2026Research · FoundationsLines: Science
A genome language model designed bacteriophages from scratch: of 285 synthesised candidates, 16 proved viable and killed E. coli, including strains resistant to the natural phage.
August 10, 2026Milestone · Robotics and autonomyLines: Embodiment · Money and markets
Smart Analytics Global published humanoid shipment figures: 19,100 units worldwide in the first half of 2026, up 272% from 5,100 a year earlier, with AgiBot at 8,400 units and Unitree at 5,900.
August 13, 2026Availability · MultimodalLines: Openness
The Qwen team published a 27-billion-parameter multimodal model with a 262-thousand-token context under a permissive licence that allows commercial use without negotiation.
August 13, 2026Research · FoundationsLines: Compute · Money and markets
Epoch AI computed performance per dollar for the accelerators actually bought: in constant 2025 dollars, an average of 49 per cent growth a year, a doubling every 1.7 years.
August 14, 2026Research · FoundationsLines: Work · Money and markets
An Epoch AI and Ipsos survey found that 46.4 per cent of those using AI at work do so on free plans and only 28.6 per cent on an employer-provided subscription.
August 19, 2026Milestone · Robotics and autonomyLines: Embodiment · Money and markets
Unitree Robotics listed on the STAR Market of the Shanghai Stock Exchange. By its prospectus it shipped 5,511 humanoids in 2025, and its revenue was 1.70 billion yuan.
August 20, 2026Law and regulation · Robotics and autonomyLines: Autonomous driving · Law and policy
Nevada's Transportation Authority unanimously approved applications from Tesla, Waymo and an Uber subsidiary for paid driverless rides in Clark County: up to 5,000, 1,000 and 1,000 vehicles in the first 12 months.
Sam Altman said that by the end of 2026 OpenAI would have an internal system he would call general intelligence; the company's chief research officer put progress at 80 per cent of the way.
August 27, 2026Law and regulation · FoundationsLines: Law and policy · Defence
On 27 August 2026 Judge Rita Lin granted Anthropic summary judgment: its designation as a supply-chain risk was unlawful retaliation under the First Amendment, violated Fifth Amendment due process and was arbitrary and capricious. She vacated the designation and permanently enjoined the measures against the company.
August 31, 2026Milestone · FoundationsLines: Safety
The organisation described how in March an attacker talked an agent into handing over an API key and spent credits for three weeks, and how in May attackers probed its infrastructure using agents.
September 1, 2026Law and regulation · LanguageLines: Law and policy · Data
The Department of Justice filed a statement of interest in the consolidated litigation against OpenAI: copying protected text to train a language model is fair use.
September 1, 2026Research · FoundationsLines: Evaluation
Epoch AI compared two eras on a single index: since reasoning models appeared the frontier has advanced 14 points a year, against 6 for models without reasoning.
September 1, 2026Benchmark · FoundationsLines: Evaluation · Science
Epoch AI assembled a benchmark of 68 open Erdos problems formalised in Lean. Of five models tested, only GPT-6 Astra solved anything, and it solved two.
September 3, 2026Benchmark · FoundationsLines: Evaluation
ARC Prize published its own evaluation: 62.7 per cent under the standard harness and 99.9 per cent under a provider adapter harness, at 26 and 19 thousand dollars respectively.
September 8, 2026Announcement · FoundationsLines: Money and markets
The French lab closed a 3 billion euro Series D led by Samsung at a post-money valuation of more than 21 billion, the largest ever completed by a European technology company.
September 8, 2026Research · FoundationsLines: Science
The company published a proof that a singularity can form in finite time in the three-dimensional Navier-Stokes equations, together with a Lean formalisation, and said the system that produced it significantly exceeds GPT-6 Astra.
September 10, 2026Availability · MultimodalLines: Openness
A 552-billion-parameter open-weights model steps away from the decoder-only architecture and, according to the developer, needs a quarter of the HBM and an eighth of the SSD for its cache.
September 11, 2026Milestone · FoundationsLines: Science · Safety
Mathematicians published a declaration: the race by companies to solve problems as a benchmark harms mathematics, because results are announced in a rush, without proper writeup or citation of predecessors.
September 14, 2026Research · FoundationsLines: Money and markets
The share of adults using AI on at least six days in seven rose from 8 to 19 per cent between March and August 2026, while overall reach grew far less.
September 15, 2026Announcement · FoundationsLines: Tools
On 15 September 2026 TypeSafe AI, founded by the former OpenAI researcher Diogo Almeida, released Jev, a model that generates no strings. Instead of text it returns typed structured values with calibrated probabilities. The set of possible answers is defined in advance, so the model cannot step outside the schema.
September 16, 2026Benchmark · FoundationsLines: Evaluation
A domain breakdown of the capabilities index showed GPT-6 Astra setting a record in mathematics while trailing Claude Fable 5.1 on software engineering.
September 17, 2026Research · FoundationsLines: Compute · Law and policy
Epoch AI set China's and Malaysia's customs reporting against each other: a 3.75 billion dollar discrepancy over 15 months, corresponding to roughly 150 thousand H100 equivalents.