All entries

Numbers from the wrong version

In three days the atlas grew from 657 records to 848: the methods the 2020s stand on, Ukraine, machine translation and retrieval, recommendation, safety, law, medicine, vision, sound, games, the 1980s outside the laboratory and language in the 1990s. The commonest error of these passes was my own: numbers taken from later versions of preprints.

Two briefs

The first brief, of 23 September, had twelve requests and closed gaps the axis itself had shown: the methods of 1991–2022 that the 2020s stand on, Kyiv and Ukraine, centres in Europe, Asia and Africa, machine translation and retrieval, recommendation and advertising, safety and alignment, law before 2021, medicine, vision before 1980 and the generative branch, multimodality and sound, games after AlphaZero and benchmarks for code. The first request, robotics from 1960 to 1999, was done on the 23rd; the other eleven went through in these days, and they are 136 records. The second brief, of 25 September, went deeper in two places: 1980–2000, and images and multimodality before 2010. It has the McGurk effect, Put-That-There, the first digital image, GrabCut and seam carving, QBIC, face recognition at the Super Bowl in Tampa, Shazam, the Blizzard Challenge, PROSPECTOR, Alvey and ESPRIT, latent semantic analysis, Brill’s tagger and Kneser-Ney smoothing. Another 55 records.

All but two of the requests in both briefs were done without an outside report. The session writes its hypotheses (dates, figures, links) into the package’s control log before opening the first source, and its verdicts after. Two requests, on Ukraine and on Europe, Asia and Africa, went through a Gemini report from which only the names were taken. The prompts sent there are kept beside the brief, so that the report can be read against what was actually asked. Of the Ukrainian report’s seven addresses, two led to the work with the content the report claimed for it.

For Ukraine the 6-in-1 made one exception to the ban on Russian sources. For information about Ukraine before 1991, where no other copy exists, such sources may be used, marked !!!ru. No source in that package needed it.

The commonest error was mine

I wrote the hypotheses myself, and their most persistent defect is the one this log had already caught in other people’s reports: figures from later versions of a preprint. In the methods of the 2020s, four of the five wrong figures came from later versions rather than the one the record is dated to: LoRA, RoPE, RAG. Alignment faking is 14 % in the second version and 12 % in the first. VQA’s quarter of a million images and ten answers per question exist only in later versions; the first has 123,285 images and three answers. R-CNN, YOLO, COCO: the same. The rule that follows is a single one: a figure belongs to the version dated the same as the record, and is read there, not remembered.

The other wrong hypotheses are of the same class, confidence without a document. Fenton’s 222,135 are women, not examinations. SWE-bench’s 1.96 % is Claude 2 with BM25, and the headline 4.8 % needs an oracle retriever. OpenAI Five’s 7,215 are wins, not games, and 3,140 of them are games the opponents abandoned. DARPA’s Strategic Computing plan never names Japan once, although the British government did call Alvey a response to the Japanese initiative.

The brief was wrong too. METEO ran experimentally from 9 December 1975, not from 1981. The audit the Watson for Oncology request pointed to is of another MD Anderson system, and the words “Watson for Oncology” occur in it zero times. The record follows the document.

The card of what was out of reach

Every pass now ends with one card: what it did not reach, at what address, which claim it was looking for and what changes if it is found. The 6-in-1 works through it in a separate chat, often with a Gemini search summary. Such a summary is read as a pointer, not a source: it mostly finds the right document and gets wrong what the document says.

That is how Starner’s 1995 thesis, which the 6-in-1 downloaded by hand, showed that the system did use coloured gloves after all, and moved the record to February 1995. Rowley’s face detector moved from the journal paper of January 1998 to the report of November 1995, which already had the same 130 images. LabelMe moved from 2007 to an MIT memo of September 2005. The winner of Blizzard 2005 was named by the entrants’ own papers, and they contradicted the summary on which system was which. Tampa got its first press report, the St. Petersburg Times of 31 January 2001, and not one of the summary’s figures was in it.

A price from a later month

The error was noticed while a video about 1997 was being put together: the atlas gave Dragon NaturallySpeaking a price of around two hundred dollars. The record’s own source, Technology Review in 1998, says $700 at launch, and $199 is where Dragon came down after IBM ViaVoice went on sale at $99 that August. It is the same flaw as with the preprints: a true number that belongs to a later moment than the record.

The record was rewritten from Dragon’s press release of 2 April 1997, found in a Wayback copy from May of that year. Announced in April, on sale from June at $695, and the hundred words a minute given as the company’s promise rather than a measurement. No source gives the nine thousand dollars for DragonDictate that the old text compared it with, so the comparison is gone. The 6-in-1 found a source for that sum the same evening: Reuters in the New York Times of 20 March 1990, where the $9,000 buys the software together with a speech recognition board. It now stands under DragonDictate’s own record, and the NaturallySpeaking text stays as it was read. The 6-in-1 read the rewritten record in both languages and accepted it, the third record confirmed by a person.

Byte bounds

Every record adds 200 to 390 bytes gzip to /uk/records/, depending on how much it says about what was not found. Over these days the bounds of that page, the axis, the axis’s data, people, organisations and lines were raised more than once. Each time it was the pass’s first action, before its first record, with the arithmetic beside it, and several times the raise turned out to have come one pass before it was needed. The tightest page now is /uk/records/ itself: 2,671 bytes under its bound, and the next pass starts there.

On that same 24 September the 6-in-1 accepted a record as a person for the second time: Wiener’s predictor, rewritten from his own final report.

Entry written September 26, 2026

Commits this entry accounts for

  • 0eb7ddc
  • 989fc58
  • 85f0d0f
  • e649692
  • e9b004f
  • b93e7a7
  • 7ab23ae
  • 257ef4a
  • d722056
  • 5f447b1
  • a96029f
  • 6e0a74d
  • 14218fb
  • d6d2a64
  • cc96971
  • 5679c9d
  • 86ded07
  • c5305d3
  • 2866ac1
  • 1e8e6b5
  • 2317787
  • ae7c1b4
  • 31fe4f3
  • 2ec9ff4
  • df5fae8
  • 1940d3b
  • 5fdd82b
  • 0f5e5dd
  • 65e12a4
  • 342ccaa
  • 61d0365
  • 0004f8d
  • cbcfaf9