Retrieval-augmented generation (RAG)
On 22 May 2020 Patrick Lewis and eleven co-authors from Facebook AI Research, University College London and New York University posted RAG: a pre-trained generator (BART-large, 400M parameters) combined with a dense vector index of Wikipedia, the December 2018 dump split into 21,015,324 passages of 100 words, with retriever and generator fine-tuned together. RAG-Sequence scored 44.5 exact match on Natural Questions, 56.1 on TriviaQA, 45.2 on WebQuestions and 52.2 on CuratedTrec; the paper claims a new state of the art on all four, for TriviaQA only on the split comparable with T5, since DPR has 57.9 on the usual one.
Why it matters
A generator could draw on documents retrieved from an external index rather than only on what its parameters store, and write free text rather than cut an answer out of a passage. The abstract names provenance and updating world knowledge as the open problems this addresses. Search-o1 in 2025 cites this paper as the source of retrieval-augmented generation.
RAG-Sequence uses one retrieved document for the whole output, RAG-Token may use a different document for each token; both treat the document as a latent variable. Table 1 of the first version: T5-11B without retrieval 34.5 on Natural Questions, REALM 40.4, DPR 41.5, RAG-Token 44.1, RAG-Sequence 44.5. For TriviaQA the first version gives 56.1 on the usual test split and 68.0 on the hidden Wiki split. The figure 56.8 for TriviaQA that appears in retellings belongs to the fourth version (April 2021) and is not this record's number.