Search inside the chain of reasoning
On 9 January 2025 Renmin University of China and Tsinghua showed Search-o1: a reasoning model that stops at the point of its own uncertainty, searches an external source and returns to reasoning. First the authors measured how often it hesitates at all.
Why it matters
Until then search was placed before reasoning: retrieve documents, then think. Here search moved inside — the model reaches for knowledge at the moment it is missing rather than in advance, and that changes where error in a long chain comes from.
First the measurement everything grew out of. On the GPQA diamond set the authors counted how often QwQ-32B-Preview uses words of uncertainty inside its own reasoning: the word "perhaps" alone appears on average over 30 times per chain. That figure is about one word and one model, not uncertainty in general. The second component is the Reason-in-Documents module. Retrieved documents are verbose, and dropping them into the chain breaks it, so a separate module first compresses the document against the current query and only then hands it to the reasoning. Tested on hard problems of three kinds — science, mathematics and coding — and on six open-domain question-answering benchmarks. What the record does not claim. Three kinds of task are not five domains; the paper names science, mathematics and code. The paper gives no single headline gain: it compares against direct reasoning and against ordinary RAG on each set separately.