Reasoning by depth rather than by tokens
On 7 February 2025 a group from the ELLIS Institute in Tübingen and the University of Maryland showed an architecture that thinks longer not in words but in repetitions: a recurrent block unrolls to arbitrary depth at answering time. A proof-of-concept model of 3.5 billion parameters was trained on 800 billion tokens.
Why it matters
Until then thinking longer meant more tokens in the context window — a chain of thought that takes up room and needs training data containing such chains. Here depth became a separate dial: the model can think longer without writing one extra word.
Three properties the authors claim over chain-of-thought: no specially collected reasoning data is needed, a small context window suffices, and in this form the model can capture reasoning that does not sit well in words. The limit is measured and stated carefully. The paper says performance on reasoning benchmarks improves, sometimes dramatically, up to a computation load equivalent to 50 billion parameters. That is about the amount of computation, not about a 3.5-billion model matching a 50-billion one in quality — the difference matters, and the source does not make the second claim. What the record does not claim. This is a research model, not a release; the authors call it a proof of concept. The gain is described as "sometimes dramatic", with no single figure across benchmarks.