RTFM: a world without a 3D model
On 16 October 2025 World Labs showed a research preview of RTFM: a model that generates in real time the video of a scene a viewer moves through, on a single H100 GPU.
Why it matters
The scene here exists not as a mesh or Gaussian splats but as a neural network’s memory of posed frames: the renderer became learned. An editorial assessment: this is the developer’s claim about its own system, not a measured result.
By the company’s post, RTFM is an autoregressive diffusion transformer over sequences of frames, trained end to end on large-scale video to predict the next frame. Input frames become activations (the KV cache) that implicitly hold the world; there is no explicit 3D representation. Each frame has a pose in space, so the memory has spatial structure, and “context juggling” picks the nearest frames for a new one. The company claims interactive framerates on a single H100 and persistence of the world. It gives its own arithmetic: a naive interactive 4K stream at 60 frames a second needs over 100,000 tokens a second, and an hour of persistence needs contexts of well over 100 million tokens. What the record does not state. The post gives no parameter count, no comparison and no measure of persistence, and only promises dynamic worlds and interaction with them as a next step. A trade outlet the same day restates the post and reports no test. The date is the day of the post; one retelling in a search result gave 17 October, but it was not opened.