o1 and compute at answering time
OpenAI showed a model trained to think before answering: it generates a long hidden chain of reasoning, and the longer it thinks the more accurate the answer.
Why it matters
A second axis of growth appeared: quality began to depend on time spent thinking, not only on model size and training volume.
The model was trained by reinforcement on problems with a checkable answer: mathematics, code, the sciences. Reward went to the correct result, and the model developed habits that look human, checking itself, backtracking, trying another approach. On the qualifying exam for the American mathematics olympiad the score jumped from a few per cent to over eighty. OpenAI hid the chain of reasoning and showed only a summary, which drew criticism. The practical consequence: the cost of an answer stopped being fixed, because accuracy is now paid for in seconds.