Back to timeline

Research · August 1, 2018

In-hand manipulation trained only in simulation

OpenAI trained a five-fingered Shadow hand to reorient a block in its palm, with the policy trained entirely in simulation under randomised physical properties; on hardware it gave a median of 13 successful rotations in a row.

Why it matters

Dexterous in-hand manipulation had resisted both hand-coding and real-robot learning. Randomising the simulation made it reachable without touching the hardware during training, and that is what made the simulation-to-reality route credible for contact-rich tasks rather than only for locomotion.

The policy was trained in an environment where friction coefficients, object appearance and other physical properties were randomised. There were no human demonstrations at all, yet behaviour familiar from the human hand emerged by itself: finger gaiting, multi-finger coordination and the deliberate use of gravity. On the physical setup the median was 13 consecutive successful rotations against 50 in simulation; in the reported table the best hardware run reached 50. The price of that transfer is also stated: learning to rotate the object in simulation without randomisation costs about 3 years of simulated experience, and about 100 years with full randomisation, corresponding to roughly 1.5 and 50 hours of wall-clock time on their setup.

Event record

Event date
August 1, 2018
Timeline date
Event date
Verification
Sources gathered automatically · September 20, 2026
Lines
ID
evt-0431

The date the first version was posted to arXiv. OpenAI’s own post is dated slightly earlier, but openai.com refuses automated fetching, so the preprint posting is taken as the date.

Sources

Related events

Records that link to this one