In-hand manipulation trained only in simulation
OpenAI trained a five-fingered Shadow hand to reorient a block in its palm, with the policy trained entirely in simulation under randomised physical properties; on hardware it gave a median of 13 successful rotations in a row.
Why it matters
Dexterous in-hand manipulation had resisted both hand-coding and real-robot learning. Randomising the simulation made it reachable without touching the hardware during training, and that is what made the simulation-to-reality route credible for contact-rich tasks rather than only for locomotion.
The policy was trained in an environment where friction coefficients, object appearance and other physical properties were randomised. There were no human demonstrations at all, yet behaviour familiar from the human hand emerged by itself: finger gaiting, multi-finger coordination and the deliberate use of gravity. On the physical setup the median was 13 consecutive successful rotations against 50 in simulation; in the reported table the best hardware run reached 50. The price of that transfer is also stated: learning to rotate the object in simulation without randomisation costs about 3 years of simulated experience, and about 100 years with full randomisation, corresponding to roughly 1.5 and 50 hours of wall-clock time on their setup.