New GPT-6 Astra Policy Annihilates Robot Manipulation Benchmark

Just as the dust seemed to be settling on the robotics benchmark wars, a new paper from researchers Yu-Mool Shu and Lipxin Zheng has arrived to thoroughly upend the status quo. Titled “GPT 6 Astra as an Embodied Policy,” the research introduces a model that doesn’t just nudge the needle; it practically snaps it off. The system has absolutely decimated the competition on the RoboDojo manipulation benchmark—a rigorous gauntlet of simulated and real-world tasks designed to push generalist robots to their limits.

The results are, to put it bluntly, a bit of a drubbing. The new policy—a sophisticated hybrid pairing a controller dubbed π₀.₅ with the powerhouse GPT-6 Astra—clocked a mean score of 62.6 across ten tasks. To give you some perspective, the nearest rival, Galaxea G0.5, could only muster a 38.26. We aren’t looking at an incremental gain here; this is a 64% performance surge that leaves established models from the likes of Xiaomi and Meituan looking like they’re stuck in second gear. The work seems to build on the burgeoning field of co-training vision-language-action models on massive, diverse datasets to finally crack the code of real-world generalisation.

What’s particularly telling is how “GPT 6 Astra Direct” control fared on its own: it managed a far more pedestrian 37.81. This highlights the “secret sauce” of the π₀.₅ hybrid controller, which acts as the crucial bridge—translating the high-level “intelligence” of the large model into the fluid, nuanced physical movements required for complex tasks. With the project being open-source (there’s a GitHub link tucked away on their site), you can bet the rest of the industry is currently pulling an all-nighter to reverse-engineer this architecture. You can pore over the project page and data for yourself here: GPT 6 Astra as an Embodied Policy.

Why does this matter?

This is about more than just bragging rights on a leaderboard. It signals a definitive pivot away from monolithic, “one-size-fits-all” policies in favour of more nuanced hybrid systems. This massive performance chasm suggests that marrying a “big brain” like GPT-6 Astra with a specialised, fine-grained motor control policy like π₀.₅ is exponentially more effective than asking a large model to do all the heavy lifting itself. For an industry that has long struggled to get robots out of the lab and into messy, unpredictable environments, this could be the architectural “lightbulb moment” needed to finally make autonomous helpers a household reality.