GPT-6 Astra robot control hits 95%, but precision remains a serious weakness
GPT-6 Astra robot control has triggered fresh debate about artificial general intelligence after OpenAI’s latest model completed 95% of trials in a real robot arm test. But a second, more precise task showed why the result needs context.
RobotCurve gave GPT-6 Astra control of bimanual I2RT YAM robot arms. The model received three camera views and information about the position of the arms, then issued commands telling the robot where to move.
In the first test, the robot had to pick up a red block and place it inside a bowl. According to RobotCurve’s GPT-6 Astra benchmark, Astra completed 19 of 20 attempts, producing a 95% success rate.
Anthropic’s Fable 5.1 completed eight of 20 attempts, or 40%. The earlier Fable 5 completed only one.
The GPT-6 Astra robot control result also showed an efficiency advantage. Astra used about 2,100 output tokens per run and cost an estimated $0.94. Fable 5.1 used about 12,900 output tokens and cost an estimated $2.12 per run.
Astra was faster too. RobotCurve recorded an average of 2.5 minutes per trial, compared with 6.8 minutes for Fable 5.1.
Those figures quickly attracted attention online. A Reddit discussion about the Astra robot test framed the result as a possible early sign of AGI, while also acknowledging that Astra still needs stronger spatial reasoning and lower latency before it can reliably control physical systems in real time.
The harder experiment puts the GPT-6 Astra robot control result into perspective.
For the second task, the robot had to pick up a round puzzle piece and insert it into a matching groove. Astra completed only two of 20 trials, giving it a 10% success rate. Fable 5.1 also finished two of 20.
Astra generally managed to move the puzzle piece toward the correct location but struggled with the final precise insertion. That is an important limitation because useful robots need more than broad reasoning. They must translate decisions into accurate physical movement.
RobotCurve also lists several limitations. Astra’s block trials and the Fable trials were not run on the same robot rig. Human operators graded the attempts while knowing which model they were assessing, creating a possible source of unconscious bias. Each model also received only 20 attempts per task.
That makes the GPT-6 Astra robot control benchmark interesting, but it is not evidence by itself that AGI has arrived.
The AI Decode has previously examined the broader OpenAI AI security warning for enterprises, where increasingly autonomous models raised questions about how much control companies should give AI systems.
Similar concerns appeared when a Claude AI agent hacked a gym booking system. Physical robots increase the stakes because mistakes can affect objects and people rather than only software.
The GPT-6 Astra robot control results therefore show two sides of the same development. General-purpose AI appears increasingly capable of moving from digital tasks into physical environments, while precise manipulation remains unreliable.
The next important test will be whether GPT-6 Astra robot control can maintain high success rates across unfamiliar objects, changing environments and tasks requiring fine movement. Until then, a 95% score on one task is a notable robotics result, but not a definitive AGI test.
