CORRECTION
Correction: Evidence Before Validation-Layer Claims
Originally published April 14, 2026 · corrected July 18, 2026
What we are correcting
The original post used RoboGate's historical learned-policy results and a scripted-controller comparison to argue for a large cross-simulator capability gap. We withdraw that numerical comparison and its interpretation.
What the audit found
All twelve historical raw successes occurred only in moving_disturbance episode 42. Under its deterministic seed, exogenous motion can move an ungrasped object into the target, while the old scorer checks only distance. The scripted and learned-policy paths were also non-equivalent, and adapters, budgets and provenance were not sufficiently controlled. Raw zeros therefore remain unqualified rather than becoming valid zero scores.
What remains valid
Physical AI teams still need deployment-specific reliability evidence, but that conclusion does not depend on the withdrawn model comparison. Useful evidence must declare the policy, embodiment, cell and operating envelope, pass negative and positive controls, attest live inference, preserve immutable provenance and validate the full manipulation state transition.
What changes now
All 20 historical model runs are quarantined. Automatic evaluation and publication are paused. Harness v2 requires grasp, lift, transport, release and stable placement; per-scenario controls; one versioned manifest and budget; strict adapter preflight; and fail-closed artifact promotion. We will publish replacement numbers only after target-GPU qualification and a complete model canary.