Back to Blog

CORRECTION

Correction: Evidence Before Validation-Layer Claims

Originally published April 14, 2026 · corrected July 18, 2026

This correction replaces every RoboGate learned-policy score and capability claim in the original post. Historical raw artifacts remain available for audit.

What we are correcting

The original post used RoboGate's historical learned-policy results and a scripted-controller comparison to argue for a large cross-simulator capability gap. We withdraw that numerical comparison and its interpretation.

What the audit found

All twelve historical raw successes occurred only in moving_disturbance episode 42. Under its deterministic seed, exogenous motion can move an ungrasped object into the target, while the old scorer checks only distance. The scripted and learned-policy paths were also non-equivalent, and adapters, budgets and provenance were not sufficiently controlled. Raw zeros therefore remain unqualified rather than becoming valid zero scores.

What remains valid

Physical AI teams still need deployment-specific reliability evidence, but that conclusion does not depend on the withdrawn model comparison. Useful evidence must declare the policy, embodiment, cell and operating envelope, pass negative and positive controls, attest live inference, preserve immutable provenance and validate the full manipulation state transition.

What changes now

All 20 historical model runs are quarantined. Automatic evaluation and publication are paused. Harness v2 requires grasp, lift, transport, release and stable placement; per-scenario controls; one versioned manifest and budget; strict adapter preflight; and fail-closed artifact promotion. We will publish replacement numbers only after target-GPU qualification and a complete model canary.