NVIDIA Isaac Sim
Harness runtimeThe qualified benchmark runtime is built as an immutable derived image and applies the same physical and rendering contract to controls and model adapters.
Physical AI reliability data engine
RoboGate does not treat one simulator score as a safety certificate. It connects pre-deployment failure maps, decision evidence, policy-aware production drift, and revisioned failure data into one evidence flow.
Benchmark status · methodology review
The public Fair board currently has no qualified model scores. Twenty legacy runs remain auditable quarantine records and are not used for ranking or capability claims. Publication resumes only after Harness v2 controls and a GPU canary pass.
One immutable scenario manifest, explicit manipulation transitions, adapter evidence, and control runs before a score can be ranked.
Maps a specific policy, robot, cell, and operating envelope to observed failure conditions instead of reducing readiness to one universal score.
Ingests CSV, MCAP, or authenticated webhook telemetry and detects drift by recipe and policy-version cohort against an explicit baseline.
Retains source, test contract, evidence, and revision history so simulation and field failures can become a compounding reliability dataset.
Declared contract
Qualified harness
Model run
Validated evidence
Deploy + monitor
Zero-action negative control: 0/68
Perfect oracle: 68/68 in three complete repetitions
Independent nonblank primary and wrist cameras
Real inference plus immutable adapter and checkpoint hashes
Exactly 68 complete episodes with scenario-application evidence
Qualification-bound validation before atomic promotion
These labels describe code present in this repository. They do not imply endorsement, partnership, production deployment, or benchmark validity.
The qualified benchmark runtime is built as an immutable derived image and applies the same physical and rendering contract to controls and model adapters.
A contribution package lives under contrib/isaaclab-arena. It is an integration surface, not evidence that Fair results have passed qualification.
The contrib/azure-physical-ai package demonstrates job submission and metric logging without changing the benchmark publication gate.
The contrib/cosmos-evaluator-plugin package exposes checker services. Its output remains subject to RoboGate result validation and provenance rules.