Physical AI reliability data engine

Evidence across the deployment lifecycle

RoboGate does not treat one simulator score as a safety certificate. It connects pre-deployment failure maps, decision evidence, policy-aware production drift, and revisioned failure data into one evidence flow.

Benchmark status · methodology review

The public Fair board currently has no qualified model scores. Twenty legacy runs remain auditable quarantine records and are not used for ranking or capability claims. Publication resumes only after Harness v2 controls and a GPU canary pass.

Product surfaces

Evidence flow

Declared contract

Qualified harness

Model run

Validated evidence

Deploy + monitor

Publication gates

1

Zero-action negative control: 0/68

2

Perfect oracle: 68/68 in three complete repetitions

3

Independent nonblank primary and wrist cameras

4

Real inference plus immutable adapter and checkpoint hashes

5

Exactly 68 complete episodes with scenario-application evidence

6

Qualification-bound validation before atomic promotion

Repository integrations

These labels describe code present in this repository. They do not imply endorsement, partnership, production deployment, or benchmark validity.

NVIDIA Isaac Sim

Harness runtime

The qualified benchmark runtime is built as an immutable derived image and applies the same physical and rendering contract to controls and model adapters.

Isaac Lab-Arena

Repository adapter

A contribution package lives under contrib/isaaclab-arena. It is an integration surface, not evidence that Fair results have passed qualification.

Azure ML + MLflow

Reference integration

The contrib/azure-physical-ai package demonstrates job submission and metric logging without changing the benchmark publication gate.

Cosmos Evaluator

Reference integration

The contrib/cosmos-evaluator-plugin package exposes checker services. Its output remains subject to RoboGate result validation and provenance rules.