# RoboGate Observatory — Assessment Protocol

**Version:** 2.0 draft
**Status:** no learned-policy rating or PARI is currently publication-eligible
**Effective date:** none — requires a qualified cohort

## 1. Primary result

The primary task result is the number of successful state-machine completions
out of 68 qualified scenarios, with a Wilson 95% binomial confidence interval.
A simulator result applies only to its declared model, adapter, harness and
operating contract.

Errors are not scored as failed episodes. A run containing a client, server,
inference, protocol, timeout, completeness, provenance or integrity error is
ineligible for a score.

## 2. Domain results

Qualified episode outcomes may be summarized across the four declared domains:
nominal, edge, adversarial and domain randomization. Domain values must
recompute from episode records; no domain score may be inferred from an
unqualified total.

## 3. Diagnostic progress

Reach and manipulation-stage diagnostics may be published beside the primary
result when derived from verified state transitions. Progress is not a success
rate, safety grade or substitute for task completion.

## 4. Ratings and PARI

Star-rating and PARI formulas are suspended. They may be reinstated only through
a versioned protocol after:

- a qualified, sufficiently sized and clearly defined learned-policy cohort
  exists;
- cohort inclusion, correlation and missing-data policies are declared;
- uncertainty and backstops are validated; and
- the formula cannot imply deployment certification beyond the test contract.

The historical v1.0 ratings, scripted comparison and PARI are withdrawn.

## 5. Corrections and revisions

A corrected result supersedes the prior revision in rankings while both remain
auditable. The public record must identify the reason, timestamp, validator
version and hashes. A retraction appears at the same visibility as the original
claim.

## 6. Current status

Qualified learned-policy results: **0**. Historical D-041 model runs:
**20 quarantined**. No score, rank, medal, star, PARI or cross-simulator
capability conclusion is currently valid.
