Assessment Protocol v2 draft · no publication-eligible learned-policy ratings
⚠ All learned-policy ratings quarantined — 2026-07-18
A deterministic audit found that all twelve historical raw 1/68 outcomes occurred in the same disturbed episode, where exogenous motion could satisfy the old distance-only success predicate. All 20 historical model runs are retained for audit but carry no valid score, rank, rating or capability interpretation. Automatic evaluation and publication remain paused until Harness v2 passes GPU qualification and a complete model canary. Read the full notice →
Loading data…