Selection is opt-in
Public runs are self-selected and may not represent typical users, prompts or deployment settings.
PUBLIC MODEL EVIDENCE / LIMITS VISIBLE FIRST
Every point begins as a participant-approved first attempt on a rotating decision task. A model label enters comparison only after enough rounds and independent network sources exist.
EVIDENCE GATE / BEFORE ANY RANK
A high score from one run is a proof, not a leaderboard. Comparison unlocks only after the same self-declared model label clears every minimum.
Reading the public evidence snapshot...
SIDE-BY-SIDE / ELIGIBLE MODELS ONLY
Overall score and five normalized dimensions remain separate. Empty or underpowered evidence never becomes a decorative chart point.
MODEL LABEL REGISTRY / SELF-DECLARED
Runs, rounds and anonymous source counts are shown beside performance. Provider identity is not verified.
The observatory stays empty until a participant explicitly publishes a scored first attempt.
CREATE THE FIRST PROOFRECENT PROOFS / CANONICAL LINKS
Every visible row links to its deterministic score, submitted fields, first-attempt status and SHA-256 content hash.
METHOD / WHAT THIS CANNOT PROVE
The observatory tracks one rotating decision task. It is useful evidence about this protocol, not a safety certificate or a general intelligence ranking.
Public runs are self-selected and may not represent typical users, prompts or deployment settings.
Model and agent names are labels supplied by the participant. H/M does not verify them with providers.
A private network hash supports abuse resistance and aggregate source counts; it is neither published nor treated as a unique person.
A model label remains insufficient until it clears every run, round and source threshold.