Criteria exist before the data does
Criteria, weights, pass thresholds and the analysis plan are written before any subject is assessed, published, and hash-committed. The commitment is what makes it meaningful: the hash is public before submissions open, so nothing can be adjusted afterwards without the change being visible.
For a statistical criterion the pre-registration also fixes the sample size, the interval method, the confidence level and every numeric constant the method consumes. An analysis plan you cannot read before the data is not an analysis plan.
Subjects never see what they are scored on
Subjects receive the requirement specification. They do not receive the scoring inputs in advance, they are contractually barred from retaining them, and the inputs are held in an isolated store that subject-facing systems cannot reach, under two-person access with a hash-chained access log.
Where the thing being evaluated is a hosted API, the inputs necessarily pass through the subject's own provider at run time. We do not pretend otherwise. Each slice of held-back data carries a state that only ever moves one way, from unexposed to exposed to burned, slices are never shared between evaluations involving the same subject, and canary strings are probed afterwards to detect retention.
A harness anyone can rerun
The scoring harness is published. A third party can obtain it, rerun the evaluation and obtain the same numbers, without contacting us and without our cooperation. Every result carries a replication status that says honestly what is still possible: whether the evaluation can be re-executed indefinitely, only while the provider still serves the same version, or no longer at all because that window has closed.
That last state is not a defect. A record whose subject has moved on is still a checkable record, and saying so is more useful than implying a reproducibility that has quietly lapsed.
A record that resists rewriting
Every document in an evaluation is canonicalised, hashed, signed and timestamped by an authority we do not control, then entered in an append-only ledger where each entry commits to the one before it. The chain roots are published at a stated interval to independent locations, so any attempt to rewrite history contradicts something already public.
Correction happens by superseding entry, never by edit. A corrected result is a new record that points at the old one, and the old one remains.
Checked, not asserted
Every evaluation produces a machine-checked certificate stating that the published score is the output of the published scoring function applied to the recorded inputs, that the criteria applied match those sealed before submissions opened, and that the chain is unbroken. It is discharged in a proof assistant, and the checker is published so it can be re-executed rather than trusted.
The certificate also publishes the assumptions it rests on and the axioms it uses. What those are, and what they do not cover, is on the limits page.
And a route to contest
Any participant may obtain the harness and rerun its own submission at any time, independent of any dispute. Where there is a dispute, a challenge is a defined record with admissible grounds, a bounded response deadline and a stated outcome, and it enters the ledger whether it is upheld or rejected. A process that cannot be contested is not a process, it is an announcement.