Antefacts

Limits

We would rather say this ourselves than have a reviewer say it for us.

Every assurance method has a boundary. Most of this market leaves theirs to be discovered. Ours is written down, published with the specification, and included in every evidence package we produce, because a boundary you state is a credential and a boundary someone finds is a scandal.

Proved

What is proved

That the published score is the output of the published scoring function applied to the recorded inputs. That the criteria applied are the criteria sealed before submissions opened. That the record has not been altered, and that the chain from any entry to a published root is unbroken.

That is a claim about the conduct of an evaluation. It is machine-checked and it is re-checkable by anyone who downloads the package.

Assumed

What is assumed

Four things, named in every package rather than buried.

Collision resistance
That a cryptographic hash is collision resistant, so that equality of recorded hashes means identity of the underlying content. This is a computational assumption, stated outside the machine-checked development, because asserting it inside would be mathematically false.
Timestamp authority
That the timestamping authority, which we do not control, is honest and uncompromised.
Execution environment
That the recorded execution environment is the one that actually executed, as attested by container digest and isolation attestation.
Proof toolchain
That the proof assistant's kernel, its toolchain and its compilation are correct.
Never proved

What is measured but never proved

Everything about the system under test. Performance is reported as a statistical interval with a pre-registered sample size and method, or as a score against pre-registered criteria. It is never described as proved, verified, certified or guaranteed, and no package we issue will say otherwise.

Out of scope

What the method does not address at all

It does not make the numbers right. It makes them checkable.

It does not make the criteria sensible. A poorly chosen set of criteria, honestly executed, passes every check we make. The defence against that is not mathematics: it is that the criteria are authored by the party whose decision it is, published in full before submissions open, and open to challenge.

It does not detect a leaked held-back set or a colluding assessor. Isolation, two-person access, hash-chained logs, disjoint slices and canary probes make either attributable after the fact. They do not make it impossible. The genuine defence is that every participant may rerun its own submission independently.

It does not close the window between published roots. A rewrite of entries made since the last independently published root is possible in principle. The interval between roots is stated publicly, because it is the size of that window.

On the proof

And one honest note about the proof itself

The part of our machine-checked layer that verifies a result is a decision procedure, and its adequacy result is close to definitional. We say so rather than dressing it up. The substance sits in two other places: a scoring function that is total, pure and exactly rational-valued, compiled from the same definition that was checked, which is what makes rerunning it produce identical numbers; and a chain check whose executable form is proved to agree with an independent characterisation of validity, so a bug in the checker cannot define validity into existence.

Antefacts is not an accredited certification body and is not a notified body. We do not issue certificates and we hold no accreditation.