Voirdire

Measurement status

The fresh v3 result below is limited to its frozen three-model test scope.

Fresh v3 confirmation: PASS within the tested scope

1,800 newly collected responses formed 300 six-response rounds: 100 rounds for each of the three model/provider combinations below. The frozen classifier was not fitted on this confirmation sample. Abstentions count as incorrect in the correct/all confidence bound.

Preregistered plan timestamp: 2026-10-03T15:05:32.077033+00:00.

Observed responses: 2026-10-03T15:09:36.407046+00:00 — 2026-10-03T15:21:41.749644+00:00.

Model / providerCorrectAbstainedWrong familyCorrect/all 95% lower boundFalse accusation 95% upper bound
openai/gpt-4o-mini
OpenAI · openai
93/10070 86.250%3.699%
meta-llama/llama-3.3-70b-instruct
Groq · groq
95/10050 88.825%3.699%
mistralai/mistral-small-3.2-24b-instruct
Mistral · mistral/eu
94/10060 87.523%3.699%

These are per-model Wilson 95% intervals. Thresholds fixed before confirmation: correct/all lower bound ≥75%; false-accusation upper bound ≤10%. Zero observed wrong-family results does not establish zero future risk.

Scope: these three endpoints and providers, the six fixed probes, and the frozen collection policy. No claim is made for unknown models, other providers, model weights or future endpoint changes. This statistical PASS does not approve payouts or establish model identity.

Probe IDs: ref-004, ref-007, stb-005, stb-006, tok-003, tok-007. Collection: temperature 1.0, maximum 600 output tokens. Token-limited responses remain partial and are handled by the preregistered policy.

Profile SHA-256: 023fabdd55527c3637be1ba0886c60a4ce586527763558b7094f13f34d4e2a71

Exact frozen profile · Fresh gate and metrics · Frozen release manifest

Original live run: UNDECIDABLE

13,230 real responses; 105 held-out observations.

Live calibration is complete, but it does not meet the evidence threshold for identity verdicts: Llama's 95% upper false-accusation bound is 22.38%, above the 10% limit.

This earlier experiment remains undecidable. Its results have not been relabelled as approved or pooled with the fresh v3 confirmation.

Original diagnostic data and integrity hashes

The application records collector-attested evidence on Bradbury. Profile results are behavioural classifications, not proof of model identity. Transaction finality and settlement must be checked independently.

Live protocol controls and native withdrawal status

Synthetic fixture scores are not published here.

Source, methodology and limitations