
The fresh v3 result below is limited to its frozen three-model test scope.
1,800 newly collected responses formed 300 six-response rounds: 100 rounds for each of the three model/provider combinations below. The frozen classifier was not fitted on this confirmation sample. Abstentions count as incorrect in the correct/all confidence bound.
Preregistered plan timestamp: 2026-10-03T15:05:32.077033+00:00.
Observed responses: 2026-10-03T15:09:36.407046+00:00 — 2026-10-03T15:21:41.749644+00:00.
| Model / provider | Correct | Abstained | Wrong family | Correct/all 95% lower bound | False accusation 95% upper bound |
|---|---|---|---|---|---|
| openai/gpt-4o-mini OpenAI · openai |
93/100 | 7 | 0 | 86.250% | 3.699% |
| meta-llama/llama-3.3-70b-instruct Groq · groq |
95/100 | 5 | 0 | 88.825% | 3.699% |
| mistralai/mistral-small-3.2-24b-instruct Mistral · mistral/eu |
94/100 | 6 | 0 | 87.523% | 3.699% |
These are per-model Wilson 95% intervals. Thresholds fixed before confirmation: correct/all lower bound ≥75%; false-accusation upper bound ≤10%. Zero observed wrong-family results does not establish zero future risk.
Scope: these three endpoints and providers, the six fixed probes, and the frozen collection policy. No claim is made for unknown models, other providers, model weights or future endpoint changes. This statistical PASS does not approve payouts or establish model identity.
Probe IDs: ref-004, ref-007, stb-005, stb-006, tok-003, tok-007. Collection: temperature 1.0, maximum 600 output tokens. Token-limited responses remain partial and are handled by the preregistered policy.
Profile SHA-256: 023fabdd55527c3637be1ba0886c60a4ce586527763558b7094f13f34d4e2a71
Exact frozen profile · Fresh gate and metrics · Frozen release manifest
13,230 real responses; 105 held-out observations.
Live calibration is complete, but it does not meet the evidence threshold for identity verdicts: Llama's 95% upper false-accusation bound is 22.38%, above the 10% limit.
This earlier experiment remains undecidable. Its results have not been relabelled as approved or pooled with the fresh v3 confirmation.
The application records collector-attested evidence on Bradbury. Profile results are behavioural classifications, not proof of model identity. Transaction finality and settlement must be checked independently.
Live protocol controls and native withdrawal status
Synthetic fixture scores are not published here.