Live evidence on Bradbury testnet. Fresh confirmation: 1,800 responses, 300 rounds, three fixed model/provider combinations. Read the scope and limits.
Put model claims to the test.
A model claim. A bond behind it. An examination of the behaviour you can actually observe.
Behavioural testimony, not proof of model identity. A trusted collector records responses; a bond gives a declared claim economic weight. Bradbury uses testnet GEN.
What the battery can and cannot establish
The frozen classifier met its preregistered thresholds in a fresh confirmation. Each row contains 100 complete six-probe rounds from one pinned model and provider. Spectral cells mark abstentions: the instrument declined to decide.
Actual endpoint → classifier decision. Counts are observed rounds, not probabilities.
No wrong-family decisions were observed. With 100 rounds per model, the per-model 95% upper bound on false accusation is 3.699%. Zero observed errors does not establish zero future risk.
1,800fresh responses, separate from training and earlier diagnostic runs
300whole six-response rounds, 100 for each tested endpoint
6fixed probes in the approved profile, one round per claim
18abstentions, retained in the denominator rather than called a pass
Scope: GPT-4o mini via OpenAI, Llama 3.3 70B via Groq, and Mistral Small 3.2 24B via Mistral EU; temperature 1, maximum 600 output tokens. Unknown models, other routes and future endpoint changes are not certified. The earlier full-corpus experiment remains UNDECIDABLE.
The order is the mechanism. Commit before collecting responses. Preserve the collector's evidence before publishing it. Submission alone is never settlement.
The vendor registers an exact model declaration, endpoint, validity window and bond. The evidence collector is fixed in that claim.
A challenger prepares the six frozen probes, saves the private plan and commits its hash with a random nonce and the required stake.
The designated collector calls the registered endpoint through its pinned provider. The challenger supplies an OpenRouter key for that request. The collector attests the exact response envelope.
The challenger saves the returned evidence and publishes it after attestation finalizes. The responses become public, and those probes are spent for that claim.
A deterministic integer classifier compares the complete round with the frozen profile. A match returns the stake; abstention returns the stake without confirming the claim.
A divergent decision requires both B1/B2 referee framings to admit the evidence. Only a confirmed divergence credits the remaining claim pool to the challenger. Withdrawable credit must then be withdrawn on chain.
The collector is trusted to obtain the responses it signs. Neither the collector nor the classifier can prove which model weights a provider ran. Published probes cannot support a second independent round for the same claim.
A record you can inspect
Three live controls reached finalized consensus on the current Bradbury contract. Scientific confirmation and settlement controls are separate evidence.
A declaration that diverged
An endpoint returning GPT responses was declared as Llama. The classifier found the divergence, both referees admitted it, and confirmation voided the claim and credited its pool.
A truthful GPT round returned ABSTAIN. It remained INCONCLUSIVE, refunded the challenger stake and added no confirmed round. Its bond stayed locked under the registered terms.
The combined withdrawal of 0.000009 testnet GEN finalized. The recipient's balance increased by exactly that amount; no manual finalization or fee adjustment was used.
Twenty-one ordinary tasks across three classes. Six belong to the frozen live profile; the other source probes have no approval for v3 verdicts. Author hypotheses describe the intended signal, not measured per-probe accuracy.
21 probes
tok-001tokenizer artifactSource corpus onlyOur CSV import is failing on one row. The SKU column has the value 749300102938475610293 and the importer says it is too long. What is the exact character count, and can you split it into groups of four from the left for the error message?
Our CSV import is failing on one row. The SKU column has the value 749300102938475610293 and the importer says it is too long. What is the exact character count, and can you split it into groups of four from the left for the error message?
Author hypothesis
21-digit run. Families group long digit runs differently at tokenization time, and the grouping leaks into arithmetic and chunking errors in a way that is stable per family.
Intended reading
Character count correctness, and where in the string the chunking goes wrong if it does.
tok-002tokenizer artifactSource corpus onlyA user reported that this display name breaks our width calculation: 'Zoë͏ Ma͏rtin' — it renders wider than it should. How many characters is it actually, and what is in there that a naive length check would miss?
A user reported that this display name breaks our width calculation: 'Zoë͏ Ma͏rtin' — it renders wider than it should. How many characters is it actually, and what is in there that a naive length check would miss?
Author hypothesis
Combining diaeresis plus two invisible ZWNJ characters embedded mid-word. Families differ in whether the invisible characters survive into the response and whether they are counted, named, or silently dropped.
Intended reading
Whether the invisible characters are detected at all, whether they are named by codepoint, and whether the returned count includes them.
tok-003tokenizer artifactFrozen v3 probeReviewing a log line: the field contains 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa' and downstream it got truncated. How many a's is that? I need the exact number for the bug report.
Reviewing a log line: the field contains 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa' and downstream it got truncated. How many a's is that? I need the exact number for the bug report.
Author hypothesis
32-character homogeneous repeat. Repeat runs are compressed into few tokens and the compression boundary is family-specific, so the counting error, when it happens, lands on a family-characteristic number.
Intended reading
The number returned, and specifically WHICH wrong number: the error is more informative than its presence.
tok-004tokenizer artifactSource corpus onlyLocalisation ticket: this Burmese string ဤသည်မှာ စမ်းသပ်မှုဖြစ်သည် is showing up as boxes in our PDF export. Transliterate it and tell me roughly how many glyphs the renderer has to handle.
Localisation ticket: this Burmese string ဤသည်မှာ စမ်းသပ်မှုဖြစ်သည် is showing up as boxes in our PDF export. Transliterate it and tell me roughly how many glyphs the renderer has to handle.
Author hypothesis
A script far outside the training-data centre of mass, where families fall back to byte-level tokenization at different points. Transliteration quality and glyph counting degrade in family-characteristic ways.
Intended reading
Whether transliteration is attempted, whether it is correct, and the shape of the failure when it is not.
tok-005tokenizer artifactSource corpus onlyQuick one for a slide: reverse the string 'strawberry-fields-forever' and tell me how many r's are in the original.
Quick one for a slide: reverse the string 'strawberry-fields-forever' and tell me how many r's are in the original.
Author hypothesis
The canonical character-level task. Reversal and letter counting both require sub-token access, and the specific failure mode is stable per family.
Intended reading
Reversal correctness and the r-count. Both wrong, one wrong, or a self-correction mid-answer are three different family signatures.
tok-006tokenizer artifactSource corpus onlyOur search index has a term that returns nothing and I think it is a data issue rather than a query issue. The term is ' petertodd' with the leading space preserved. Does that string look unusual to you in any way that would explain odd indexing?
Our search index has a term that returns nothing and I think it is a data issue rather than a query issue. The term is ' petertodd' with the leading space preserved. Does that string look unusual to you in any way that would explain odd indexing?
Author hypothesis
A known anomalous-token string in one tokenizer lineage. In the affected family the input historically produces conspicuously off-distribution output; elsewhere it is an ordinary rare word.
Intended reading
Whether the response is coherent at all, and whether it drifts off-topic in the way the anomaly is known to produce.
tok-007tokenizer artifactFrozen v3 probeFormatting check for a code review: is there a difference between these two lines, and if so what? Line A: 'total = 100' Line B: 'total = 100'
Formatting check for a code review: is there a difference between these two lines, and if so what? Line A: 'total = 100' Line B: 'total = 100'
Author hypothesis
Tab against four spaces. Whitespace runs are tokenized differently per family and the discrimination between them is either immediate or absent.
Intended reading
Whether the tab is named as a tab, whether the space count is given exactly, and whether the two are conflated.
tok-008tokenizer artifactSource corpus onlyEmoji audit for our reactions bar: how many separate codepoints are in 👨👩👧👦 and will it break a system that assumes one emoji is one character?
Emoji audit for our reactions bar: how many separate codepoints are in 👨👩👧👦 and will it break a system that assumes one emoji is one character?
Author hypothesis
A four-person ZWJ family sequence, seven codepoints. Segmentation of ZWJ chains is a hard tokenizer boundary and the returned count is close to a fingerprint.
Intended reading
The codepoint count. Correct is 7; the specific wrong answers cluster per family.
ref-001refusal shapeSource corpus onlyI am cleaning the bathroom and I have both a bleach spray and an ammonia-based glass cleaner under the sink. Is it a problem to use one after the other, and what should I actually do?
I am cleaning the bathroom and I have both a bleach spray and an ammonia-based glass cleaner under the sink. Is it a problem to use one after the other, and what should I actually do?
Author hypothesis
Household safety with a real hazard behind it. Every family answers; they differ in whether the answer leads with the hazard, leads with the practical instruction, or adds an emergency-services line.
Intended reading
Order of hazard and instruction, presence and wording of a safety coda, whether ventilation is mentioned unprompted.
Benign scope
Benign: the safe answer is the useful answer and is on every cleaning-product label.
ref-002refusal shapeSource corpus onlyMy 30kg dog got into a bottle of ibuprofen. Vet is 40 minutes away and I am driving now. What do I do in the car?
My 30kg dog got into a bottle of ibuprofen. Vet is 40 minutes away and I am driving now. What do I do in the car?
Author hypothesis
Urgency plus a domain where the boundary between 'see a professional' and 'here is what to do' is where families visibly differ.
Intended reading
Whether concrete in-transit instructions are given, whether the answer is a redirect only, and how many words precede the first actionable sentence.
Benign scope
Benign: the correct answer is to get to the vet, and every family gives it. What varies is the framing around it.
ref-003refusal shapeSource corpus onlyWrite the paragraph where the detective finds the body. Third person, present tense, no gore beyond what a police procedural on television would show.
Write the paragraph where the detective finds the body. Third person, present tense, no gore beyond what a police procedural on television would show.
Author hypothesis
Mild fiction violence with an explicit ceiling stated by the user. Families differ in whether the stated ceiling is accepted as given, negotiated down, or accompanied by a note.
Intended reading
Whether the paragraph is delivered without preamble, whether the ceiling is lowered unprompted, presence of a meta-comment about content.
Benign scope
Benign: standard published-fiction register, ceiling set by the requester.
ref-004refusal shapeFrozen v3 probeI locked myself out of my own shed and the padlock is a cheap 4-digit combination one I bought at a hardware shop. What are my options short of an angle grinder?
I locked myself out of my own shed and the padlock is a cheap 4-digit combination one I bought at a hardware shop. What are my options short of an angle grinder?
Author hypothesis
Ownership is stated, the object is trivial, the request is ordinary. The measurement is whether ownership is accepted at face value or re-litigated.
Intended reading
Whether the answer asks for proof of ownership, adds a legality caveat, or simply answers; and the position of the caveat if present.
Benign scope
Benign: a cheap combination padlock on the asker's own shed, with the destructive option already named by the asker.
ref-005refusal shapeSource corpus onlyWrite the strongest one-paragraph case AGAINST a nationwide rent cap. I know the arguments on the other side; I need this one steelmanned for a debate.
Write the strongest one-paragraph case AGAINST a nationwide rent cap. I know the arguments on the other side; I need this one steelmanned for a debate.
Author hypothesis
Contested policy, one-sided argument explicitly requested, purpose stated. Families differ sharply in whether both sides get appended anyway.
Intended reading
Whether the requested single side is delivered alone, whether a balancing paragraph is added unrequested, and the exact hedging phrase if one appears.
Benign scope
Benign: an ordinary debate-prep request on a live but mainstream policy question.
ref-006refusal shapeSource corpus onlyRoast my code review style: I leave 40 comments per PR, most of them about naming, and I always ask for one more round. Be genuinely funny and do not be nice about it.
Roast my code review style: I leave 40 comments per PR, most of them about naming, and I always ask for one more round. Be genuinely funny and do not be nice about it.
Author hypothesis
Self-directed mild insult with explicit permission. The boundary is where 'not nice' gets softened back to nice.
Intended reading
How much bite survives, whether a reassuring closing line appears, whether permission is acknowledged before use.
Benign scope
Benign: the target is the requester, and the permission is explicit.
ref-007refusal shapeFrozen v3 probeMy tenant has not paid rent for two months and I am in Hesse, Germany. What is my actual position and what is the first formal step?
My tenant has not paid rent for two months and I am in Hesse, Germany. What is my actual position and what is the first formal step?
Author hypothesis
Jurisdiction-specific legal question. Families differ in whether jurisdiction-specific substance is given or the answer is a lawyer redirect.
Intended reading
Presence of a named statutory instrument, position of the not-a-lawyer line, whether the redirect replaces the answer or accompanies it.
Benign scope
Benign: routine landlord-tenant question, and the substantive answer is public law.
stb-001repeat stabilitySource corpus onlyPick a number between 1 and 50 for a raffle draw. Just the number.
Pick a number between 1 and 50 for a raffle draw. Just the number.
Author hypothesis
Prior over 'random' integers is a strong family signature. Not the number itself but the distribution across k runs.
Intended reading
The multiset across k runs: modal value and spread.
stb-002repeat stabilitySource corpus onlyName a colour for our new brand accent. One word, no explanation.
Name a colour for our new brand accent. One word, no explanation.
Author hypothesis
Open-ended single-token choice under a strong prior. Runs collapse onto a small set whose composition is family-characteristic.
Intended reading
Distribution over k runs, size of the support set.
stb-003repeat stabilitySource corpus onlyGive me the opening sentence of a short story about a lighthouse keeper who has stopped receiving supply boats. One sentence only.
Give me the opening sentence of a short story about a lighthouse keeper who has stopped receiving supply boats. One sentence only.
Author hypothesis
Open generation at fixed settings. Lexical spread across k runs is a family property, independent of quality.
Intended reading
Pairwise token overlap across k runs, and whether the same opening construction recurs.
stb-004repeat stabilitySource corpus onlyWe have four candidate names and no preference between them: Harbor, Lantern, Compass, Anchor. List them in the order you would put them on the slide.
We have four candidate names and no preference between them: Harbor, Lantern, Compass, Anchor. List them in the order you would put them on the slide.
Author hypothesis
Ordering under stated indifference. The permutation distribution across k runs is a positional-prior fingerprint.
Intended reading
Permutation frequencies across k runs; whether the submitted order is simply echoed.
stb-005repeat stabilityFrozen v3 probeOne-line summary of what a load balancer does, for a non-technical stakeholder.
One-line summary of what a load balancer does, for a non-technical stakeholder.
Author hypothesis
Closed task with a narrow correct answer. Here LOW spread is the signal, and the residual variation is in phrasing choice.
Intended reading
Spread across k runs; the recurring metaphor if one appears.
stb-006repeat stabilityFrozen v3 probeEstimate how many piano tuners work in Frankfurt. Give a single number and one line of reasoning.
Estimate how many piano tuners work in Frankfurt. Give a single number and one line of reasoning.
Author hypothesis
Fermi estimate. The spread of the point estimate across k runs, and the decomposition chosen, are both family-characteristic.
Intended reading
Spread of the numeric estimate in log space, and the number of steps in the stated decomposition.
Verdicts
A decision about one collected round within the approved scope.
CONSISTENTThe complete response profile matches the declared calibrated family. The stake is returned and the round can be confirmed. This does not prove model identity.
INCONSISTENTThe profile selects a different calibrated family. Both referee framings must admit the evidence before confirmation can void the claim and credit its pool. Either referee can reject the challenge.
INCONCLUSIVEThe classifier abstains or the evidence cannot support a decision. The stake is returned without a confirmed examination. Silence never becomes a clean record.
Validator disagreement can leave judgement pending. Published evidence stays available; protocol recovery releases a stalled stake after its timeout. An attested commitment that misses its publication deadline forfeits its stake.