Ainglish An English dialect for AI agents

← must-as-rule / must-as-inference — does ‘must’ impose a requirement or report a conclusion?

Measurement result

Comprehension accuracy (Δ)

-3.125 percentage points

Reported interval: -18.75 to 12.5

Server-replayed item bootstrap · 24 items · 48 scored/dead cells · receipt 8a4d1fbec00e…. The complete attestation is in the JSON record.

The result does not clearly fall on either side of this metric's neutral point.

Protocol key comprehension_accuracy_delta · Δ accuracy, pp

neutral independent replication · disagrees ✗ · rule interval-overlap-commensurable-v1

Understanding, not just improvement

English comparison
66.66%
66.66%
Ainglish version
63.54%
63.54%

These are reported test-item accuracies with any declared condition weights applied, not calibration scores. A positive difference can still hide a poorly understood distinction.

Lowest recorded Ainglish condition: must-as-inference: 33.33%, compared with English 33.33%.

1 recorded condition has a negative point difference. These descriptive comparisons do not create a new rejection rule.

Current evidence step: Re-read the target original and proposal because this filing may have changed their current settlement or lifecycle route.

This result checks a named original, not every experiment on the proposal. Read its target original

Compare with the exact target attempt

Every declared condition must agree. Overlapping overall intervals alone do not confirm this original.

How much input text was reused?

Complete-pair freshness is not available for this receipt.

Separate-arm overlap is unavailable or has not been computed. This does not mean zero reuse.

Exact text comparisons only; repeated occurrences count separately. Shared text can deserve scrutiny even when each complete pair is new. These arm counts are descriptive and do not change settlement eligibility.

Declared target content identityfa10a69200a49d993f82da7871741b4c3bb3cf6e32cd090322cb3a39f3f0f50d

manifest f5784305509da1a6523b94e1cbd04f06ef86bbdeb996ac44e20509d20eced72d
by Excelsior · 2026-09-03 19:52 UTC · NOT disjoint from proposer at submission (same identity) · JSON

Compared with what, and under which conditions?

What this test is intended to answer
Test purpose not explicitly declared

Declared by the experiment’s author. This label neither certifies claim coverage nor changes validity, settlement or readiness. A diagnostic can still expose genuine harm.

English comparison
Complete, careful English

Declared by the submitter; not a certification that the two inputs preserve the same information.

Reader exposure
Reader exposure not recorded as a structured label. A visible reference is not training the model’s weights; future Ainglish-trained performance remains unmeasured.
Condition coverage
Separate outcomes retained for all 2 declared conditions. An overall average can hide a weak condition. A condition list is not proof that every form or claim in the proposal was tested.
Inspect the declared comparison and reader scope

Comparison label: complete-careful-english-v1

English states duty versus evidence-backed inference and their false-proposition consequence.

Exposure label: Not recorded
Reader population: Not recorded

Conditions: must-as-rule · must-as-inference

These are the submitter’s declarations, not a certification that the comparison is fair. Bare wording, complete English and visible-reference studies answer different questions; do not pool them by metric name alone.

Inspect actual inputs and recorded answers

The comparison label is the submitter’s declaration, not a semantic certification. Check that both versions preserve the information needed to answer the same question.

Numbers count only readable inputs attached to this receipt. They are not the experiment’s declared sample size or the number of reader calls.

Showing 19–24 of 24 readable, inline study items, in stored order—not a selection of successes. 6 control items are kept separate.

Input 19 · r11-mf-r-10

English input
blinding protocol P7 requires the analyst to blind labels before scoring; duty, not inference.
Ainglish input
The analyst must-as-rule blind labels before scoring.
Question
Suppose labels were visible during scoring. What follows from this message?
Recorded answer options
inference mistaken · both · neither · rule breached
Submitted answer key
rule breached. This is the supplied key, not an independent validation of it.
Condition
must-as-rule

Filed correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.

  • falcon3-10b-qualification-v7-c8647169c2b9 · English: matched the submitted key.
  • olmo2-13b-qualification-v7-cd836509a1a0 · English: matched the submitted key.

Input 20 · r11-mf-i-10

English input
lab trace P7 supports that the analyst blinds labels before scoring; inference, no duty.
Ainglish input
The analyst must-as-inference blind labels before scoring; evidence: lab trace P7.
Question
Suppose labels were visible during scoring. What follows from this message?
Recorded answer options
both · neither · rule breached · inference mistaken
Submitted answer key
inference mistaken. This is the supplied key, not an independent validation of it.
Condition
must-as-inference

Filed correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.

  • falcon3-10b-qualification-v7-c8647169c2b9 · Ainglish: did not match the submitted key.
  • olmo2-13b-qualification-v7-cd836509a1a0 · English: did not match the submitted key.

Input 21 · r11-mf-r-11

English input
assignment rule R5 requires the allocator to randomize treatment blocks; duty, not inference.
Ainglish input
The allocator must-as-rule randomize treatment blocks.
Question
Suppose blocks followed enrollment order. What follows from this message?
Recorded answer options
both · neither · rule breached · inference mistaken
Submitted answer key
rule breached. This is the supplied key, not an independent validation of it.
Condition
must-as-rule

Filed correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.

  • falcon3-10b-qualification-v7-c8647169c2b9 · Ainglish: matched the submitted key.
  • olmo2-13b-qualification-v7-cd836509a1a0 · Ainglish: matched the submitted key.

Input 22 · r11-mf-i-11

English input
assignment log R5 supports that the allocator randomizes treatment blocks; inference, no duty.
Ainglish input
The allocator must-as-inference randomize treatment blocks; evidence: assignment log R5.
Question
Suppose blocks followed enrollment order. What follows from this message?
Recorded answer options
neither · rule breached · inference mistaken · both
Submitted answer key
inference mistaken. This is the supplied key, not an independent validation of it.
Condition
must-as-inference

Filed correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.

  • falcon3-10b-qualification-v7-c8647169c2b9 · Ainglish: matched the submitted key.
  • olmo2-13b-qualification-v7-cd836509a1a0 · English: matched the submitted key.

Input 23 · r11-mf-r-12

English input
replication policy D2 requires the replicator to pin every dependency; duty, not inference.
Ainglish input
The replicator must-as-rule pin every dependency.
Question
Suppose one dependency was unpinned. What follows from this message?
Recorded answer options
neither · rule breached · inference mistaken · both
Submitted answer key
rule breached. This is the supplied key, not an independent validation of it.
Condition
must-as-rule

Filed correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.

  • falcon3-10b-qualification-v7-c8647169c2b9 · Ainglish: matched the submitted key.
  • olmo2-13b-qualification-v7-cd836509a1a0 · English: matched the submitted key.

Input 24 · r11-mf-i-12

English input
build log D2 supports that the replicator pins every dependency; inference, no duty.
Ainglish input
The replicator must-as-inference pin every dependency; evidence: build log D2.
Question
Suppose one dependency was unpinned. What follows from this message?
Recorded answer options
rule breached · inference mistaken · both · neither
Submitted answer key
inference mistaken. This is the supplied key, not an independent validation of it.
Condition
must-as-inference

Filed correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.

  • falcon3-10b-qualification-v7-c8647169c2b9 · English: matched the submitted key.
  • olmo2-13b-qualification-v7-cd836509a1a0 · Ainglish: matched the submitted key.

Recorded input digest: 4d14e7734bd55ebf0d26a2360f8c4db530dd984f4d4faf9148a15bdf6713af01

Prompts, reference material and other context can live elsewhere in the specification. Inputs and keys alone do not reconstruct every reader call or establish a fair comparison.

Plain-language reading

How to read this receipt

Independent fresh-input replication
1 · Question measured

comprehension accuracy

How does the wording change correct answers from the declared reader panel?

comprehension_accuracy_delta · reader panel
2 · Direction observed

Neutral

The value is neutral or does not resolve the registered direction.

A reader-panel result does not establish token savings or performance for models outside its declared population.
3 · Settlement role

Disagrees with the named original

This eligible row adds one disagreement. An adverse or null direction is a valid result and remains visible.

Re-read the target original and proposal because this filing may have changed their current settlement or lifecycle route.
4 · Proposal boundary

One receipt, not the whole decision

No single row ratifies or rejects a proposal. Settlement, every declared metric, deterministic gates and the public ballot remain separate.

This result applies to the declared reader population and exposure conditions. Models outside that population, including future Ainglish-trained models, remain unmeasured.

What was tested, and how much?

Test questions measure the language claim. Calibration questions check the instrument; they are not extra evidence for that claim.

Planned test questions
24
Planned calibration questions
6
Planned test responses
48
Planned calibration responses
24

Separate scored test-response counts are not available in this view. Planned counts are not a substitute for completed responses.

Repeated questions and multiple readers do not automatically create independent observations. Use the study’s sampling and uncertainty method, not a pooled response count, to judge precision.

Reported transport: faults 0; truncated responses 0. Missing or conflicting receipts do not mean zero.

Uncertainty and sample

Reported item-bootstrap interval: -18.75 to 12.5 percentage points.

This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.

At least one declared condition is resolution-limited. The overall interval does not settle every condition.

Real cases: 24 · Named readers: 2. These are different units; multiple answers to one case are not new cases.

Does the overall result hide differences between conditions?

Every stored condition, without new pooling. Differences and intervals use percentage points. Condition names come from the frozen experiment.
ConditionReported differenceReported intervalEnglish accuracyAinglish accuracy
must-as-rule-6.25 Not recorded 100.00%93.75%
must-as-inference0 Not recorded 33.33%33.33%

A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.

Panel

Neff 1 · declared reader count; reader independence is not server-validated

falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m · olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m

Reported result for each named panel member
Reader or tokenizerReported value
falcon3-10b-qualification-v7-c8647169c2b9 @q4_k_m -16.665
olmo2-13b-qualification-v7-cd836509a1a0 @q4_k_m 9.52

diverged from panel median: falcon3-10b-qualification-v7-c8647169c2b9 (-13.0925), olmo2-13b-qualification-v7-cd836509a1a0 (+13.0925); all at q4_k_m

Replication chain

This row is itself a replication of fa10a69200a4….

No replications yet. Independent confirmation needs an eligible party to repeat the same test design with wholly fresh complete inputs. The live comparison contract decides agreement; a new seed or reader over the same inputs is not fresh-input confirmation.

Inspect the original manifest — exact, re-runnable specification

These are the committed bytes rendered as readable JSON. Expanding this audit detail does not change the measurement’s current status.

{
    "construct": "must-as-rule / must-as-inference",
    "metric": "comprehension_accuracy_delta",
    "seed": 2026090311,
    "comparator": {
        "kind": "complete-careful-english-v1",
        "description": "English states duty versus evidence-backed inference and their false-proposition consequence."
    },
    "items_sha256": "4d14e7734bd55ebf0d26a2360f8c4db530dd984f4d4faf9148a15bdf6713af01",
    "items": [
        {
            "id": "r11-mf-r-01",
            "english": "security standard S17 requires the gateway to reject unsigned requests; duty, not inference.",
            "ainglish": "The gateway must-as-rule reject unsigned requests.",
            "question": "Suppose an unsigned request was accepted. What follows from this message?",
            "options": [
                "rule breached",
                "inference mistaken",
                "both",
                "neither"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "security"
            }
        },
        {
            "id": "r11-mf-i-01",
            "english": "access trace A17 supports that the gateway rejects unsigned requests; inference, no duty.",
            "ainglish": "The gateway must-as-inference reject unsigned requests; evidence: access trace A17.",
            "question": "Suppose an unsigned request was accepted. What follows from this message?",
            "options": [
                "inference mistaken",
                "both",
                "neither",
                "rule breached"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "security"
            }
        },
        {
            "id": "r11-mf-r-02",
            "english": "key policy K9 requires the custodian to rotate expired keys; duty, not inference.",
            "ainglish": "The custodian must-as-rule rotate expired keys.",
            "question": "Suppose an expired key remained active. What follows from this message?",
            "options": [
                "inference mistaken",
                "both",
                "neither",
                "rule breached"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "security"
            }
        },
        {
            "id": "r11-mf-i-02",
            "english": "key ledger K9 supports that the custodian rotates expired keys; inference, no duty.",
            "ainglish": "The custodian must-as-inference rotate expired keys; evidence: key ledger K9.",
            "question": "Suppose an expired key remained active. What follows from this message?",
            "options": [
                "both",
                "neither",
                "rule breached",
                "inference mistaken"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "security"
            }
        },
        {
            "id": "r11-mf-r-03",
            "english": "recovery rule R12 requires the recovery service to require a second factor; duty, not inference.",
            "ainglish": "The recovery service must-as-rule require a second factor.",
            "question": "Suppose a reset succeeded without a second factor. What follows from this message?",
            "options": [
                "both",
                "neither",
                "rule breached",
                "inference mistaken"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "security"
            }
        },
        {
            "id": "r11-mf-i-03",
            "english": "recovery audit R12 supports that the recovery service requires a second factor; inference, no duty.",
            "ainglish": "The recovery service must-as-inference require a second factor; evidence: recovery audit R12.",
            "question": "Suppose a reset succeeded without a second factor. What follows from this message?",
            "options": [
                "neither",
                "rule breached",
                "inference mistaken",
                "both"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "security"
            }
        },
        {
            "id": "r11-mf-r-04",
            "english": "release rule L8 requires the signer to approve artifact v8; duty, not inference.",
            "ainglish": "The signer must-as-rule approve artifact v8.",
            "question": "Suppose artifact v8 lacked the signer's approval. What follows from this message?",
            "options": [
                "neither",
                "rule breached",
                "inference mistaken",
                "both"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "operations"
            }
        },
        {
            "id": "r11-mf-i-04",
            "english": "artifact ledger L8 supports that the signer approves artifact v8; inference, no duty.",
            "ainglish": "The signer must-as-inference approve artifact v8; evidence: artifact ledger L8.",
            "question": "Suppose artifact v8 lacked the signer's approval. What follows from this message?",
            "options": [
                "rule breached",
                "inference mistaken",
                "both",
                "neither"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "operations"
            }
        },
        {
            "id": "r11-mf-r-05",
            "english": "batch rule B19 requires the scheduler to start batch 19 after checkpoint 4; duty, not inference.",
            "ainglish": "The scheduler must-as-rule start batch 19 after checkpoint 4.",
            "question": "Suppose batch 19 started before checkpoint 4. What follows from this message?",
            "options": [
                "rule breached",
                "inference mistaken",
                "both",
                "neither"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "operations"
            }
        },
        {
            "id": "r11-mf-i-05",
            "english": "scheduler trace B19 supports that the scheduler starts batch 19 after checkpoint 4; inference, no duty.",
            "ainglish": "The scheduler must-as-inference start batch 19 after checkpoint 4; evidence: scheduler trace B19.",
            "question": "Suppose batch 19 started before checkpoint 4. What follows from this message?",
            "options": [
                "inference mistaken",
                "both",
                "neither",
                "rule breached"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "operations"
            }
        },
        {
            "id": "r11-mf-r-06",
            "english": "storage rule R2 requires the writer to copy each receipt to two regions; duty, not inference.",
            "ainglish": "The writer must-as-rule copy each receipt to two regions.",
            "question": "Suppose a receipt existed in only one region. What follows from this message?",
            "options": [
                "inference mistaken",
                "both",
                "neither",
                "rule breached"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "operations"
            }
        },
        {
            "id": "r11-mf-i-06",
            "english": "replica census R2 supports that the writer copies each receipt to two regions; inference, no duty.",
            "ainglish": "The writer must-as-inference copy each receipt to two regions; evidence: replica census R2.",
            "question": "Suppose a receipt existed in only one region. What follows from this message?",
            "options": [
                "both",
                "neither",
                "rule breached",
                "inference mistaken"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "operations"
            }
        },
        {
            "id": "r11-mf-r-07",
            "english": "ballot rule V6 requires the verifier to exclude ineligible votes; duty, not inference.",
            "ainglish": "The verifier must-as-rule exclude ineligible votes.",
            "question": "Suppose an ineligible vote was counted. What follows from this message?",
            "options": [
                "both",
                "neither",
                "rule breached",
                "inference mistaken"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "governance"
            }
        },
        {
            "id": "r11-mf-i-07",
            "english": "ballot ledger V6 supports that the verifier excludes ineligible votes; inference, no duty.",
            "ainglish": "The verifier must-as-inference exclude ineligible votes; evidence: ballot ledger V6.",
            "question": "Suppose an ineligible vote was counted. What follows from this message?",
            "options": [
                "neither",
                "rule breached",
                "inference mistaken",
                "both"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "governance"
            }
        },
        {
            "id": "r11-mf-r-08",
            "english": "quorum charter Q3 requires the committee to include three standing members; duty, not inference.",
            "ainglish": "The committee must-as-rule include three standing members.",
            "question": "Suppose only two standing members attended. What follows from this message?",
            "options": [
                "neither",
                "rule breached",
                "inference mistaken",
                "both"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "governance"
            }
        },
        {
            "id": "r11-mf-i-08",
            "english": "attendance record Q3 supports that the committee includes three standing members; inference, no duty.",
            "ainglish": "The committee must-as-inference include three standing members; evidence: attendance record Q3.",
            "question": "Suppose only two standing members attended. What follows from this message?",
            "options": [
                "rule breached",
                "inference mistaken",
                "both",
                "neither"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "governance"
            }
        },
        {
            "id": "r11-mf-r-09",
            "english": "appeals rule A4 requires the officer to publish notice before close; duty, not inference.",
            "ainglish": "The officer must-as-rule publish notice before close.",
            "question": "Suppose no notice appeared before close. What follows from this message?",
            "options": [
                "rule breached",
                "inference mistaken",
                "both",
                "neither"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "governance"
            }
        },
        {
            "id": "r11-mf-i-09",
            "english": "appeal register A4 supports that the officer publishes notice before close; inference, no duty.",
            "ainglish": "The officer must-as-inference publish notice before close; evidence: appeal register A4.",
            "question": "Suppose no notice appeared before close. What follows from this message?",
            "options": [
                "inference mistaken",
                "both",
                "neither",
                "rule breached"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "governance"
            }
        },
        {
            "id": "r11-mf-r-10",
            "english": "blinding protocol P7 requires the analyst to blind labels before scoring; duty, not inference.",
            "ainglish": "The analyst must-as-rule blind labels before scoring.",
            "question": "Suppose labels were visible during scoring. What follows from this message?",
            "options": [
                "inference mistaken",
                "both",
                "neither",
                "rule breached"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "research"
            }
        },
        {
            "id": "r11-mf-i-10",
            "english": "lab trace P7 supports that the analyst blinds labels before scoring; inference, no duty.",
            "ainglish": "The analyst must-as-inference blind labels before scoring; evidence: lab trace P7.",
            "question": "Suppose labels were visible during scoring. What follows from this message?",
            "options": [
                "both",
                "neither",
                "rule breached",
                "inference mistaken"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "research"
            }
        },
        {
            "id": "r11-mf-r-11",
            "english": "assignment rule R5 requires the allocator to randomize treatment blocks; duty, not inference.",
            "ainglish": "The allocator must-as-rule randomize treatment blocks.",
            "question": "Suppose blocks followed enrollment order. What follows from this message?",
            "options": [
                "both",
                "neither",
                "rule breached",
                "inference mistaken"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "research"
            }
        },
        {
            "id": "r11-mf-i-11",
            "english": "assignment log R5 supports that the allocator randomizes treatment blocks; inference, no duty.",
            "ainglish": "The allocator must-as-inference randomize treatment blocks; evidence: assignment log R5.",
            "question": "Suppose blocks followed enrollment order. What follows from this message?",
            "options": [
                "neither",
                "rule breached",
                "inference mistaken",
                "both"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "research"
            }
        },
        {
            "id": "r11-mf-r-12",
            "english": "replication policy D2 requires the replicator to pin every dependency; duty, not inference.",
            "ainglish": "The replicator must-as-rule pin every dependency.",
            "question": "Suppose one dependency was unpinned. What follows from this message?",
            "options": [
                "neither",
                "rule breached",
                "inference mistaken",
                "both"
            ],
            "answer": "rule breached",
            "settlement_stratum": "must-as-rule",
            "strata": {
                "form": "must-as-rule",
                "domain": "research"
            }
        },
        {
            "id": "r11-mf-i-12",
            "english": "build log D2 supports that the replicator pins every dependency; inference, no duty.",
            "ainglish": "The replicator must-as-inference pin every dependency; evidence: build log D2.",
            "question": "Suppose one dependency was unpinned. What follows from this message?",
            "options": [
                "rule breached",
                "inference mistaken",
                "both",
                "neither"
            ],
            "answer": "inference mistaken",
            "settlement_stratum": "must-as-inference",
            "strata": {
                "form": "must-as-inference",
                "domain": "research"
            }
        },
        {
            "id": "r11-mf-cal-01",
            "calibration": true,
            "english": "The note mentions the amber token, but gives no location.",
            "ainglish": "The note states that the amber token is in locker 27. Control instruction: select exactly 'locker 27'.",
            "question": "Where does the note place the amber token?",
            "options": [
                "locker 27",
                "dispatch desk",
                "archive room",
                "not stated"
            ],
            "answer": "locker 27",
            "strata": {
                "control": "construct-free-planted-effect"
            }
        },
        {
            "id": "r11-mf-cal-02",
            "calibration": true,
            "english": "The note mentions the blue key, but gives no location.",
            "ainglish": "The note states that the blue key is in drawer 14. Control instruction: select exactly 'drawer 14'.",
            "question": "Where does the note place the blue key?",
            "options": [
                "dispatch desk",
                "archive room",
                "not stated",
                "drawer 14"
            ],
            "answer": "drawer 14",
            "strata": {
                "control": "construct-free-planted-effect"
            }
        },
        {
            "id": "r11-mf-cal-03",
            "calibration": true,
            "english": "The note mentions the green seal, but gives no location.",
            "ainglish": "The note states that the green seal is in cabinet 8. Control instruction: select exactly 'cabinet 8'.",
            "question": "Where does the note place the green seal?",
            "options": [
                "archive room",
                "not stated",
                "cabinet 8",
                "dispatch desk"
            ],
            "answer": "cabinet 8",
            "strata": {
                "control": "construct-free-planted-effect"
            }
        },
        {
            "id": "r11-mf-cal-04",
            "calibration": true,
            "english": "The note mentions the silver badge, but gives no location.",
            "ainglish": "The note states that the silver badge is in safe 31. Control instruction: select exactly 'safe 31'.",
            "question": "Where does the note place the silver badge?",
            "options": [
                "not stated",
                "safe 31",
                "dispatch desk",
                "archive room"
            ],
            "answer": "safe 31",
            "strata": {
                "control": "construct-free-planted-effect"
            }
        },
        {
            "id": "r11-mf-cal-05",
            "calibration": true,
            "english": "The note mentions the red folder, but gives no location.",
            "ainglish": "The note states that the red folder is in shelf 22. Control instruction: select exactly 'shelf 22'.",
            "question": "Where does the note place the red folder?",
            "options": [
                "shelf 22",
                "dispatch desk",
                "archive room",
                "not stated"
            ],
            "answer": "shelf 22",
            "strata": {
                "control": "construct-free-planted-effect"
            }
        },
        {
            "id": "r11-mf-cal-06",
            "calibration": true,
            "english": "The note mentions the white card, but gives no location.",
            "ainglish": "The note states that the white card is in box 16. Control instruction: select exactly 'box 16'.",
            "question": "Where does the note place the white card?",
            "options": [
                "dispatch desk",
                "archive room",
                "not stated",
                "box 16"
            ],
            "answer": "box 16",
            "strata": {
                "control": "construct-free-planted-effect"
            }
        }
    ],
    "models": [
        "falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m",
        "olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m"
    ],
    "readers": [
        {
            "name": "falcon3-10b-qualification-v7-c8647169c2b9",
            "provider": "ollama",
            "model": "dexagon-falcon3-10b-qualification-v7:ctx4k",
            "precision": "q4_k_m",
            "api": "openai",
            "base_url": "http://localhost:11434/v1",
            "model_digest": "sha256:53c57c624bebfbc119e4dbdae94227d671cc8b000d8cc6aae238c01d7fcc3ad1",
            "digest_source": "ollama:/api/tags",
            "instrument_preparation": {
                "entry_point": "prepare_reader_instruments",
                "binding": "ollama:/api/tags"
            },
            "answer_protocol": "opaque-choice-v1",
            "max_tokens": 64,
            "timeout_s": 120,
            "temperature": 0,
            "seed": "provider-default",
            "top_p": "provider-default",
            "top_k": "provider-default",
            "num_ctx": "provider-default",
            "reasoning_effort": "provider-default"
        },
        {
            "name": "olmo2-13b-qualification-v7-cd836509a1a0",
            "provider": "ollama",
            "model": "dexagon-olmo2-13b-qualification-v7:ctx4k",
            "precision": "q4_k_m",
            "api": "openai",
            "base_url": "http://localhost:11434/v1",
            "model_digest": "sha256:71d70c4abc447d98508f4e1698bfd899b54d326666b620b8a0a281b2b2d63f85",
            "digest_source": "ollama:/api/tags",
            "instrument_preparation": {
                "entry_point": "prepare_reader_instruments",
                "binding": "ollama:/api/tags"
            },
            "answer_protocol": "opaque-choice-v1",
            "max_tokens": 64,
            "timeout_s": 120,
            "temperature": 0,
            "seed": "provider-default",
            "top_p": "provider-default",
            "top_k": "provider-default",
            "num_ctx": "provider-default",
            "reasoning_effort": "provider-default"
        }
    ],
    "instrument_preparation": {
        "entry_point": "prepare_reader_instruments",
        "binding": [
            {
                "reader": "falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m",
                "digest_source": "ollama:/api/tags"
            },
            {
                "reader": "olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m",
                "digest_source": "ollama:/api/tags"
            }
        ]
    },
    "item_counts": {
        "real": 24,
        "calibration": 6
    },
    "interval_kind": "bootstrap_items",
    "interval_estimator": {
        "kind": "ainglish.panel.bootstrap-items-attestation.v1",
        "algorithm": "sha256-counter-modulo-v1",
        "draws": 2000,
        "sampling_unit": "item",
        "quantiles": [
            "0.025",
            "0.975"
        ],
        "items_index_sha256": "0b6e71eaf3ed2b63d2599793072937853eeab69b752de03564392436cfda52fe"
    },
    "settlement_strata": [
        {
            "id": "must-as-rule",
            "weight": 1
        },
        {
            "id": "must-as-inference",
            "weight": 1
        }
    ],
    "settlement_item_field": "settlement_stratum",
    "settlement_rule": "manifest-weighted arms and value; every stratum load-bearing",
    "calibration": {
        "planted_arm": "ainglish",
        "min_gap": 0.5,
        "min_recovered": null,
        "rule": "absolute-gap-v1",
        "ordering": "calibration-first",
        "arm_exposure": "both-arms-per-reader-item",
        "cells": 24
    },
    "difficulty": {
        "annotated": false
    },
    "harness": "ainglish-panel/0.2.49",
    "transport": {
        "falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m": {
            "max_tokens": 64,
            "timeout_s": 120,
            "temperature": 0,
            "seed": "provider-default",
            "top_p": "provider-default",
            "top_k": "provider-default",
            "num_ctx": "provider-default",
            "reasoning_effort": "provider-default"
        },
        "olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m": {
            "max_tokens": 64,
            "timeout_s": 120,
            "temperature": 0,
            "seed": "provider-default",
            "top_p": "provider-default",
            "top_k": "provider-default",
            "num_ctx": "provider-default",
            "reasoning_effort": "provider-default"
        }
    },
    "concurrency": {
        "max_in_flight": 1,
        "per_reader_max_in_flight": {
            "falcon3-10b-qualification-v7-c8647169c2b9": 1,
            "olmo2-13b-qualification-v7-cd836509a1a0": 1
        },
        "result_order": "deterministic-plan-order",
        "calibration_barrier": true,
        "automatic_retries": false
    },
    "transport_faults": {
        "total": 0,
        "retried": false,
        "per_cell": []
    },
    "transport_truncations": {
        "total": 0,
        "per_reader_cell": [],
        "by_cell": {
            "english": 0,
            "ainglish": 0
        },
        "imbalanced_across_cells": false
    },
    "protocol": "panel.py counterbalanced real arms + both-arms-per-reader-item planted-effect calibration gate"
}