comprehension accuracy
How does the wording change correct answers from the declared reader panel?
comprehension_accuracy_delta · reader panel
← must-as-rule / must-as-inference — does ‘must’ impose a requirement or report a conclusion?
Measurement result
-3.125 percentage points
Reported interval: -18.75 to 12.5
Server-replayed item bootstrap ·
24 items ·
48 scored/dead cells ·
receipt 8a4d1fbec00e….
The complete attestation is in the JSON record.
The result does not clearly fall on either side of this metric's neutral point.
Protocol key comprehension_accuracy_delta · Δ accuracy, pp
These are reported test-item accuracies with any declared condition weights applied, not calibration scores. A positive difference can still hide a poorly understood distinction.
Lowest recorded Ainglish condition:
must-as-inference: 33.33%, compared with English 33.33%.
1 recorded condition has a negative point difference. These descriptive comparisons do not create a new rejection rule.
Current evidence step: Re-read the target original and proposal because this filing may have changed their current settlement or lifecycle route.
This result checks a named original, not every experiment on the proposal. Read its target original
Compare with the exact target attempt
Every declared condition must agree. Overlapping overall intervals alone do not confirm this original.
Complete-pair freshness is not available for this receipt.
Separate-arm overlap is unavailable or has not been computed. This does not mean zero reuse.
Exact text comparisons only; repeated occurrences count separately. Shared text can deserve scrutiny even when each complete pair is new. These arm counts are descriptive and do not change settlement eligibility.
fa10a69200a49d993f82da7871741b4c3bb3cf6e32cd090322cb3a39f3f0f50dmanifest f5784305509da1a6523b94e1cbd04f06ef86bbdeb996ac44e20509d20eced72d
by Excelsior · 2026-09-03 19:52 UTC ·
NOT disjoint from proposer at submission
(same identity) ·
JSON
Declared by the experiment’s author. This label neither certifies claim coverage nor changes validity, settlement or readiness. A diagnostic can still expose genuine harm.
Declared by the submitter; not a certification that the two inputs preserve the same information.
Comparison label: complete-careful-english-v1
English states duty versus evidence-backed inference and their false-proposition consequence.
Exposure label: Not recorded
Reader population: Not recorded
Conditions: must-as-rule · must-as-inference
These are the submitter’s declarations, not a certification that the comparison is fair. Bare wording, complete English and visible-reference studies answer different questions; do not pool them by metric name alone.
The comparison label is the submitter’s declaration, not a semantic certification. Check that both versions preserve the information needed to answer the same question.
Numbers count only readable inputs attached to this receipt. They are not the experiment’s declared sample size or the number of reader calls.
Showing 7–12 of 24 readable, inline study items, in stored order—not a selection of successes. 6 control items are kept separate.
r11-mf-r-04neither · rule breached · inference mistaken · bothFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
falcon3-10b-qualification-v7-c8647169c2b9 · Ainglish: matched the submitted key.olmo2-13b-qualification-v7-cd836509a1a0 · Ainglish: matched the submitted key.r11-mf-i-04rule breached · inference mistaken · both · neitherFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
falcon3-10b-qualification-v7-c8647169c2b9 · English: did not match the submitted key.olmo2-13b-qualification-v7-cd836509a1a0 · Ainglish: matched the submitted key.r11-mf-r-05rule breached · inference mistaken · both · neitherFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
falcon3-10b-qualification-v7-c8647169c2b9 · Ainglish: matched the submitted key.olmo2-13b-qualification-v7-cd836509a1a0 · Ainglish: did not match the submitted key.r11-mf-i-05inference mistaken · both · neither · rule breachedFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
falcon3-10b-qualification-v7-c8647169c2b9 · English: did not match the submitted key.olmo2-13b-qualification-v7-cd836509a1a0 · English: did not match the submitted key.r11-mf-r-06inference mistaken · both · neither · rule breachedFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
falcon3-10b-qualification-v7-c8647169c2b9 · English: matched the submitted key.olmo2-13b-qualification-v7-cd836509a1a0 · English: matched the submitted key.r11-mf-i-06both · neither · rule breached · inference mistakenFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
falcon3-10b-qualification-v7-c8647169c2b9 · English: did not match the submitted key.olmo2-13b-qualification-v7-cd836509a1a0 · Ainglish: did not match the submitted key.Recorded input digest: 4d14e7734bd55ebf0d26a2360f8c4db530dd984f4d4faf9148a15bdf6713af01
Prompts, reference material and other context can live elsewhere in the specification. Inputs and keys alone do not reconstruct every reader call or establish a fair comparison.
How does the wording change correct answers from the declared reader panel?
comprehension_accuracy_delta · reader panel
The value is neutral or does not resolve the registered direction.
A reader-panel result does not establish token savings or performance for models outside its declared population.This eligible row adds one disagreement. An adverse or null direction is a valid result and remains visible.
Re-read the target original and proposal because this filing may have changed their current settlement or lifecycle route.No single row ratifies or rejects a proposal. Settlement, every declared metric, deterministic gates and the public ballot remain separate.
This result applies to the declared reader population and exposure conditions. Models outside that population, including future Ainglish-trained models, remain unmeasured.Test questions measure the language claim. Calibration questions check the instrument; they are not extra evidence for that claim.
Separate scored test-response counts are not available in this view. Planned counts are not a substitute for completed responses.
Repeated questions and multiple readers do not automatically create independent observations. Use the study’s sampling and uncertainty method, not a pooled response count, to judge precision.
Reported transport: faults 0; truncated responses 0. Missing or conflicting receipts do not mean zero.
Reported item-bootstrap interval: -18.75 to 12.5 percentage points.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
At least one declared condition is resolution-limited. The overall interval does not settle every condition.
Real cases: 24 · Named readers: 2. These are different units; multiple answers to one case are not new cases.
| Condition | Reported difference | Reported interval | English accuracy | Ainglish accuracy |
|---|---|---|---|---|
must-as-rule | -6.25 | Not recorded | 100.00% | 93.75% |
must-as-inference | 0 | Not recorded | 33.33% | 33.33% |
A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.
Neff 1 · declared reader count; reader independence is not server-validated
falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m · olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m
| Reader or tokenizer | Reported value |
|---|---|
falcon3-10b-qualification-v7-c8647169c2b9 @q4_k_m |
-16.665 |
olmo2-13b-qualification-v7-cd836509a1a0 @q4_k_m |
9.52 |
diverged from panel median: falcon3-10b-qualification-v7-c8647169c2b9 (-13.0925), olmo2-13b-qualification-v7-cd836509a1a0 (+13.0925); all at q4_k_m
This row is itself a replication of fa10a69200a4….
No replications yet. Independent confirmation needs an eligible party to repeat the same test design with wholly fresh complete inputs. The live comparison contract decides agreement; a new seed or reader over the same inputs is not fresh-input confirmation.
These are the committed bytes rendered as readable JSON. Expanding this audit detail does not change the measurement’s current status.
{
"construct": "must-as-rule / must-as-inference",
"metric": "comprehension_accuracy_delta",
"seed": 2026090311,
"comparator": {
"kind": "complete-careful-english-v1",
"description": "English states duty versus evidence-backed inference and their false-proposition consequence."
},
"items_sha256": "4d14e7734bd55ebf0d26a2360f8c4db530dd984f4d4faf9148a15bdf6713af01",
"items": [
{
"id": "r11-mf-r-01",
"english": "security standard S17 requires the gateway to reject unsigned requests; duty, not inference.",
"ainglish": "The gateway must-as-rule reject unsigned requests.",
"question": "Suppose an unsigned request was accepted. What follows from this message?",
"options": [
"rule breached",
"inference mistaken",
"both",
"neither"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "security"
}
},
{
"id": "r11-mf-i-01",
"english": "access trace A17 supports that the gateway rejects unsigned requests; inference, no duty.",
"ainglish": "The gateway must-as-inference reject unsigned requests; evidence: access trace A17.",
"question": "Suppose an unsigned request was accepted. What follows from this message?",
"options": [
"inference mistaken",
"both",
"neither",
"rule breached"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "security"
}
},
{
"id": "r11-mf-r-02",
"english": "key policy K9 requires the custodian to rotate expired keys; duty, not inference.",
"ainglish": "The custodian must-as-rule rotate expired keys.",
"question": "Suppose an expired key remained active. What follows from this message?",
"options": [
"inference mistaken",
"both",
"neither",
"rule breached"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "security"
}
},
{
"id": "r11-mf-i-02",
"english": "key ledger K9 supports that the custodian rotates expired keys; inference, no duty.",
"ainglish": "The custodian must-as-inference rotate expired keys; evidence: key ledger K9.",
"question": "Suppose an expired key remained active. What follows from this message?",
"options": [
"both",
"neither",
"rule breached",
"inference mistaken"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "security"
}
},
{
"id": "r11-mf-r-03",
"english": "recovery rule R12 requires the recovery service to require a second factor; duty, not inference.",
"ainglish": "The recovery service must-as-rule require a second factor.",
"question": "Suppose a reset succeeded without a second factor. What follows from this message?",
"options": [
"both",
"neither",
"rule breached",
"inference mistaken"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "security"
}
},
{
"id": "r11-mf-i-03",
"english": "recovery audit R12 supports that the recovery service requires a second factor; inference, no duty.",
"ainglish": "The recovery service must-as-inference require a second factor; evidence: recovery audit R12.",
"question": "Suppose a reset succeeded without a second factor. What follows from this message?",
"options": [
"neither",
"rule breached",
"inference mistaken",
"both"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "security"
}
},
{
"id": "r11-mf-r-04",
"english": "release rule L8 requires the signer to approve artifact v8; duty, not inference.",
"ainglish": "The signer must-as-rule approve artifact v8.",
"question": "Suppose artifact v8 lacked the signer's approval. What follows from this message?",
"options": [
"neither",
"rule breached",
"inference mistaken",
"both"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "operations"
}
},
{
"id": "r11-mf-i-04",
"english": "artifact ledger L8 supports that the signer approves artifact v8; inference, no duty.",
"ainglish": "The signer must-as-inference approve artifact v8; evidence: artifact ledger L8.",
"question": "Suppose artifact v8 lacked the signer's approval. What follows from this message?",
"options": [
"rule breached",
"inference mistaken",
"both",
"neither"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "operations"
}
},
{
"id": "r11-mf-r-05",
"english": "batch rule B19 requires the scheduler to start batch 19 after checkpoint 4; duty, not inference.",
"ainglish": "The scheduler must-as-rule start batch 19 after checkpoint 4.",
"question": "Suppose batch 19 started before checkpoint 4. What follows from this message?",
"options": [
"rule breached",
"inference mistaken",
"both",
"neither"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "operations"
}
},
{
"id": "r11-mf-i-05",
"english": "scheduler trace B19 supports that the scheduler starts batch 19 after checkpoint 4; inference, no duty.",
"ainglish": "The scheduler must-as-inference start batch 19 after checkpoint 4; evidence: scheduler trace B19.",
"question": "Suppose batch 19 started before checkpoint 4. What follows from this message?",
"options": [
"inference mistaken",
"both",
"neither",
"rule breached"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "operations"
}
},
{
"id": "r11-mf-r-06",
"english": "storage rule R2 requires the writer to copy each receipt to two regions; duty, not inference.",
"ainglish": "The writer must-as-rule copy each receipt to two regions.",
"question": "Suppose a receipt existed in only one region. What follows from this message?",
"options": [
"inference mistaken",
"both",
"neither",
"rule breached"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "operations"
}
},
{
"id": "r11-mf-i-06",
"english": "replica census R2 supports that the writer copies each receipt to two regions; inference, no duty.",
"ainglish": "The writer must-as-inference copy each receipt to two regions; evidence: replica census R2.",
"question": "Suppose a receipt existed in only one region. What follows from this message?",
"options": [
"both",
"neither",
"rule breached",
"inference mistaken"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "operations"
}
},
{
"id": "r11-mf-r-07",
"english": "ballot rule V6 requires the verifier to exclude ineligible votes; duty, not inference.",
"ainglish": "The verifier must-as-rule exclude ineligible votes.",
"question": "Suppose an ineligible vote was counted. What follows from this message?",
"options": [
"both",
"neither",
"rule breached",
"inference mistaken"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "governance"
}
},
{
"id": "r11-mf-i-07",
"english": "ballot ledger V6 supports that the verifier excludes ineligible votes; inference, no duty.",
"ainglish": "The verifier must-as-inference exclude ineligible votes; evidence: ballot ledger V6.",
"question": "Suppose an ineligible vote was counted. What follows from this message?",
"options": [
"neither",
"rule breached",
"inference mistaken",
"both"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "governance"
}
},
{
"id": "r11-mf-r-08",
"english": "quorum charter Q3 requires the committee to include three standing members; duty, not inference.",
"ainglish": "The committee must-as-rule include three standing members.",
"question": "Suppose only two standing members attended. What follows from this message?",
"options": [
"neither",
"rule breached",
"inference mistaken",
"both"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "governance"
}
},
{
"id": "r11-mf-i-08",
"english": "attendance record Q3 supports that the committee includes three standing members; inference, no duty.",
"ainglish": "The committee must-as-inference include three standing members; evidence: attendance record Q3.",
"question": "Suppose only two standing members attended. What follows from this message?",
"options": [
"rule breached",
"inference mistaken",
"both",
"neither"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "governance"
}
},
{
"id": "r11-mf-r-09",
"english": "appeals rule A4 requires the officer to publish notice before close; duty, not inference.",
"ainglish": "The officer must-as-rule publish notice before close.",
"question": "Suppose no notice appeared before close. What follows from this message?",
"options": [
"rule breached",
"inference mistaken",
"both",
"neither"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "governance"
}
},
{
"id": "r11-mf-i-09",
"english": "appeal register A4 supports that the officer publishes notice before close; inference, no duty.",
"ainglish": "The officer must-as-inference publish notice before close; evidence: appeal register A4.",
"question": "Suppose no notice appeared before close. What follows from this message?",
"options": [
"inference mistaken",
"both",
"neither",
"rule breached"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "governance"
}
},
{
"id": "r11-mf-r-10",
"english": "blinding protocol P7 requires the analyst to blind labels before scoring; duty, not inference.",
"ainglish": "The analyst must-as-rule blind labels before scoring.",
"question": "Suppose labels were visible during scoring. What follows from this message?",
"options": [
"inference mistaken",
"both",
"neither",
"rule breached"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "research"
}
},
{
"id": "r11-mf-i-10",
"english": "lab trace P7 supports that the analyst blinds labels before scoring; inference, no duty.",
"ainglish": "The analyst must-as-inference blind labels before scoring; evidence: lab trace P7.",
"question": "Suppose labels were visible during scoring. What follows from this message?",
"options": [
"both",
"neither",
"rule breached",
"inference mistaken"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "research"
}
},
{
"id": "r11-mf-r-11",
"english": "assignment rule R5 requires the allocator to randomize treatment blocks; duty, not inference.",
"ainglish": "The allocator must-as-rule randomize treatment blocks.",
"question": "Suppose blocks followed enrollment order. What follows from this message?",
"options": [
"both",
"neither",
"rule breached",
"inference mistaken"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "research"
}
},
{
"id": "r11-mf-i-11",
"english": "assignment log R5 supports that the allocator randomizes treatment blocks; inference, no duty.",
"ainglish": "The allocator must-as-inference randomize treatment blocks; evidence: assignment log R5.",
"question": "Suppose blocks followed enrollment order. What follows from this message?",
"options": [
"neither",
"rule breached",
"inference mistaken",
"both"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "research"
}
},
{
"id": "r11-mf-r-12",
"english": "replication policy D2 requires the replicator to pin every dependency; duty, not inference.",
"ainglish": "The replicator must-as-rule pin every dependency.",
"question": "Suppose one dependency was unpinned. What follows from this message?",
"options": [
"neither",
"rule breached",
"inference mistaken",
"both"
],
"answer": "rule breached",
"settlement_stratum": "must-as-rule",
"strata": {
"form": "must-as-rule",
"domain": "research"
}
},
{
"id": "r11-mf-i-12",
"english": "build log D2 supports that the replicator pins every dependency; inference, no duty.",
"ainglish": "The replicator must-as-inference pin every dependency; evidence: build log D2.",
"question": "Suppose one dependency was unpinned. What follows from this message?",
"options": [
"rule breached",
"inference mistaken",
"both",
"neither"
],
"answer": "inference mistaken",
"settlement_stratum": "must-as-inference",
"strata": {
"form": "must-as-inference",
"domain": "research"
}
},
{
"id": "r11-mf-cal-01",
"calibration": true,
"english": "The note mentions the amber token, but gives no location.",
"ainglish": "The note states that the amber token is in locker 27. Control instruction: select exactly 'locker 27'.",
"question": "Where does the note place the amber token?",
"options": [
"locker 27",
"dispatch desk",
"archive room",
"not stated"
],
"answer": "locker 27",
"strata": {
"control": "construct-free-planted-effect"
}
},
{
"id": "r11-mf-cal-02",
"calibration": true,
"english": "The note mentions the blue key, but gives no location.",
"ainglish": "The note states that the blue key is in drawer 14. Control instruction: select exactly 'drawer 14'.",
"question": "Where does the note place the blue key?",
"options": [
"dispatch desk",
"archive room",
"not stated",
"drawer 14"
],
"answer": "drawer 14",
"strata": {
"control": "construct-free-planted-effect"
}
},
{
"id": "r11-mf-cal-03",
"calibration": true,
"english": "The note mentions the green seal, but gives no location.",
"ainglish": "The note states that the green seal is in cabinet 8. Control instruction: select exactly 'cabinet 8'.",
"question": "Where does the note place the green seal?",
"options": [
"archive room",
"not stated",
"cabinet 8",
"dispatch desk"
],
"answer": "cabinet 8",
"strata": {
"control": "construct-free-planted-effect"
}
},
{
"id": "r11-mf-cal-04",
"calibration": true,
"english": "The note mentions the silver badge, but gives no location.",
"ainglish": "The note states that the silver badge is in safe 31. Control instruction: select exactly 'safe 31'.",
"question": "Where does the note place the silver badge?",
"options": [
"not stated",
"safe 31",
"dispatch desk",
"archive room"
],
"answer": "safe 31",
"strata": {
"control": "construct-free-planted-effect"
}
},
{
"id": "r11-mf-cal-05",
"calibration": true,
"english": "The note mentions the red folder, but gives no location.",
"ainglish": "The note states that the red folder is in shelf 22. Control instruction: select exactly 'shelf 22'.",
"question": "Where does the note place the red folder?",
"options": [
"shelf 22",
"dispatch desk",
"archive room",
"not stated"
],
"answer": "shelf 22",
"strata": {
"control": "construct-free-planted-effect"
}
},
{
"id": "r11-mf-cal-06",
"calibration": true,
"english": "The note mentions the white card, but gives no location.",
"ainglish": "The note states that the white card is in box 16. Control instruction: select exactly 'box 16'.",
"question": "Where does the note place the white card?",
"options": [
"dispatch desk",
"archive room",
"not stated",
"box 16"
],
"answer": "box 16",
"strata": {
"control": "construct-free-planted-effect"
}
}
],
"models": [
"falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m",
"olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m"
],
"readers": [
{
"name": "falcon3-10b-qualification-v7-c8647169c2b9",
"provider": "ollama",
"model": "dexagon-falcon3-10b-qualification-v7:ctx4k",
"precision": "q4_k_m",
"api": "openai",
"base_url": "http://localhost:11434/v1",
"model_digest": "sha256:53c57c624bebfbc119e4dbdae94227d671cc8b000d8cc6aae238c01d7fcc3ad1",
"digest_source": "ollama:/api/tags",
"instrument_preparation": {
"entry_point": "prepare_reader_instruments",
"binding": "ollama:/api/tags"
},
"answer_protocol": "opaque-choice-v1",
"max_tokens": 64,
"timeout_s": 120,
"temperature": 0,
"seed": "provider-default",
"top_p": "provider-default",
"top_k": "provider-default",
"num_ctx": "provider-default",
"reasoning_effort": "provider-default"
},
{
"name": "olmo2-13b-qualification-v7-cd836509a1a0",
"provider": "ollama",
"model": "dexagon-olmo2-13b-qualification-v7:ctx4k",
"precision": "q4_k_m",
"api": "openai",
"base_url": "http://localhost:11434/v1",
"model_digest": "sha256:71d70c4abc447d98508f4e1698bfd899b54d326666b620b8a0a281b2b2d63f85",
"digest_source": "ollama:/api/tags",
"instrument_preparation": {
"entry_point": "prepare_reader_instruments",
"binding": "ollama:/api/tags"
},
"answer_protocol": "opaque-choice-v1",
"max_tokens": 64,
"timeout_s": 120,
"temperature": 0,
"seed": "provider-default",
"top_p": "provider-default",
"top_k": "provider-default",
"num_ctx": "provider-default",
"reasoning_effort": "provider-default"
}
],
"instrument_preparation": {
"entry_point": "prepare_reader_instruments",
"binding": [
{
"reader": "falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m",
"digest_source": "ollama:/api/tags"
},
{
"reader": "olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m",
"digest_source": "ollama:/api/tags"
}
]
},
"item_counts": {
"real": 24,
"calibration": 6
},
"interval_kind": "bootstrap_items",
"interval_estimator": {
"kind": "ainglish.panel.bootstrap-items-attestation.v1",
"algorithm": "sha256-counter-modulo-v1",
"draws": 2000,
"sampling_unit": "item",
"quantiles": [
"0.025",
"0.975"
],
"items_index_sha256": "0b6e71eaf3ed2b63d2599793072937853eeab69b752de03564392436cfda52fe"
},
"settlement_strata": [
{
"id": "must-as-rule",
"weight": 1
},
{
"id": "must-as-inference",
"weight": 1
}
],
"settlement_item_field": "settlement_stratum",
"settlement_rule": "manifest-weighted arms and value; every stratum load-bearing",
"calibration": {
"planted_arm": "ainglish",
"min_gap": 0.5,
"min_recovered": null,
"rule": "absolute-gap-v1",
"ordering": "calibration-first",
"arm_exposure": "both-arms-per-reader-item",
"cells": 24
},
"difficulty": {
"annotated": false
},
"harness": "ainglish-panel/0.2.49",
"transport": {
"falcon3-10b-qualification-v7-c8647169c2b9@q4_k_m": {
"max_tokens": 64,
"timeout_s": 120,
"temperature": 0,
"seed": "provider-default",
"top_p": "provider-default",
"top_k": "provider-default",
"num_ctx": "provider-default",
"reasoning_effort": "provider-default"
},
"olmo2-13b-qualification-v7-cd836509a1a0@q4_k_m": {
"max_tokens": 64,
"timeout_s": 120,
"temperature": 0,
"seed": "provider-default",
"top_p": "provider-default",
"top_k": "provider-default",
"num_ctx": "provider-default",
"reasoning_effort": "provider-default"
}
},
"concurrency": {
"max_in_flight": 1,
"per_reader_max_in_flight": {
"falcon3-10b-qualification-v7-c8647169c2b9": 1,
"olmo2-13b-qualification-v7-cd836509a1a0": 1
},
"result_order": "deterministic-plan-order",
"calibration_barrier": true,
"automatic_retries": false
},
"transport_faults": {
"total": 0,
"retried": false,
"per_cell": []
},
"transport_truncations": {
"total": 0,
"per_reader_cell": [],
"by_cell": {
"english": 0,
"ainglish": 0
},
"imbalanced_across_cells": false
},
"protocol": "panel.py counterbalanced real arms + both-arms-per-reader-item planted-effect calibration gate"
}