Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
include-both → includes-both (d=1 · visible)
include-start-only → includes-start-only (d=1 · visible)
include-end-only → includes-end-only (d=1 · visible)
exclude-both → excludes-both (d=1 · visible)
include-both → include both (d=1 · visible)
include-start-only → include start only (d=2 · visible)
include-end-only → include end only (d=2 · visible)
exclude-both → exclude both (d=1 · visible)
include-both → exclude-both (d=2 · silent)
-
slot cross-product
min distance within slot 2
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
Primary: a preregistered comprehension panel balanced across the four endpoint states and across numeric ascending, numeric descending, dates, timestamps, identifiers, alphabetic spans, and pagination. Each lexical frame appears with all four states so domain convention cannot reveal the answer. For every instruction ask two independently scored questions: “Would a value exactly equal to the first written endpoint be selected?” and the same for the second, with yes/no/cannot-tell.
Compare (1) the Ainglish qualifier, (2) its full careful-English mapping, and (3) a bare-range descriptive arm. The confirmatory claim is non-inferiority of the marked arm to careful English within 5 percentage points on exact two-bit accuracy, with token_delta < 0; report each marker and direction stratum separately. The bare arm measures residual ambiguity and forced endpoint assumptions but is not allowed to make an easy “better than ambiguity” result stand in for the careful-English comparison. Do not use mathematical interval brackets as the English control; those are a competing notation, not the declared mapping.
Secondary: measure robustness after hyphen loss, single-character insertions/deletions, and the specifically disclosed two-substitution `include-both` → `exclude-both` channel. Hyphen loss should be non-degrading because it yields the careful instruction. For corrupted valid markers, score both detection and semantic recovery; silently interpreting the opposite as intended is a failure. A tag-fidelity audit compares the marked range with the set actually selected, including values exactly equal to A and B. REFUTED IF the marked arm is more than 5 points worse than careful English, readers systematically treat written “start” as the numeric lower bound in descending cases, the disclosed polarity corruption passes silently at a material rate, fidelity falls below the register floor, or observed adoption remains zero under the no-adoption sweep.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Token cost: lower · Comprehension accuracy: no settled result
Technical aggregate assessment: helps. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.