Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
all-or-nothing → all or nothing (d=2 · visible)
all-or-nothing → all-for-nothing (d=1 · visible)
all-or-nothing → all-or-nothings (d=1 · visible)
keep-successes → keep successes (d=1 · visible)
keep-successes → kept-successes (d=2 · visible)
-
slot cross-product
min distance within slot 14
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
PRIMARY: preregister a paired agent-comprehension panel comparing each marked form with its complete careful-English mapping under the same bounded action set, per-member outcomes, and effect model. Use at least 100 paired items per form and report the forms separately. Cross permissions, file operations, data migration, publication, notification, archival, indexing, and reversible external actions. Every scenario template appears with both policies, and success/failure positions are balanced so domain, order, or which member fails cannot reveal the answer.
For each item ask held-out operational questions using short opaque answer labels whose maximum lengths are exercised by equal-length calibration: (1) after one required member fails, which successful member effects remain authoritative at terminal handoff; (2) must a prior successful member be withheld or reversed solely because its sibling failed; and (3) is the set's terminal state full success, partial result, or failed-with-no-retained-effects? Exact joint recovery is primary. Prediction: each marked form is non-inferior to its full careful-English mapping within 5 percentage points, clears the register's absolute accuracy floor, and has token_delta < 0 against that complete mapping. Report paired delta and interval, absolute accuracy, discordant cells, each form, domain, reversibility, failure position, and reader separately.
COMPARATORS AND OVER-READING: bare unqualified batch language is a descriptive ambiguity arm, never the confirmatory denominator. Include “perform no changes unless every member succeeds,” “roll back every successful member if any member fails,” “keep each successful result even if another member fails,” `atomic`, “best effort,” and “partial success allowed” as practical competitors. Narrow or reject the pair if a competitor carries the same boundary more clearly and reliably at equal or lower cost. Ask separate questions showing that the marker does not determine sequential versus parallel execution, stop-on-first-failure versus attempt-all, retry safety, delegation, or whether an individual member met its own success criterion. Include a known-positive trap that should elicit each named over-read; an all-negative instrument is undiagnostic.
REQUIRED HARD CELLS: include failure before any effect, failure after one staged success, failure after one committed but reversibly compensable success, an irreversible member that makes `all-or-nothing` invalid, remaining members not attempted after a catastrophic stop, nested action sets with different inner and outer policies, a successful action later invalidated for an independent reason, and partial progress that is not yet a successful member effect. Correct readers must distinguish an impossible policy from permission to improvise a partial result.
ROBUSTNESS AND FIDELITY: repeat matched cells after hyphen-to-space conversion, punctuation loss, ordinary single-character edits, and especially `all-for-nothing`. Hyphen loss should preserve direction; the one-insertion idiom must be rejected as an invalid qualifier. For fidelity, use auditable per-member status and effect logs plus a declared terminal handoff. `all-or-nothing` is false if any successful sibling remains authoritative after a required failure, or if an executor knowingly starts an irreversible set without a no-partial guarantee. `keep-successes` is false if a valid success is reversed solely because a sibling failed, or if failure disclosure is suppressed. Hidden or unauditable effects are UNKNOWN, not faithful.
REFUTED IF either form is inferior to careful English beyond 5 points; readers confuse “all-or-nothing” with a prediction that all will succeed; `keep-successes` is read as ignore-errors or mandatory continue-on-error; either form leaks into execution order, retry, delegation, or action-count judgments at material rates; impossible atomicity is silently promised; `all-for-nothing` is accepted as a policy; fidelity falls below the register floor; a practical competitor dominates in clarity and length; or an eligible post-ratification scan finds no adoption.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Comprehension accuracy: no settled result
Technical aggregate assessment: unmeasured. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.