Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
start-by → start by (d=1 · visible)
start-by → starts-by (d=1 · visible)
start-by → star-by (d=1 · visible)
complete-by → complete by (d=1 · visible)
complete-by → completes-by (d=1 · visible)
complete-by → compete-by (d=1 · visible)
-
slot cross-product
min distance within slot 7
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
Primary: a preregistered paired comprehension panel compares each marked form with its full careful-English mapping under the same determinate ground truth. Use durative tasks for which start and successful completion are distinct, balanced across uploads, builds, reviews, migrations, payments, physical dispatch, and asynchronous jobs. Cross each task frame with both markers so domain expectations cannot reveal the answer. Keep t as an explicit UTC instant to prevent time-zone or deictic ambiguity from contaminating the phase test.
Present four diagnostic states relative to t: (1) acknowledgement/queueing only; (2) genuine execution started but unfinished; (3) declared success condition satisfied; and (4) execution ended in failure. Ask whether the deadline obligation has been met and which fact—start, successful completion, both, or neither—is required by the instruction. Exact phase-state accuracy is primary. Prediction: marked language is non-inferior to careful English within 5 percentage points for each polarity and has token_delta < 0. Report absolute accuracy, paired delta with interval, each marker separately, and unresolved when the interval cannot exclude the margin.
A third bare arm uses “do X by t.” It descriptively measures which event readers assume and the cannot-tell rate; it is not the confirmatory accuracy denominator. Beating deliberately underspecified prose cannot substitute for matching careful English. Include positive controls with ordinary explicit prose and negative controls where no deadline is present.
Secondary robustness channels: hyphen-to-space, parenthesis loss, single-character edits, and the disclosed `complete-by` → `compete-by` corruption. Hyphen-to-space should be non-degrading. For a lexical corruption, detection is required; silently interpreting an invalid different word as the intended marker is not credited as semantic recovery. Tag-fidelity samples real uses against event evidence: acknowledgements and queue records cannot substantiate `start-by`, and terminal failure cannot substantiate `complete-by`. REFUTED IF either marker is inferior to careful English beyond 5 points, acknowledgement is routinely accepted as a start, failure is routinely accepted as completion, readers treat `start-by` as a completion deadline at material rates, the disclosed corruption passes silently, fidelity falls below the register floor, or observed adoption is zero under the no-adoption sweep.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Comprehension accuracy: no settled result
Technical aggregate assessment: unmeasured. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.