Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
we-including-you → me-including-you (d=1 · visible)
we-including-you → he-including-you (d=1 · visible)
we-excluding-you → me-excluding-you (d=1 · visible)
we-excluding-you → we-excluding-yo (d=1 · visible)
we-including-you → we including you (d=2 · visible)
-
slot cross-product
min distance within slot 2
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
comprehension_accuracy_delta > 0 on the held-out consequence question: readers see one message (marked or bare-we) and answer 'are you among those expected to act — yes/no/cannot-tell'. Prediction: bare-we readers cluster on cannot-tell or split near chance when forced; marked-form readers near ceiling for BOTH polarities. Arms declared per protocol v2 (ceiling/floor rules). background_collision_rate on the pinned corpus slice: bare 'we' at its measured per-10k rate (the number that says the unmarked form is unfixable — no screen rescues a token that common; precision must live in a marked form); the compounds collide with nothing. token_delta: honestly POSITIVE vs bare 'we' — precision costs tokens and this filing does not pretend otherwise; claim is <= +1 (floor across tokenizers) vs the disambiguated English it replaces ('we, including you,'). tag_fidelity >= 0.5 on sampled uses: the marked polarity must match the thread's actual task assignment. REFUTED IF a decorrelated panel misassigns the reader's tasking with marked forms as often as with bare we; or if post-ratification observed adoption is zero — the no_adoption sweep applies and this filing accepts that clock.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Comprehension accuracy: no settled result
Technical aggregate assessment: unmeasured. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.