Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
fact-not-known → fact not-known (d=1 · visible)
fact-not-known → fact-not known (d=1 · visible)
fact-not-known → facts-not-known (d=1 · visible)
fact-not-known → fact-not-know (d=1 · visible)
choice-not-made → choice not-made (d=1 · visible)
choice-not-made → choice-not made (d=1 · visible)
choice-not-made → choices-not-made (d=1 · visible)
choice-not-made → choice-not-make (d=1 · visible)
-
slot cross-product
min distance within slot 10
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
PRIMARY: a pre-registered paired comprehension panel compares each marked form with its full careful-English mapping under the same ground truth. Use at least 100 paired items per marker (200 total). For every item ask two held-out questions whose vocabulary appears in neither surface: (1) “Does an operative answer already exist independently of a new selection?” and (2) “What can close the gap: retrieving/deriving evidence, an authorized selection, or neither?” Exact joint classification is primary. Prediction: each marked form is non-inferior to careful English within 5 percentage points, each absolute accuracy clears the protocol floor, and token_delta < 0 against the full honest mapping. Report each marker separately, paired delta with 95% interval, discordant-pair count, and the resolution bound; if the interval cannot exclude the margin, report UNRESOLVED.
ITEM DESIGN: cross domains and lexical expectations so topic cannot reveal the answer—software state, payments, schedules, policy, procurement, physical inventory, mathematical results, and release planning each appear under both markers. Required hard cells include: (a) an authorized decision already made but not learned by the speaker (`fact-not-known`); (b) every relevant fact retrieved but authority has not selected (`choice-not-made`); (c) a preference exists but is not operative; (d) a decision exists but is not applied; (e) a future contingency fixed by neither current fact nor authorized choice (neither); (f) human-required and agent-authorized choices; (g) negative and nested issues; and (h) a named criterion whose output exists but has not been computed. Balance answer positions and keep the deciding authority out of the held-out question text.
A third bare arm uses “TBD,” “open,” or “we don't know yet.” It is a descriptive ambiguity arm, not the confirmatory accuracy denominator: report evidence/selection/neither/cannot-tell distributions and forced-guess splits. A perfect reader may correctly answer cannot-tell when bare prose omits the resolution mode; beating that omission cannot replace matching careful English.
ROBUSTNESS: repeat the panel after first-hyphen loss, second-hyphen loss, all-hyphen loss, ordinary single-character edits, and whole-token `not` deletion. Hyphen loss should be non-degrading. `fact-known` and `choice-made` are opposite-state phrases, not recoverable aliases: readers must surface the corruption rather than silently supply the missing negation. Report the token-deletion channel separately from character-edit robustness so its distance does not hide its semantic severity.
TAG FIDELITY: instrumentable cases only. `fact-not-known` is false when no criterion currently fixes an answer or when the declared information available to the speaker already contains it. `choice-not-made` is false when an operative selection already exists, even if the speaker has not retrieved it. Hidden mental state with no auditable trace is UNKNOWN and excluded, never counted as faithful. REFUTED IF either marker is inferior to careful English beyond 5 points, readers systematically treat made-but-unlearned choices as still unmade, readers infer human authority from `choice-not-made`, negation loss passes unnoticed at meaningful rates, fidelity is below 0.5 on auditable cases, or post-ratification observed adoption is zero.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Token cost: lower · Comprehension accuracy: no settled result
Technical aggregate assessment: helps. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.