Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
true-as-worded → true as worded (d=2 · visible)
true-as-worded → trues-as-worded (d=1 · visible)
true-as-worded → true-is-worded (d=1 · visible)
false-as-worded → false as worded (d=2 · visible)
false-as-worded → falses-as-worded (d=1 · visible)
false-as-worded → false-is-worded (d=1 · visible)
-
slot cross-product
min distance within slot 4
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
Primary: a preregistered paired comprehension panel compares each marker with its full careful-English mapping under identical question and world-state ground truth. Balance positive questions, contracted negative questions, uncontracted `not`, lexical negatives (`fail`, `lack`, `reject`), scoped quantifiers, and two negations. Every question frame appears with both truth states and both markers, so desirability or lexical polarity cannot reveal the answer. Exclude tag, alternative, bundled, and internally ambiguous questions in the confirmatory set because the construct declares them out of scope.
Ask a held-out real-world consequence rather than “was the answer true?” For “Didn't node A reject build 7?” followed by a marker, ask whether node A accepted or rejected build 7. Exact denotation accuracy is primary. Prediction: marked answers are non-inferior to the full mapping within 5 percentage points for each marker and negation stratum, with token_delta < 0 against that mapping. Report absolute accuracy, paired delta and interval, polarity-specific cells, and UNRESOLVED when the interval cannot exclude the margin.
Two secondary comparators keep the claim honest. Bare yes/no is a descriptive ambiguity arm: report interpretation entropy and cross-model/dialect splits, but do not use it as the confirmatory accuracy denominator. A full declarative echo answer (“The backup did not finish”) is the practical competitor. Stratify questions by proposition length and compare tokens and comprehension. REFUTE OR NARROW the construct if echo answers dominate it in both clarity and length across representative exchanges; do not cherry-pick only long propositions to manufacture compression.
Robustness channels include hyphen-to-space, punctuation loss, one-character edits, and distractors that ask about wording quality. Hyphen-to-space should be non-degrading. The pair is distance 4, so no single edit reaches the opposite marker. Tag-fidelity audits whether a use has exactly one salient determinate P and whether any accompanying restatement/evidence agrees with the selected truth value. REFUTED IF either pole is inferior to careful English beyond 5 points, readers reverse negative questions at material rates, “as-worded” is routinely read as grammaticality rather than truth, scoped-negation cells fail, the explicit echo baseline dominates, fidelity falls below the register floor, or observed adoption is zero under the no-adoption sweep.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Token cost: mixed results · Comprehension accuracy: no settled result
Technical aggregate assessment: helps. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.