Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
some-or-all → some or all (d=2 · visible)
some-but-not-all → some but not all (d=3 · visible)
some-or-all → some-nor-all (d=1 · visible)
some-but-not-all → some-but-all (d=4 · visible)
-
slot cross-product
min distance within slot 6
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
PRIMARY: preregister a paired comprehension panel with at least 100 meaning-matched items per form. Cross incident tests, replicas, permissions, recipients, alerts, inventory, and ordinary human situations. Every action frame appears with both meanings so topic and consequence cannot reveal the key. Compare each marked form with its full careful-English mapping and with bare ‘some’; bare ‘some’ is a descriptive ambiguity arm, not the easy confirmatory denominator.
Use two held-out consequence questions per item whose wording does not repeat ‘or all’ or ‘but not all’: (1) ‘Must at least one member of the set fail to satisfy the predicate?’ and (2) ‘Would the sentence be contradicted if every member satisfied the predicate?’ For some-or-all the keyed answers are no/no. For some-but-not-all they are yes/yes. Exact joint recovery is primary. Report absolute accuracy and paired delta with intervals for each form separately; never pool the two forms so an easy arm can hide a failing one.
Prediction: each marker is non-inferior to its careful-English mapping within a preregistered 5-percentage-point margin, materially more accurate than bare ‘some’ on the all-case question, and has token_delta <= 0 against its meaning-matched expansion across both maintained tokenizer lineages. Token delta versus bare ‘some’ is honestly positive: precision costs surface.
OVER-READING CONTROLS: ask whether some-or-all claims the writer has not counted (it does not), whether some-but-not-all identifies which members satisfy the predicate (it does not), whether either gives an exact count (it does not), and whether the marker itself fixes the population boundary (it does not). Include invalid controls with an unbounded set or a set known to contain fewer than two members; the correct response is invalid or unresolved, not an invented complement.
ROBUSTNESS: repeat matched cells after hyphen-to-space conversion, punctuation stripping, ordinary single-character edits, and whole-token ‘not’ deletion. Hyphen loss should preserve direction. ‘some-but-all’ must be surfaced as malformed rather than silently repaired or executed. Test the nearest live-register forms returned by preflight as named distractors, not only self-chosen corruptions.
TAG FIDELITY on auditable cases: some-but-not-all is false when zero or every member satisfies the predicate; some-or-all is false when zero members satisfy it. Underinformativeness is not falsity: a true some-or-all in an all-members world remains truth-conditionally faithful, while the separate over-reading panel measures whether readers mistake it for ignorance.
REFUTED IF either form is inferior to careful English by more than 5 points; marked readers recover the all-case no better than bare-some readers; the two markers collapse into the same interpretation; some-or-all is systematically read as an assertion of speaker ignorance; the complement or exact-count over-readings persist materially; whole-token negation loss passes silently at a material rate; fidelity falls below the register floor; a simpler existing form dominates both clarity and length; or observed adoption is zero under the no-adoption sweep.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Comprehension accuracy: no settled result
Technical aggregate assessment: unmeasured. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.