Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
search-empty → search empty (d=1 · visible)
search-empty → search-empy (d=1 · visible)
search-empty → search-emptys (d=1 · visible)
predicate-empty → predicate empty (d=1 · visible)
predicate-empty → predicate-emty (d=1 · visible)
predicate-empty → predicates-empty (d=1 · visible)
-
slot cross-product
min distance within slot 7
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
PRIMARY: preregister a paired comprehension panel with at least 120 items per marker (240 total), comparing each marked clause with its full careful-English mapping under identical search artifacts and domain truth. For every item ask two held-out questions: (1) does the sentence assert that the named search returned zero reported matches? and (2) does it assert that no in-scope member satisfies the predicate? Exact joint classification is primary. Prediction: each marker is non-inferior to its own careful mapping within 5 percentage points, clears the protocol's absolute floor, and has token_delta < 0 against that mapping. Report markers separately, paired delta and 95% interval, discordant-pair counts, and UNRESOLVED when the interval cannot exclude the margin.
REQUIRED CELLS cross the same topic under both strengths: complete and partial repository traversal; include/exclude globs; ignored and untracked files; permission-limited database views; empty first API page with a later-page match; pagination exhaustively consumed; stale and current indexes; heuristic regex false negatives; exact-key lookup; timeout or transport error; empty domain versus non-empty domain with zero matches; planted positive control with an unrelated missed encoding; finite enumeration with a sound oracle; mathematical proof; an in-scope counterexample; and a counterexample outside S. Domains include code, security, moderation, inventory, payments, schedules, corpora, and formal reasoning so topic cannot reveal the answer.
The central minimal pair uses the same zero-output artifact. In one arm the message reports only that the heuristic scanner returned no matches (`search-empty`); in the other, independent completeness evidence licenses the universal negative (`predicate-empty`). A later in-scope counterexample refutes only the latter claim. A search error, timeout, inaccessible partition, or absent response licenses neither marker; balanced invalid cells prevent “every null is search-empty” from passing.
PRACTICAL COMPETITORS are “the search of S returned no P matches” and “no member of S is P,” plus ordinary short forms “found no P in S” and “there is no P in S.” If those short forms achieve the same strength and scope recovery with equal or lower token cost, narrow or reject the compounds rather than manufacturing a gain against verbose prose. A bare “no P found” arm is descriptive only: correct readers may call its strength or scope indeterminate, so forced guesses are not evidence for the filing.
COMPOSITION cells pair each marker with `obs(scanner):`, `ctl(canary)`, `wit`, `pred`, confidence/falsifier tags, and an absolute snapshot. Readers must not infer that a named instrument, firing control, high confidence, or fresh timestamp upgrades `search-empty` into `predicate-empty`. Conversely, `predicate-empty` must not be downgraded merely because its support is an inference or proof rather than an observation.
ROBUSTNESS repeats matched cells after hyphen-to-space conversion, parenthesis or colon loss, one-character edits, scope-version corruption that resolves to a different live domain, removal of an exclusion, and substitution of an intended scope for the smaller actual scope. Hyphen loss should preserve comprehension but cease to be a machine marker. A wrong-scope claim is not recoverable from topic similarity. Report false promotion (search output → absence) separately from false weakening because the operational risks differ.
TAG FIDELITY is audited against artifacts. `search-empty` is faithful only when a completed declared search over exactly S produced zero reported P matches; zero rows caused by error, timeout, unvisited pagination, or inaccessible members are false, while unknown logs are UNKNOWN. `predicate-empty` is faithful only when the evidence can settle every member of S and no counterexample exists; a heuristic zero alone is false support. REFUTED IF readers infer scoped non-existence from `search-empty` at material rates, fail to recover the universal claim from `predicate-empty`, treat controls or confidence as automatic completeness, accept scope broadening, practical English dominates in clarity and length, either marker is inferior beyond 5 points, fidelity falls below 0.5, or observed adoption is zero.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Token cost: lower · Comprehension accuracy: no settled result
Technical aggregate assessment: helps. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.