Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
supersedes( → supercedes( (d=1 · visible)
supersedes( → supersede( (d=1 · visible)
supersedes( → superseded( (d=1 · visible)
supplements( → supplement( (d=1 · visible)
supplements( → supplement's( (d=1 · visible)
-
slot cross-product
min distance within slot 6
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
PRIMARY: build a pre-registered paired instruction-state panel with at least 120 items per relation (240 total). Each item contains two or more immutable clause IDs, their action-bearing contents and issuer identities, a marked follow-up, and an otherwise identical full careful-English expansion. Ask the held-out reader to return (1) the exact set of clauses active after the update, (2) the exact set newly inactive, (3) whether any realised effect must be undone or repeated, (4) whether a conflict or invalid reference must be surfaced, and (5) the resulting action set. Exact joint state is primary; per-field scores diagnose the failure.
Prediction: each marker is non-inferior to its full careful-English expansion within 5 percentage points, clears the protocol's absolute floor, and has token_delta < 0 against that expansion. A decorrelated bare-English arm uses ordinary “actually,” “instead,” “also,” adjacency, and unmarked follow-ups. On items where bare English admits both accumulation and replacement, the marked arm predicts at least a 10-point exact-state improvement. Bare ambiguity is reported rather than forced into a single gold answer where the author supplied none.
REQUIRED STATE CELLS: (a) simple one-clause replacement and addition; (b) several active clauses with only one referenced; (c) explicit multi-reference updates; (d) partial prior execution, proving no implicit rollback or repetition; (e) B supplements A, then C supersedes only A; (f) A superseded by B, then B superseded by C; (g) a contradictory supplement; (h) stale, missing, ambiguous, self, cyclic, and mixed-validity reference lists; (i) a different speaker without update authority; (j) authored order different from delivery order; and (k) a duplicated/retried follow-up whose stable ID must not create a second state transition. Score all-or-nothing reference validity separately from semantic recovery.
COMPOSITION CELLS: place `req:`, `will:`, `start-by/complete-by`, `no-delegation`, `given_c/except_l`, and `in-parallel/in-sequence` inside X. Include the relation string inside `force-suspended`, where it must remain inert. Require clause-level references when only one member of a grouped instruction is replaced; whole-message guessing is an error. A factual correction and a fired falsifier are negative controls: readers must not use these action-lifecycle markers as truth-status operators.
PRACTICAL COMPETITORS: compare `supersedes(id)` with “ignore instruction id and use this instead; completed effects remain,” and `supplements(id)` with “keep instruction id active and also do this; neither overrides the other.” Also test the shorter “replace id” and “also.” If a practical competitor reaches the same exact state more reliably at lower token cost, narrow or reject the filed surface rather than claiming value against only a verbose expansion.
ROBUSTNESS: repeat matched cells after colon loss, parenthesis loss, ordinary single-character marker edits, reference transposition, one-character reference corruption, delayed delivery, duplicated delivery, and summarisation that preserves IDs but changes adjacency. Colon/parenthesis loss and malformed marker spellings are invalid, not recovery aliases. A corrupted reference that resolves to a different active clause is the dangerous wrong-target class and must be reported separately from an unresolved reference. Marker robustness cannot rescue an unauthenticated or transport-corrupted identifier.
FIDELITY: sample auditable uses against message IDs, issuer authority, receipt order, task traces, and realised effects. A `supersedes` use is false if any named active clause remains treated as obligatory after receipt, if an unnamed clause is retired, or if a completed effect is claimed undone without an explicit compensating action. A `supplements` use is false if a named clause is silently displaced or a conflict is silently resolved by recency. Hidden state is UNKNOWN, never faithful by assumption.
REFUTED IF either marker is inferior to careful English beyond 5 points; the marked arm fails to improve exact active-set recovery over ambiguous bare follow-ups; readers routinely infer rollback, partial-reference application, conversation-wide scope, or last-write-wins; contradictory supplements are silently resolved; unauthorised or wrong-target updates are accepted at material rates; a practical competitor dominates in clarity and length; fidelity falls below 0.5; or observed adoption is zero.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Comprehension accuracy: no settled result
Technical aggregate assessment: unmeasured. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.