Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
no-delegation → no delegation (d=1 · visible)
no-delegation → non-delegation (d=1 · visible)
no-delegation → no-delegations (d=1 · visible)
one-hop-delegation-allowed → one hop delegation allowed (d=3 · visible)
one-hop-delegation-allowed → none-hop-delegation-allowed (d=1 · visible)
one-hop-delegation-allowed → one-hop-delegations-allowed (d=1 · visible)
-
slot cross-product
min distance within slot 13
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
PRIMARY: a pre-registered paired comprehension panel compares each marked qualifier with its full careful-English mapping under the same task, actors, authority, and external-policy ground truth. Use at least 100 paired items per qualifier (200 total), balanced across software changes, private-data review, research, payments, physical work, moderation, and publication. Cross each task frame with both qualifiers so topic sensitivity cannot reveal the delegation policy.
For every item ask three held-out questions: (1) may the responsible principal assign a completion-bearing subtask to an immediate delegate? (2) if an immediate delegate is used, may that delegate pass the subtask to a further principal? and (3) which principal still owes the issuer the completed result? Exact joint classification is primary. Prediction: each marked qualifier is non-inferior to its full mapping within 5 percentage points, clears the protocol's absolute floor, and has token_delta < 0 against that mapping. Report each qualifier separately, paired delta and 95% interval, discordant-pair count, and the v2 resolution bound; an interval that cannot exclude the margin is UNRESOLVED.
REQUIRED HARD CELLS: (a) multiple sibling delegates, so “one hop” is not misread as “one delegate”; (b) an immediate delegate attempting a second hop; (c) a named plural level-zero actor set; (d) deterministic tools versus independently deciding principals; (e) advice or reported evidence versus an assigned completion-bearing subtask; (f) delegation of an unprivileged subtask when the final step requires the original principal's authority; (g) a permitted delegate that lacks the required capability; and (h) composition with `req:`, `will:`, `allowed-to`, `each-alone/as-one`, and `in-parallel/in-sequence`. Predeclare the identity/policy rule that classifies instruments and principals; do not let panel scorers choose it after seeing answers.
A bare action arm—“please do X” or “I will do X”—is descriptive only. Correct readers may answer that delegation is unspecified, so it is not an easy accuracy denominator. Add two practical-English competitors: “do it yourself” and “you may use subagents.” The first may over-prohibit tools; the second may fail to bound recursive delegation or accountability. If either competitor matches the filed semantics in comprehension while being reliably shorter, narrow or reject the construct rather than manufacturing compression against only a verbose paraphrase.
ROBUSTNESS: repeat the panel after hyphen-to-space conversion, ordinary single-character edits, whole-word `no` deletion, `dis` insertion before `allowed`, and the declared d=1 `none-hop` corruption. Hyphen-to-space should be non-degrading. The polarity attacks are not recoverable aliases: readers must surface the corruption rather than silently infer the safer policy. Report permission expansion and over-restriction separately; pooling them would hide the dangerous direction.
TAG FIDELITY: score only auditable cases with task-assignment traces and a predeclared principal/instrument boundary. `no-delegation` is false if another principal performs a completion-bearing subtask. `one-hop-delegation-allowed` is misused if a second-hop assignment occurs, if original constraints are broadened, or if the original responsible principal represents accountability as transferred. Hidden handoffs are UNKNOWN, not faithful. REFUTED IF either qualifier is inferior to careful English beyond 5 points, readers confuse hop depth with delegate count, infer that first-hop delegates may redelegate, treat the permission as credential-sharing authority, interpret `no-delegation` as banning ordinary tools at material rates, a practical competitor dominates in clarity and length, dangerous polarity corruption passes unnoticed, fidelity is below 0.5, or observed adoption is zero.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Token cost: lower · Comprehension accuracy: no settled result
Technical aggregate assessment: helps. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.