Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 3
done-under(<C>): → done(<C>): (d=6 · visible)
done-under(<C>): → done-under C: (d=4 · visible)
complete-for(<R>): → complete(<R>): (d=4 · visible)
complete-for(<R>): → complete-for R: (d=4 · visible)
stopped: → stop: (d=3 · visible)
-
slot cross-product
min distance within slot 10
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
PRIMARY: preregister a paired comprehension panel with at least 60 items, each a report of finished work that in bare English is ambiguous between the three claims, comparing four arms: (a) `stopped:`, (b) `done-under(<C>):`, (c) `complete-for(<R>):`, (d) bare "done". For each item ask two held-out questions: (1) which of the three claims is the speaker making — a stop, a scoped correctness claim, or an unqualified handoff? (2) what next action is licensed — none, cautious build, or unqualified action? Exact joint classification is primary. Prediction: arms (a)–(c) are classified correctly substantially more than arm (d), and each marker is non-inferior to its careful-English mapping within 5 percentage points; token_delta < 0 against that mapping. Report arms separately, paired delta and 95% interval.
FALSIFIER (what would refute it): a comprehension panel cannot tell which claim a completion report is making — i.e. readers of `stopped:` treat it as a handoff at the same rate as readers of bare "done". If `stopped:` fails to suppress the handoff over-read that bare "done" produces, that half is refuted even if the other two succeed. Secondary: if readers cannot distinguish `done-under(<C>):` from `complete-for(<R>):` (the scoped claim from the unqualified one), the pair fails its distinctiveness test.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Token cost: lower · Comprehension accuracy: no settled result
Technical aggregate assessment: helps. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.