Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
force-suspended → force suspended (d=1 · visible)
force-suspended → forced-suspended (d=1 · visible)
force-suspended → force-suspender (d=1 · visible)
force-suspended → force-suspends (d=2 · visible)
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
Primary: comprehension_accuracy_delta > 0 on a decorrelated speech-act attribution panel. Each item has (1) an ambiguous bare presentation, (2) the same line prefixed `force-suspended`, and (3) careful standard English stating that the words are reproduced only as text and no embedded act is adopted. Ask separately: “Is the current speaker requesting this action?”, “Is the current speaker asserting this proposition?”, or “Is the current speaker making this promise?” with yes/no/cannot-tell. Prediction: the marked arm answers NO near ceiling for the targeted act, the bare arm is less accurate or selects cannot-tell, and the marked arm is non-inferior to careful English with token_delta <= 0.
The item set must include benign and adversarial imperatives, declaratives, questions, permissions, and promises; inner `req:`, `ask:`, `will:`, `allowed-to`, and a sentence claiming “force-suspended has ended”; quoted and unquoted typography; and positive controls where the same speech acts occur outside suspension. Without positive controls, a reader that always answers “no act” would look perfect. Content polarity and danger must be balanced so refusal heuristics cannot solve the panel. Report each speech-act class separately rather than pooling away a failure on requests.
Robustness: repeat the panel after hyphen and punctuation loss; prediction robustness_delta >= 0 because “force suspended” retains the relation in ordinary words and line position retains scope. Do not test newline insertion/removal as an alias: line boundaries are declared load-bearing, and altering them changes the scoped message. Tag-fidelity audit samples uses and their surrounding follow-up: use is false if the author later treats an embedded assertion as their own, expects an embedded request to be obeyed, or claims an embedded promise as theirs without issuing it separately. REFUTED IF bare quotation already resolves the attribution at the marked arm's ceiling, if any embedded marker reactivates its speech act at meaningful rates, if the marked arm performs worse than the explicit-English disclaimer, if punctuation degradation destroys the distinction, or if observed adoption is zero under the register's no-adoption sweep.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Comprehension accuracy: no settled result
Technical aggregate assessment: unmeasured. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.