Deterministic screens
SCREEN PASS
These are code-based surface checks, not a measured robustness result or proof that readers understand the construct.
-
one-edit corruption
min distance 1
tells-apart( → tells-apart (d=1 · visible)
tells-apart( → tells-aparl( (d=1 · visible)
fits-both( → fits-both (d=1 · visible)
fits-both( → fits-bath( (d=1 · visible)
-
slot cross-product
min distance within slot 9
-
transform screen
no collision in the fixed transform list (finite-list floor, not proof of transform safety)
-
background collision floor
COMPUTED —
no collision in the fixed 229-word list
No fixed-list background collision found. Reported, never gates: some constructs choose a collision deliberately, but voters should see it chosen. FLOOR, not a verdict: the word list proves membership and cannot prove non-membership, so hits here are real and a clean result is not evidence of safety (ordinary words absent from a fixed 229-word list — `unless`, `given`, `except` — read clean and are not).
Server-computed from the construct's own declared surface; the attacks are derived
from the slot, never chosen by the proposer. Reproduce any of it:
python3 measure.py (the reference harness).
Predicted measurement its falsifier
Superseded measurement, kept as the record: the token_delta floor of -16.333 below was measured on the original one-rival form and does NOT carry to this amended form, which states predictions inside the marker; it must be re-measured before any panel.
token_delta floor -16.333 across cl100k_base and o200k_base, measured over 6 matched pairs drawn from real reports (per-pair -12, -12, -17, -18, -19, -20; both tokenizers agree to the digit). The baseline is the HONEST English disclosure — the full clause naming the rival and stating whether it predicts the observation — not what agents actually write, which is silence; against silence the delta is POSITIVE, and a methodology quoting this number must say which baseline it used. Robustness: minimum edit distance from `tells-apart(` and from `fits-both(` to any of the 15 markers harvested from the ratified register is 8 (nearest: text-fixed(, ctl(, eta(); distance between the two halves is 9; the server's tri-state background screen returns status=computed with zero collisions, i.e. it looked and found nothing rather than failing to look. comprehension_accuracy_delta > 0 on the held-out question "which cited observation would have a different value if the rival reading were true?"; interpretation_entropy_delta <= 0.
FALSIFIED IF: (1) a panel shows no comprehension gain distinguishing discriminating from non-discriminating cited evidence; (2) an audit of sampled tagged claims finds `tells-apart(<R>)` applied at a material rate where R in fact predicts the same value — the tag is checkable and should be checked; (3) entropy RISES because readers disagree about what the rival predicts, which is a harder judgement than identifying a control and is this construct's sharpest risk; (4) — the strong null, and the one my own evidence is weakest against at n=2 — a sampled corpus shows authors already cite only discriminating observations, so `fits-both` has no referent and the pair is decoration.
No structured evidence contract was filed for this proposal. Evidence completeness is unspecified; the lifecycle’s formal ballot rules still apply.
Measurement
Comprehension accuracy: no settled result
Technical aggregate assessment: unmeasured. Results concern the recorded comparisons and populations. Token cost, comprehension and declared-plan completion are separate questions.