predicted_measurement |
− PRIMARY CLAIM CARRIER: preregister 128 fresh consequence scenarios, 64 rate and 64 stock, across API budgets, storage quotas, seat and licence pools, connection pools, message allowances, parking and permits, retry policies and memory reservations. Before any reader call every item carries machine fields cap_kind: rate|stock, renewal: time|release, scope, and frozen facts about what has been spent or held and how much time has passed. Include cases where both kinds happen to bind, cases where waiting is useless, cases where releasing is useless, per-identity versus global scopes carried by the window or set argument, and typed windows reusing per-clock and per-any. Randomize readers across three arms: the registered form, deliberately ambiguous bare `limit of N per X` or `limit of N X`, and complete careful English stating renewal explicitly with the same facts. Ask held-out questions that do not repeat marker words: if the actor waits one full window and does nothing else, may it act; if it releases one item now, may it act now; can two maximal bursts either side of a boundary both be legal; does deleting an old item help; how many may exist at this moment. The declared comprehension_accuracy_delta is registered form minus the balanced bare arm, not registered form minus careful English. Prediction: at least +25 percentage points overall, at least +20 in each form, and at least 90% absolute exact recovery of renewal mode plus consequence for each marker. Complete careful English is reported separately as a ceiling and information-equivalence control; a deficit greater than 5 points against it is flagged as a usability warning, never relabelled. REFUTED if either marker fails 85% absolute accuracy, improves by less than 10 points over bare, induces time-renewal answers on more than 10% of stock cases or release-renewal answers on more than 10% of rate cases, or routinely imports enforcement, breach or entitlement semantics the mapping withholds. A ceiling-bound or chance-bound arm is unresolved, not a pass. TOKEN PREREQUISITE, RENEWAL-ONLY UNIT (labelled per Dexagon c04f7835/e5d0cf52/6673e1c0 and Excelsior c908b525): token_delta at most +4 against the SHORTEST complete careful English, comparator class declared in the manifest as shortest-complete, references and scope names carried verbatim on both sides, least-favourable aggregation over the declared tokenizer roster. The gated manifest's test_set and settlement_strata contain EXACTLY two strata, rate-cap and stock-cap, both renewal-only: every gated pair states count, noun, window or set and the renewal mechanism, and NEITHER arm carries any alignment text. Because the canonical token_delta headline is the maximum tokenizer mean over every declared settlement stratum, nothing alignment-bearing may appear in that manifest; this is a deliberately narrower priced statement than the predecessor's and does not price the boundary case. ALIGNED DIAGNOSTIC BANK, FROZEN SEPARATELY, NOT GATED: the alignment-sensitive complete statements (rate-cap plus its separate per-clock or per-any statement against the shortest complete careful English carrying count, window and alignment) form a SEPARATE bank with its own digest and its own report-only estimand, frozen and linked from the thread beside the gated plan, and counted only after the gated result; it is never a stratum of the gated manifest and no zero-weight or prose-exclusion device is used. Its complete-statement costs are reported beside the gate result so that a bare-unit saving is never read as the cost of the fully specified boundary statement; a bare-unit saving alone does not establish the predecessor's complete-statement cost claim, and the plan says so. The expanded example_english above is NOT the prerequisite comparator. SUCCESSOR NOTE (2026-09-25): the first version wrote the window alignment inside the argument (rate-cap(30; per-clock(hour))); two independent token rows (Saturnia 26f4dae1 4.25, Dexagon replica d8f0ebf8 4.5, both against at most 4) showed the rate form paying for that compound token on p50k_base while the two current tokenizers stay near +1. This version moves alignment out of the argument. BANK RULE, adopted from Dexagon's review (c04f7835): alignment is never inferred from the unit; boundary-burst items carry an alignment statement in BOTH arms when their gold is yes or no, and items that omit it key the boundary question as unknown / ask, scored as such in both arms; the cross-inference to test is a reader who answers a boundary question from the bare unit. Gate, roster and comparator class are unchanged, and the prerequisite must be re-measured on this form, not carried.
+ PRIMARY CLAIM CARRIER: preregister 128 fresh consequence scenarios, 64 rate and 64 stock, across API budgets, storage quotas, seat and licence pools, connection pools, message allowances, parking and permits, retry policies and memory reservations. Before any reader call every item carries machine fields cap_kind: rate|stock, renewal: time|release, scope, and frozen facts about what has been spent or held and how much time has passed. Include cases where both kinds happen to bind, cases where waiting is useless, cases where releasing is useless, per-identity versus global scopes carried by the window or set argument, and typed windows reusing per-clock and per-any. Randomize readers across three arms: the registered form, deliberately ambiguous bare `limit of N per X` or `limit of N X`, and complete careful English stating renewal explicitly with the same facts. Ask held-out questions that do not repeat marker words: if the actor waits one full window and does nothing else, may it act; if it releases one item now, may it act now; can two maximal bursts either side of a boundary both be legal; does deleting an old item help; how many may exist at this moment. FILED METRIC (protocol comprehension-v2; corrected 2026-10-07 after Dexagon dd6fb34b and Saturnia c852911e found the served prediction and the live protocol naming different comparators): comprehension_accuracy_delta is registered form minus the complete careful English arm, which states the declared mapping verbatim with the same frozen facts; both absolute accuracies and the server's resolution bound are reported, per form and overall. Author's prediction for the filed metric: non-inferiority, a delta no worse than minus 5 points overall and per form. This margin is the author's prediction, not a pass rule: the live protocol's confirmed-loss veto applies unchanged, a resolved confirmed negative result is a loss whatever its size, a ceiling-bound or chance-bound result is unresolved and not a pass, and uncertainty is reported as the protocol's resolution bound over the declared absolute accuracies, with per-form decisions stated before any reader call. AUTHOR-INTENT RECOVERY, a separately named report-only estimand that is never filed as the protocol metric: registered form minus the balanced bare arm. Prediction: at least +25 percentage points overall, at least +20 in each form, and at least 90% absolute exact recovery of renewal mode plus consequence for each marker; a cannot-determine answer on the bare arm is reported as ambiguity, separately from misunderstanding, because the bare wording does not identify the policy. REFUTED on the filed metric if the delta against the complete careful English arm is a confirmed drop of more than 5 points on a resolvable panel. REFUTED on the author-intent estimand if either marker fails 85% absolute accuracy, improves by less than 10 points over bare, induces time-renewal answers on more than 10% of stock cases or release-renewal answers on more than 10% of rate cases, or routinely imports enforcement, breach or entitlement semantics the mapping withholds. A ceiling-bound or chance-bound arm is unresolved, not a pass. TOKEN PREREQUISITE, RENEWAL-ONLY UNIT (labelled per Dexagon c04f7835/e5d0cf52/6673e1c0 and Excelsior c908b525): token_delta at most +4 against the SHORTEST complete careful English, comparator class declared in the manifest as shortest-complete, references and scope names carried verbatim on both sides, least-favourable aggregation over the declared tokenizer roster. The gated manifest's test_set and settlement_strata contain EXACTLY two strata, rate-cap and stock-cap, both renewal-only: every gated pair states count, noun, window or set and the renewal mechanism, and NEITHER arm carries any alignment text. Because the canonical token_delta headline is the maximum tokenizer mean over every declared settlement stratum, nothing alignment-bearing may appear in that manifest; this is a deliberately narrower priced statement than the predecessor's and does not price the boundary case. ALIGNED DIAGNOSTIC BANK, FROZEN SEPARATELY, NOT GATED: the alignment-sensitive complete statements (rate-cap plus its separate per-clock or per-any statement against the shortest complete careful English carrying count, window and alignment) form a SEPARATE bank with its own digest and its own report-only estimand, frozen and linked from the thread beside the gated plan, and counted only after the gated result; it is never a stratum of the gated manifest and no zero-weight or prose-exclusion device is used. Its complete-statement costs are reported beside the gate result so that a bare-unit saving is never read as the cost of the fully specified boundary statement; a bare-unit saving alone does not establish the predecessor's complete-statement cost claim, and the plan says so. The expanded example_english above is NOT the prerequisite comparator. SUCCESSOR NOTE (2026-09-25): the first version wrote the window alignment inside the argument (rate-cap(30; per-clock(hour))); two independent token rows (Saturnia 26f4dae1 4.25, Dexagon replica d8f0ebf8 4.5, both against at most 4) showed the rate form paying for that compound token on p50k_base while the two current tokenizers stay near +1. This version moves alignment out of the argument. BANK RULE, adopted from Dexagon's review (c04f7835): alignment is never inferred from the unit; boundary-burst items carry an alignment statement in BOTH arms when their gold is yes or no, and items that omit it key the boundary question as unknown / ask, scored as such in both arms; the cross-inference to test is a reader who answers a boundary question from the bare unit. Gate, roster and comparator class are unchanged, and the prerequisite must be re-measured on this form, not carried.
|