How many tokens does it cost now?
token_delta compares matched wording on the named current tokenizers. It can run locally without a GPU. It says nothing by itself about understanding.
Evidence orientation
A token count asks how current software encodes a form. A comprehension test asks whether a reader recovers the intended meaning. Neither substitutes for the other, and one original result is not yet confirmed evidence.
token_delta compares matched wording on the named current tokenizers. It can run locally without a GPU. It says nothing by itself about understanding.
comprehension_accuracy_delta compares the Ainglish form with its complete standard-English mapping on held-out questions. It needs model or human readers, locally or through remote inference.
robustness_delta measures differential degradation under a declared corruption surface. It must not quietly charge the construct again for a baseline comprehension gap.
learnability, tag_fidelity and background-collision measurements answer narrower questions where they apply. They do not become interchangeable merely because all are numbers.
Name the falsifiable prediction, metric, comparator and stopping rule before seeing the result.
Fix the exact inputs, roster, versions and seed in a content-addressed manifest.
Mint one attempt, run once under the declared method, and publish favourable, null or adverse output honestly.
A different principal tests the same estimand with disjoint inputs. Reusing the original inputs is reproduction, not independent confirmation.
The register counts eligible principal-level voices. Agreement can confirm; disagreement remains visible and may require another independent result.
It establishes the exact claim, inputs and observed value. On its own it is not independently confirmed, however large its panel or attractive its result.
It names the original manifest it addresses, preserves the estimand, uses a disjoint principal and different metric inputs, and files the result even when it disagrees.
The default confirmation threshold is 1 eligible independent replication. A same-input rerun can verify code or arithmetic, but cannot supply the missing independent evidence voice.
A declared metric is complete only when its supporting result survives eligible independent settlement. Every required metric must be complete before the evidence gate is clear.
Eligible replications disagree. The system preserves both outcomes and counts one settlement voice per principal; it does not choose the nicer run or average away the conflict.
A confirmed adverse comprehension, clarity, robustness or applicable fidelity result can veto ratification and move the version to rejection. Token cost alone never vetoes.
A missing replication, ceiling-bound null, invalid instrument or incomplete declared metric does not become evidence of no effect. It remains named work.