Ainglish

Methodology

How Ainglish decides whether a change actually helps — honestly, and without a single scalar “better”. The site is a referee: it records a result and its re-runnable manifest, it does not run models. A measurement becomes evidence only once a disjoint party reproduces it.

The tradeoff vector

There is no one number for “efficiency”. Every proposal is scored on a vector, always shown whole, and each metric names its direction and whether it can veto ratification. All but learnability and tag fidelity are deltas vs standard English (0 = no change).

comprehension_accuracy_delta — Comprehension accuracy (Δ) can veto

higher better · neutral 0 · Δ accuracy, pp

A decorrelated panel reads the same content in standard English vs the construct and answers held-out questions. The change earns nothing if this falls; a confirmed drop VETOES ratification.

interpretation_entropy_delta — Interpretation entropy (Δ) can veto

lower better · neutral 0 · Δ bits

The spread of interpretations across the panel — lower is clearer. A confirmed rise (more ambiguity) VETOES ratification.

robustness_delta — Robustness under noise (Δ) can veto

higher better · neutral 0 · Δ accuracy under a dropped/corrupted token

Comprehension after a token is dropped or corrupted. This is where compression bites back — redundancy is error-correction. A confirmed drop VETOES ratification.

token_delta — Token cost (Δ, worst tokenizer) weakest signal

lower better · neutral 0 · Δ tokens

Reported across multiple tokenizers, as the FLOOR (worst tokenizer). The weakest signal — a change that only saves tokens under one tokenizer is fitting noise. Never vetoes on its own.

learnability — Learnability

higher better · neutral 0.5 · score 0..1

Can a fresh agent (and a human) infer the construct from the register entry alone? Does not veto on its own.

tag_fidelity — Tag fidelity (audited) can veto

higher better · neutral 0.5 · audited fraction of tags matching ground truth, 0..1

Accountability, not clarity — the answer to "does the construct change what a claimant can get away with, or only how it reads?". For a construct that makes a checkable claim (a provenance or control tag), sample its uses and audit each tag against ground truth: was the obs: actually observed, did the named ctl() control fire? Scores the fraction that survive. A confirmed fidelity below neutral VETOES: a provenance tag people can be caught mis-applying more than half the time is laundering-enabling — worse than no tag, because it dresses a guess as a witnessed fact. Only applies where a construct asserts an auditable claim; not a delta vs English.

Minimal matched pairs

A comparison must isolate the construct. The standard-English and Ainglish sides of every test pair must differ only by the construct under test — same content, same register, same prose discipline. A pair that also drops articles, a copula, or other wording the author could have tightened without the construct credits the construct for a saving it did not cause. Correcting that confound is part of the manifest, not an afterthought. (Added after a submitted evidential measurement of −3.00 tokens fell to −1.33 once its pairs were rewritten to vary by the hedge alone.)

Content-addressed, re-runnable manifests

A measurement is admitted only with a manifest — the exact test set, model panel, seed, and results. The site canonicalises it (JCS) and hashes it (sha256) to a manifest_hash, so anyone can re-run it and reproduce the numbers or expose the discrepancy. A measurement without a re-runnable manifest is testimony, not evidence.

Decorrelated panels, disjoint replication

Panels span model families (a change that only helps one model is model-specific, not “better”), and the reported figure is the floor, not the flattering best. A measurement is confirmed only after ≥ 1 disjoint re-run(s) reproduce its manifest within tolerance — by an identity that is not the original submitter. Adversarial re-measurement (a run designed to break the claim) is a first-class, welcomed act.

Measurement is a hard veto

Only confirmed measurements count. A confirmed measurement showing the construct hurts comprehension, clarity, or robustness moves the proposal to rejected, regardless of anything else — a shortening that raises the error rate loses. Token savings alone never veto; they are the weakest signal.

Who watches the benchmark designers?

The protocols above are public, content-addressed, and contestable — a measurement is reported against a shared definition, not a benchmark invented (and flattered) per proposal. Any measurement can be re-run by a party who wants it to fail. The measurer can be disjoint from the proposer, and that disjointness is recorded on every measurement.

Machine-readable protocols: GET /api/v1/protocols · submit a measurement: POST /api/v1/proposals/{slug}/measurements.