The interesting question is not just whether a new phrase sounds useful. It is what a fair test reveals—and what we change when it disappoints.
These editorial examples explain observations filed on 5 September 2026. They are not a shortlist for adoption. Current receipt status is shown separately below.
A useful distinction can still lose to careful English
“The check failed” can mean that it found a problem or that it never reached a result. “Smoke suite: verdict-fail” and “Smoke suite: no-verdict (timeout)” mark that distinction.
Plain English remains an option. Ordinary English can say “The smoke suite ran and found the deployment defective” or “The smoke suite timed out before reaching a result.”
Two matched studies kept the Ainglish cases but changed the English comparison. The filed result was favourable against bare “failed” and adverse against the complete wording. These are different questions, not two independent confirmations. Clarity of the distinction does not establish that these readers benefit from its new spelling.
Against bare “failed”, with a shared log
9.6 percentage points; reported bounds 3.8179 to 15.7058.
Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result.
Read the separate outcomes and current settlement role
Awaiting independent settlement. An original reports one result. It does not confirm itself.
How often did each version lead to the right answer?
English comparison
81.10%
Ainglish version
90.70%
Reported real-item accuracy, not the separate calibration score. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.
Real cases: 256 · Named readers: 2. These are different units; multiple answers to one case are not new cases.
Does the overall result hide differences between conditions?
Every declared condition, with its stored result. Condition names come from the frozen experiment.
-6.545 percentage points; reported bounds -10.5283 to -2.6532.
Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result.
Read the separate outcomes and current settlement role
Awaiting independent settlement. An original reports one result. It does not confirm itself.
How often did each version lead to the right answer?
English comparison
97.25%
Ainglish version
90.70%
Reported real-item accuracy, not the separate calibration score. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.
Real cases: 256 · Named readers: 2. These are different units; multiple answers to one case are not new cases.
Does the overall result hide differences between conditions?
Every declared condition, with its stored result. Condition names come from the frozen experiment.
With a timeout of 10 seconds, “set-to(30 seconds)” requests 30 seconds; “adjust-by(+30 seconds)” requests 40 seconds.
Plain English remains an option. The direct English alternatives are “Set the timeout to 30 seconds” and “Increase the timeout by 30 seconds.”
The filed overall result was uncertain, while the known-starting-value set-to condition was worse for Ainglish. Unknown starting values and ordered updates were also retained. Read the separate conditions before treating the overall average as a reason to adopt or reject both forms.
Fixed-reader, six-condition study
3.9433 percentage points; reported bounds -5.4759 to 13.8449.
Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result.
Read the separate outcomes and current settlement role
Awaiting independent settlement. An original reports one result. It does not confirm itself.
How often did each version lead to the right answer?
English comparison
62.34%
Ainglish version
66.28%
Reported real-item accuracy, not the separate calibration score. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.
Real cases: 192 · Named readers: 2. These are different units; multiple answers to one case are not new cases.
Does the overall result hide differences between conditions?
Every declared condition, with its stored result. Condition names come from the frozen experiment.
“I inspected 37 of 284 invoices” includes the size of the whole population. “part-chosen(amount-band): the 37 invoices” does not carry the missing 284.
Plain English remains an option. Keep “37 of 284 invoices” in both versions, then compare the explanation of why that subset was examined: deliberately chosen, or restricted by a named limit.
The earlier arithmetic can be correct while the comparison omits information. A new complete-information cost study retains both counts and the reason for the boundary. It is a new original on different inputs, not a numerical correction or a confirmation of the earlier result. Neither cost study establishes comprehension.
Earlier source, with the information-loss concern
-15.5 tokens per complete pair; reported bounds -15.5625 to -15.5.
Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result.
Read the separate outcomes and current settlement role
Disputed. Eligible replications disagree and this original does not hold a settlement majority.
Does the overall result hide differences between conditions?
Every declared condition, with its stored result. Condition names come from the frozen experiment.
These studies use current models and tokenizers with extensive exposure to English. A visible Ainglish reference is a separate condition, not a substitute for training on the language. Future-trained efficiency is a hypothesis to test, not a benefit measured by these receipts.