Agent-task benchmark
Sentence comprehension is not the end goal. The receiving agent has to choose the right action, avoid the wrong one and know when it needs clarification. This frozen benchmark tests that operational question against both ambiguous and equally explicit English.
Across 22 already-installed model artifacts, prompt-cold Ainglish usually led to fewer correct immediate actions than careful English that stated the same intent in full. It usually did better than ordinary ambiguous English. One short definition narrowed, but did not close, the careful-English gap overall.
Project-operated descriptive result, published . “Cold” means absent from the served prompt; the models' actual pretraining exposure is unknown. Completely failed and schema-invalid calls remain in the denominators.
What happened
The primary comparator is careful English, because it carries the same complete source intent. A positive effect means Ainglish produced more zero-repair successes on the same frozen items; a negative effect means careful or bare English did better. These are reader-level descriptive effects, not confidence intervals or independent replications.
| Comparison | Track | Mean effect | Median effect | Readers + / = / − |
|---|---|---|---|---|
| Ainglish minus careful English | Cold | -13.6 pp | -15.9 pp | 1 / 4 / 17 |
| Ainglish minus careful English | One exposure | -3.5 pp | -2.3 pp | 5 / 6 / 11 |
| Ainglish minus bare English | Cold | +7.4 pp | +4.5 pp | 13 / 7 / 2 |
| Ainglish minus bare English | One exposure | +16.5 pp | +18.2 pp | 19 / 2 / 1 |
In the cold track, Ainglish trailed equally explicit careful English by 13.6 percentage points on the mean reader effect, while leading bare ambiguous English by 7.4 points. After one local definition, those contrasts were −3.5 and +16.5 points. The comparator changes the conclusion: beating an ambiguity baseline is useful, but does not show that the compact form is yet as usable as saying the full meaning in familiar English.
Reader-family robustness
Six of the 22 artifacts are Qwen variants, so a post-hoc audit also grouped the roster into 16 declared model lineages and gave each lineage equal weight. This reduces one obvious source of pseudo-replication; it does not make the lineages independent or turn a convenience roster into a model-population sample.
| Comparison | Track | Equal-lineage mean | Descriptive bootstrap range | Leave one lineage out |
|---|---|---|---|---|
| Ainglish minus careful English | Cold | -11.7 pp | -15.9 to -7.4 pp | -12.5 to -10.6 pp |
| Ainglish minus careful English | One exposure | -3.6 pp | -8.5 to +1.4 pp | -5.1 to -2.6 pp |
| Ainglish minus bare English | Cold | +5.1 pp | +0.9 to +9.9 pp | +3.6 to +6.1 pp |
| Ainglish minus bare English | One exposure | +11.8 pp | +6.8 to +16.7 pp | +10.6 to +13.2 pp |
No leave-one-lineage-out calculation reversed any of the four aggregate directions. One supplied
definition improved Ainglish's equal-lineage effect by
+8.1 pp against careful English and
+6.7 pp against bare English. Those ranges
are exploratory stability summaries, not confidence intervals or population-level inference.
Construct results were heterogeneous: the largest persistent deficit was
true-as-worded / false-as-worded
(-47.8 pp cold;
-43.8 pp after one definition),
with only two frozen items per construct. Inspect
the deterministic post-hoc audit.
All-row immediate-action rates
| Track | Wording arm | Correct without repair | Rate |
|---|---|---|---|
| Cold | Ainglish | 221 / 484 | 45.7% |
| Cold | Bare English | 185 / 484 | 38.2% |
| Cold | Careful English | 287 / 484 | 59.3% |
| One exposure | Ainglish | 267 / 484 | 55.2% |
| One exposure | Bare English | 187 / 484 | 38.6% |
| One exposure | Careful English | 284 / 484 | 58.7% |
Failures stayed visible
One reader, solar-pro:22b, returned HTTP 500 in all 132 cells. Another,
deepseek-v2:16b, violated the exact decision schema in 127 of 132 first outputs.
Neither was removed after the result was known. Invalid outputs across the full roster were common
and similarly distributed among arms, so the run measures operational compatibility as well as
semantic understanding.
Present interaction cost
| Track | Ainglish | Careful English |
|---|---|---|
| Cold | 229 | 223 |
| One exposure | 246 | 223 |
No present token-efficiency win appeared: mean raw interaction tokens were 229 versus 223 in the cold track and 246 versus 223 after one exposure. The definition is correctly charged to Ainglish; provider token coverage was 462 / 484 rows in each group. These counts come from different model tokenizers, so they are descriptive—not one tokenizer-independent efficiency estimate or a forecast for models trained on Ainglish.
What this says about training exposure
Familiar English has an incumbent training-data advantage; Ainglish was not known to be present in these readers' training corpora. That makes future training exposure an important hypothesis, but does not make today's adverse comparison disappear. A separate 264-observation comparison held one Qwen 2.5 7B base model and tokenizer fixed and added a tiny existing Ainglish LoRA. It did not improve immediate Ainglish task execution, its final gains were not selective to Ainglish, and it used more interaction tokens in every group.
That adapter had only 76 training rows. It is not a test of including a release in foundation-model pretraining, nor of tokenizer integration. The future claim therefore remains open and testable: larger, declared training exposure may close the familiarity gap, but the project has not shown that yet. Inspect the controlled-exposure receipt.
A separate semantics stress test
A 540-cell adversarial gauntlet supplied each construct's exact meaning to 3 installed readers. Their accuracy ranged from 88.9–95.6%, but 19 of 38 errors occurred when the correct answer was “underdetermined.” The dominant failure was treating “not stated” as “stated false.” This is a documentation and regression signal, not cold-comprehension, human-intuitiveness or governance evidence. Inspect the frozen gauntlet.
The question it asks
A source intent is fixed outside the message. A receiving agent sees one wording and a small set of executable actions. It may act immediately or ask for clarification. If it asks, it receives one scripted repair turn that states the same source intent without changing the task.
Bare English
The relevant distinction is left implicit. This is an ambiguity baseline, not a uniquely correct interpretation test.
Careful English
The complete source intent is expressed in ordinary English. This is the load-bearing comparator.
Ainglish
The same intent is expressed using the registered compact form. A win over bare English alone is insufficient.
One task, all three arms
Source intent: the receiving agent is included in the group expected to review the draft.
“We will review the release draft before publication.”
“We, including you as the receiving agent, will review the release draft before publication.”
we-including-you will review the release draft before publication.
What gets reported
- Zero-repair task success: the first action set exactly matches a valid set.
- Final task success: the action is correct after at most one scripted clarification.
- Wrong-action rate: the agent acts, but its action set is unsafe, incomplete or excessive.
- Clarification and invalid-output rates: neither disappears from the denominator.
- Repair-inclusive tokens and latency: reported only with coverage, provider and tokenizer or usage source.
There is no composite score: shorter text cannot compensate for a wrong action. Results remain item-level and should be stratified by construct and reader family before anyone reports an overall figure.
Exposure tracks in the frozen design
| Track | What the reader gets | What it can mean |
|---|---|---|
| Cold | No Ainglish definition in the served prompt. | Immediate compatibility only. It does not prove that Ainglish was absent from training data. |
| One exposure | The Ainglish arm receives one short definition, whose tokens count toward cost. | Practical use after one explicit teaching exposure. It is not zero-shot or pretrained adoption. |
A future model-adaptation track would need a new frozen task release, training receipts and held-out items. It cannot be silently added to this version.
Audit or run it
The result is pinned to commit bbea0f8528484f780e591a727241f0fbe54e8d62.
Its raw-response SHA-256 is d23e8da6573e05fe1fe79c762ffb4bc7eaddd169feeff0a78319a91044b93b2f and its
analysis SHA-256 is 78dacc62b827dd053a6e38c894a62b9e36e5d4f347fc8f34d5ec54598e86623d. The original
protocol artifact remains pinned to commit
028052715cfa61744fab0ca92268f71073de2246 and register
0.35.0. The post-hoc lineage audit is pinned to commit
116d7bef0b423ce6c53811516288f9344df85708
with result SHA-256 5cb40b9e79ff6421ae361351791b7cadcaab648a451937252f472653621ae56e.
python3 benchmark.py validate
python3 benchmark.py export --track cold --arm all --seed 20260828 > prompts.jsonl
python3 benchmark.py score responses.jsonl
A project-operated run remains internal evidence even when it uses several model names. This benchmark cannot establish human intuitiveness, external adoption, independent validation, pretraining inclusion, future tokenizer efficiency or superiority of Ainglish overall. It also does not license projecting today's adverse result unchanged onto future trained models.
Read the research status, inspect the measurement methodology, or use the frozen packet to publish a result that disagrees with the project.