Ainglish An English dialect for AI agents

Experimental results

Can agents learn unfamiliar Ainglish forms, and does that help them communicate? Here are actual tests, including the results that did not meet our safeguards.

Reviewed . Editorial selection of synthetic local research, not proposal-settlement evidence or an independent evaluation. A positive aggregate score does not erase a weaker result for an individual distinction.

Download these research summariesInspect proposal measurements instead

Local learning experiment · 6 September 2026

Does training on Ainglish help more than teaching the same ideas in English?

No selective benefit established in the small pilot.

One Qwen2.5-7B-Instruct model, one training seed, three model conditions. Read Ainglish without a reference in the prompt. The frozen score weights 96 rows containing 84 distinct cases; it is not 96 independent tasks.

Reading Ainglish without a prompt reference
Model conditionCorrect answersAccuracy
Base model, no project training72/9675.00%
Trained on Ainglish examples77/9680.21%
Trained on matched English examples77/9680.21%

Ainglish-trained minus English-trained, reading Ainglish: 0.00 percentage points. Reported 95% interval: -12.50 to 12.50 points.

Declared safeguards did not all pass

  • The prespecified boundary screen failed: -5.56 percentage points after Ainglish training versus matched English training.
  • The interval is exploratory and clustered by 12 authored frames. The later duplicate audit did not change the frozen score.

What this does not establish

  • One cached model, one seed, small synthetic task families.
  • Held-out framings and names, not held-out concepts or independent human tasks.
  • Closed answer selection, not execution of real work. Token counts cover one reading turn only.
  • No tokenizer change, no external-lab training receipt, no governance progression claim.

Inspect the complete resultRead the frozen designBrowse inputs, code and receipts

Verify this snapshotExact source version and content digest

Source commit: 54d8d282762e7904a2441f883a46ca555a3b06aa

SHA-256 of RESULT.json: 777010b3778b934ee538323f3ea02527afbc530fdeb15dca20341658f122fb4a

This page is generated from that version. Later work may add a successor entry; it does not silently replace this experiment’s result.

Local learning experiment · 6 September 2026

Can a broader teaching set improve unfamiliar wording without losing other skills?

An aggregate gain, but two family-level safeguards failed.

One Qwen2.5-7B-Instruct model and one training seed. The primary holdout has 252 distinct cases from 42 authored frames across six ratified families. The full campaign made 5,808 calls across five model conditions and several studies; those calls are not 5,808 independent test cases.

Reading Ainglish without a prompt reference
Model conditionCorrect answersAccuracy
Base model, no project training177/25270.24%
Trained on Ainglish examples226/25289.68%
Trained on matched English examples198/25278.57%

Ainglish-trained minus English-trained, reading Ainglish: +11.11 percentage points. Reported 95% interval: 2.38 to 20.24 points.

Declared safeguards did not all pass

  • Updating instructions: -6.25 points versus matched English training, below the −5-point screen.
  • English retention for alternatives: -8.33 points versus the base model, also below that screen.
  • These point-estimate screens are not statistical proofs of non-inferiority. The aggregate gain does not cancel them.

What this does not establish

  • Synthetic closed-answer tasks authored within the project, not independent human tasks or a lab replication.
  • The 95% interval resamples 42 authored frames, not 252 unrelated observations.
  • One base-model family and a fixed tokenizer. Training weights cannot change that tokenizer’s segmentation.
  • No demonstrated external adoption, governance settlement or general superiority claim.

Inspect the complete resultRead the frozen designBrowse inputs, code and receipts

Verify this snapshotExact source version and content digest

Source commit: 54d8d282762e7904a2441f883a46ca555a3b06aa

SHA-256 of RESEARCH-RESULTS.json: 945f33e674061a068a6984a8fb94410d51d5dd1b09f4028e1f60b32c19357990

This page is generated from that version. Later work may add a successor entry; it does not silently replace this experiment’s result.

Local learning experiment · 6 September 2026

Does the learning advantage survive harder reasoning and different training seeds?

No repeatable advantage established over matched English training.

336 teaching cases per language, then 216 held-out cases from 36 newly authored reasoning frames across six families. Each language was trained with seeds 17, 29 and 43 on the same Qwen2.5-7B-Instruct base. These are three training seeds, not three model families. The wider campaign also tested 120 composition cases per language and condition: 4,704 target answers plus 84 controls.

Reading Ainglish without a prompt reference
Model conditionCorrect answersAccuracy
Base model, no project training129/21659.72%
Ainglish training, seed 17139/21664.35%
English training, seed 17156/21672.22%
Ainglish training, seed 29151/21669.91%
English training, seed 29150/21669.44%
Ainglish training, seed 43143/21666.20%
English training, seed 43148/21668.52%

Seed 17: Ainglish-trained minus English-trained, reading Ainglish: -7.87 percentage points. Reported 95% interval: -15.74 to -1.39 points.

Seed 29: Ainglish-trained minus English-trained, reading Ainglish: +0.46 percentage points. Reported 95% interval: -12.50 to 13.89 points.

Seed 43: Ainglish-trained minus English-trained, reading Ainglish: -2.31 percentage points. Reported 95% interval: -17.59 to 13.43 points.

Declared safeguards did not all pass

  • Seed 17 also failed the overall −5-point screen. None of the three seeds passed every family-level screen.
  • Seed 17, Ainglish reading versus matched English training: alternatives -13.89 points; deadline -5.56 points; multiplicity -5.56 points; participants -5.56 points; unknown -16.67 points. These all fall below the −5-point screen.
  • Seed 17, English retention versus the base model: alternatives -13.89 points; unknown -50.00 points. These all fall below the −5-point screen.
  • Seed 29, Ainglish reading versus matched English training: deadline -30.56 points. These all fall below the −5-point screen.
  • Seed 29, English retention versus the base model: deadline -16.67 points; unknown -33.33 points. These all fall below the −5-point screen.
  • Seed 43, Ainglish reading versus matched English training: participants -11.11 points; unknown -19.44 points; update -27.78 points. These all fall below the −5-point screen.
  • Seed 43, English retention versus the base model: unknown -22.22 points. These all fall below the −5-point screen.

What this does not establish

  • One base model, three training seeds, six ratified families, synthetic authored frames.
  • Frame bootstrap describes this held-out frame collection, not all language or human understanding.
  • Every seed and family is retained. English incumbent exposure differs; a future learning hypothesis does not nullify current costs or harm.
  • A -5pp point screen is not statistical proof of non-inferiority.
  • The newly authored reasoning frames still draw on known distinctions. Neither new names nor more seeds establish broad transfer.
  • The changed holdout and curriculum prevent a causal comparison with the earlier +11.11-point study. Its result is preserved rather than overwritten.
  • Composition results remain mixed and often weak. A correct answer on one distinction does not establish a correct joint plan.

Inspect the complete resultRead the frozen designBrowse inputs, code and receipts

Verify this snapshotExact source version and content digest

Source commit: 925ffb89d13e7da6ac901095d72d95448f942199

SHA-256 of RESULTS.json: 4e2a794843868019b4eff4435a66f218b556f24cb195962e10ef2dfbfb62e092

This page is generated from that version. Later work may add a successor entry; it does not silently replace this experiment’s result.

What should change our confidence next?

The three-seed test makes transfer a sharper open question. Next priorities are independently authored tasks, different reader families and sender–receiver tasks that count the full cost of instructions and repairs. These are research questions, not promises of success. Read the plan for future training and tokenizer exposure.