Can agents learn unfamiliar Ainglish forms, and does that help them communicate? Here are actual tests, including the results that did not meet our safeguards.
Reviewed . Editorial selection of synthetic local research, not proposal-settlement evidence or an independent evaluation. A positive aggregate score does not erase a weaker result for an individual distinction.
Does training on Ainglish help more than teaching the same ideas in English?
No selective benefit established in the small pilot.
One Qwen2.5-7B-Instruct model, one training seed, three model conditions. Read Ainglish without a reference in the prompt. The frozen score weights 96 rows containing 84 distinct cases; it is not 96 independent tasks.
Reading Ainglish without a prompt reference
Model condition
Correct answers
Accuracy
Base model, no project training
72/96
75.00%
Trained on Ainglish examples
77/96
80.21%
Trained on matched English examples
77/96
80.21%
Ainglish-trained minus English-trained, reading Ainglish: 0.00 percentage points. Reported 95% interval: -12.50 to 12.50 points.
Declared safeguards did not all pass
The prespecified boundary screen failed: -5.56 percentage points after Ainglish training versus matched English training.
The interval is exploratory and clustered by 12 authored frames. The later duplicate audit did not change the frozen score.
What this does not establish
One cached model, one seed, small synthetic task families.
Held-out framings and names, not held-out concepts or independent human tasks.
Closed answer selection, not execution of real work. Token counts cover one reading turn only.
No tokenizer change, no external-lab training receipt, no governance progression claim.
SHA-256 of RESULT.json: 777010b3778b934ee538323f3ea02527afbc530fdeb15dca20341658f122fb4a
This page is generated from that version. Later work may add a successor entry; it does not silently replace this experiment’s result.
Local learning experiment · 6 September 2026
Can a broader teaching set improve unfamiliar wording without losing other skills?
An aggregate gain, but two family-level safeguards failed.
One Qwen2.5-7B-Instruct model and one training seed. The primary holdout has 252 distinct cases from 42 authored frames across six ratified families. The full campaign made 5,808 calls across five model conditions and several studies; those calls are not 5,808 independent test cases.
Reading Ainglish without a prompt reference
Model condition
Correct answers
Accuracy
Base model, no project training
177/252
70.24%
Trained on Ainglish examples
226/252
89.68%
Trained on matched English examples
198/252
78.57%
Ainglish-trained minus English-trained, reading Ainglish: +11.11 percentage points. Reported 95% interval: 2.38 to 20.24 points.
Declared safeguards did not all pass
Updating instructions: -6.25 points versus matched English training, below the −5-point screen.
English retention for alternatives: -8.33 points versus the base model, also below that screen.
These point-estimate screens are not statistical proofs of non-inferiority. The aggregate gain does not cancel them.
What this does not establish
Synthetic closed-answer tasks authored within the project, not independent human tasks or a lab replication.
The 95% interval resamples 42 authored frames, not 252 unrelated observations.
One base-model family and a fixed tokenizer. Training weights cannot change that tokenizer’s segmentation.
No demonstrated external adoption, governance settlement or general superiority claim.
SHA-256 of RESEARCH-RESULTS.json: 945f33e674061a068a6984a8fb94410d51d5dd1b09f4028e1f60b32c19357990
This page is generated from that version. Later work may add a successor entry; it does not silently replace this experiment’s result.
Local learning experiment · 6 September 2026
Does the learning advantage survive harder reasoning and different training seeds?
No repeatable advantage established over matched English training.
336 teaching cases per language, then 216 held-out cases from 36 newly authored reasoning frames across six families. Each language was trained with seeds 17, 29 and 43 on the same Qwen2.5-7B-Instruct base. These are three training seeds, not three model families. The wider campaign also tested 120 composition cases per language and condition: 4,704 target answers plus 84 controls.
Reading Ainglish without a prompt reference
Model condition
Correct answers
Accuracy
Base model, no project training
129/216
59.72%
Ainglish training, seed 17
139/216
64.35%
English training, seed 17
156/216
72.22%
Ainglish training, seed 29
151/216
69.91%
English training, seed 29
150/216
69.44%
Ainglish training, seed 43
143/216
66.20%
English training, seed 43
148/216
68.52%
Seed 17: Ainglish-trained minus English-trained, reading Ainglish: -7.87 percentage points. Reported 95% interval: -15.74 to -1.39 points.
Seed 29: Ainglish-trained minus English-trained, reading Ainglish: +0.46 percentage points. Reported 95% interval: -12.50 to 13.89 points.
Seed 43: Ainglish-trained minus English-trained, reading Ainglish: -2.31 percentage points. Reported 95% interval: -17.59 to 13.43 points.
Declared safeguards did not all pass
Seed 17 also failed the overall −5-point screen. None of the three seeds passed every family-level screen.
Seed 17, Ainglish reading versus matched English training: alternatives -13.89 points; deadline -5.56 points; multiplicity -5.56 points; participants -5.56 points; unknown -16.67 points. These all fall below the −5-point screen.
Seed 17, English retention versus the base model: alternatives -13.89 points; unknown -50.00 points. These all fall below the −5-point screen.
Seed 29, Ainglish reading versus matched English training: deadline -30.56 points. These all fall below the −5-point screen.
Seed 29, English retention versus the base model: deadline -16.67 points; unknown -33.33 points. These all fall below the −5-point screen.
Seed 43, Ainglish reading versus matched English training: participants -11.11 points; unknown -19.44 points; update -27.78 points. These all fall below the −5-point screen.
Seed 43, English retention versus the base model: unknown -22.22 points. These all fall below the −5-point screen.
What this does not establish
One base model, three training seeds, six ratified families, synthetic authored frames.
Frame bootstrap describes this held-out frame collection, not all language or human understanding.
Every seed and family is retained. English incumbent exposure differs; a future learning hypothesis does not nullify current costs or harm.
A -5pp point screen is not statistical proof of non-inferiority.
The newly authored reasoning frames still draw on known distinctions. Neither new names nor more seeds establish broad transfer.
The changed holdout and curriculum prevent a causal comparison with the earlier +11.11-point study. Its result is preserved rather than overwritten.
Composition results remain mixed and often weak. A correct answer on one distinction does not establish a correct joint plan.
SHA-256 of RESULTS.json: 4e2a794843868019b4eff4435a66f218b556f24cb195962e10ef2dfbfb62e092
This page is generated from that version. Later work may add a successor entry; it does not silently replace this experiment’s result.
What should change our confidence next?
The three-seed test makes transfer a sharper open question. Next priorities are independently authored tasks, different reader families and sender–receiver tasks that count the full cost of instructions and repairs. These are research questions, not promises of success. Read the plan for future training and tokenizer exposure.