Separate outcomes retained for all 2 declared conditions
verdict-fail, no-verdict
An overall average can hide a weak condition. A condition list is not proof that every form or claim in the proposal was tested.
Settlement role
Awaiting independent settlement
An original reports one result. It does not confirm itself.
How often did each version lead to the right answer?
English comparison
81.10%
Ainglish version
90.70%
Reported real-item accuracy, not the separate calibration score. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.
Real cases: 256 · Named readers: 2. These are different units; multiple answers to one case are not new cases.
Does the overall result hide differences between conditions?
Every declared condition, with its stored result. Condition names come from the frozen experiment.
Condition
Reported difference
English accuracy
Ainglish accuracy
verdict-fail
1.8
84.25%
86.05%
no-verdict
17.4
77.95%
95.35%
Inspect actual inputs and recorded answers
The comparison label is the submitter’s declaration, not a semantic certification. Check that both versions preserve the information needed to answer the same question.
No readable non-control input pairs are stored inline in this receipt. This does not mean the experiment used no inputs.
Open the declared external input artifact. The website has not fetched or verified it. Verify the declared digest recipe before relying on it: SDK item digests use canonical JSON, not the raw pretty-printed file bytes.
Recorded input digest: 750d2327a6774ae032d4d82c1882499e7606e3924b84f3f8954462e9efae3cf6
Prompts, reference material and other context can live elsewhere in the specification. Inputs and keys alone do not reconstruct every reader call or establish a fair comparison.
Declared population, method and retained outcomes
No structured study scope is declared here. Inspect the immutable manifest; do not infer a comparator or population from the headline.
Absolute arm results, reader-specific results and condition results below are retained values, not a newly pooled analysis. Accuracy arms use fractions from 0 to 1; their difference uses percentage points.
Separate outcomes retained for all 2 declared conditions
verdict-fail, no-verdict
An overall average can hide a weak condition. A condition list is not proof that every form or claim in the proposal was tested.
Settlement role
Awaiting independent settlement
An original reports one result. It does not confirm itself.
How often did each version lead to the right answer?
English comparison
97.25%
Ainglish version
90.70%
Reported real-item accuracy, not the separate calibration score. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.
Real cases: 256 · Named readers: 2. These are different units; multiple answers to one case are not new cases.
Does the overall result hide differences between conditions?
Every declared condition, with its stored result. Condition names come from the frozen experiment.
Condition
Reported difference
English accuracy
Ainglish accuracy
verdict-fail
-13.95
100.00%
86.05%
no-verdict
0.86
94.49%
95.35%
Inspect actual inputs and recorded answers
The comparison label is the submitter’s declaration, not a semantic certification. Check that both versions preserve the information needed to answer the same question.
No readable non-control input pairs are stored inline in this receipt. This does not mean the experiment used no inputs.
Open the declared external input artifact. The website has not fetched or verified it. Verify the declared digest recipe before relying on it: SDK item digests use canonical JSON, not the raw pretty-printed file bytes.
Recorded input digest: 927bcf9a7ad6bed3fac03b17c9b24dec5367823a61c4bf11c918cdab889053fb
Prompts, reference material and other context can live elsewhere in the specification. Inputs and keys alone do not reconstruct every reader call or establish a fair comparison.
Declared population, method and retained outcomes
No structured study scope is declared here. Inspect the immutable manifest; do not infer a comparator or population from the headline.
Absolute arm results, reader-specific results and condition results below are retained values, not a newly pooled analysis. Accuracy arms use fractions from 0 to 1; their difference uses percentage points.
Different wording, readers, exposure or populations can legitimately produce different results. A visible reference is not training the model’s weights. Current models and tokenizers have learned English; future Ainglish-trained performance remains a research question.