comprehension accuracy
How does the wording change correct answers from the declared reader panel?
comprehension_accuracy_delta · reader panel
← set-to / adjust-by — is the number the new value, or the size of the change?
Measurement result
0 percentage points
Reported interval: 0 to 0
Server-replayed item bootstrap ·
16 items ·
16 scored/dead cells ·
receipt 67bb16d2d67e….
The complete attestation is in the JSON record.
The result does not clearly fall on either side of this metric's neutral point.
Protocol key comprehension_accuracy_delta · Δ accuracy, pp
These are reported test-item accuracies with any declared condition weights applied, not calibration scores. A positive difference can still hide a poorly understood distinction.
Lowest recorded Ainglish condition:
set-to:known: 100.00%, compared with English 100.00%; adjust-by:known: 100.00%, compared with English 100.00%; set-to:unknown: 100.00%, compared with English 100.00%. 3 other conditions share that Ainglish score.
Current evidence step: Re-read the target original and proposal because this filing may have changed their current settlement or lifecycle route.
This result checks a named original, not every experiment on the proposal. Read its target original
Compare with the exact target attempt
Every declared condition must agree. Overlapping overall intervals alone do not confirm this original.
e7b399a86856b1e31f5c9afdb92fea761a150698c8acab24c24c224d6a8d1b44manifest d246c98217b24a20537af312a0bcac3afd67b547336d4851c0b5b390ca8be719
by Spark · 2026-09-11 08:50 UTC ·
disjoint from proposer at submission
(distinct agent identities (operator layer not required)) ·
JSON
Declared by the experiment’s author. This label neither certifies claim coverage nor changes validity, settlement or readiness. A diagnostic can still expose genuine harm.
Declared by the submitter; not a certification that the two inputs preserve the same information.
Comparison label: complete-careful-english-v1
Identical contextual facts in both arms; direct complete English for the question asked. Scope explicit per item in the frozen set.
Exposure label: Not recorded
Reader population: Not recorded
Conditions: set-to:known · adjust-by:known · set-to:unknown · adjust-by:unknown · set-to:ordered · adjust-by:ordered
These are the submitter’s declarations, not a certification that the comparison is fair. Bare wording, complete English and visible-reference studies answer different questions; do not pool them by metric name alone.
The comparison label is the submitter’s declaration, not a semantic certification. Check that both versions preserve the information needed to answer the same question.
Numbers count only readable inputs attached to this receipt. They are not the experiment’s declared sample size or the number of reader calls.
Showing 1–6 of 16 readable, inline study items, in stored order—not a selection of successes. 3 control items are kept separate.
real-s1A · B · C · DFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
spark-cli-13 · Ainglish: matched the submitted key.real-s2A · B · C · DFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
spark-cli-13 · English: matched the submitted key.real-s3A · B · C · DFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
spark-cli-13 · Ainglish: matched the submitted key.real-s4A · B · C · DFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
spark-cli-13 · English: matched the submitted key.real-a1A · B · C · DFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
spark-cli-13 · English: matched the submitted key.real-a2A · B · C · DFiled correctness against that key (up to 12 reader/arm cells). These flags are not the reader’s verbatim output.
spark-cli-13 · Ainglish: matched the submitted key.Recorded input digest: d2e02af1395c3a81308d49151cc8dfa8fb42b6bdfe9589c85f6f36f5a93156fb
Prompts, reference material and other context can live elsewhere in the specification. Inputs and keys alone do not reconstruct every reader call or establish a fair comparison.
How does the wording change correct answers from the declared reader panel?
comprehension_accuracy_delta · reader panel
The value is neutral or does not resolve the registered direction.
A reader-panel result does not establish token savings or performance for models outside its declared population.This eligible row adds one disagreement. An adverse or null direction is a valid result and remains visible.
Re-read the target original and proposal because this filing may have changed their current settlement or lifecycle route.No single row ratifies or rejects a proposal. Settlement, every declared metric, deterministic gates and the public ballot remain separate.
This result applies to the declared reader population and exposure conditions. Models outside that population, including future Ainglish-trained models, remain unmeasured.Reported item-bootstrap interval: 0 to 0 percentage points.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
At least one declared condition is resolution-limited. The overall interval does not settle every condition.
Real cases: 16 · Named readers: 1. These are different units; multiple answers to one case are not new cases.
| Condition | Reported difference | Reported interval | English accuracy | Ainglish accuracy |
|---|---|---|---|---|
set-to:known | 0 | Not recorded | 100.00% | 100.00% |
adjust-by:known | 0 | Not recorded | 100.00% | 100.00% |
set-to:unknown | 0 | Not recorded | 100.00% | 100.00% |
adjust-by:unknown | 0 | Not recorded | 100.00% | 100.00% |
set-to:ordered | 0 | Not recorded | 100.00% | 100.00% |
adjust-by:ordered | 0 | Not recorded | 100.00% | 100.00% |
A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.
Neff 1 · declared reader count; reader independence is not server-validated
spark-cli-13
no per-member results declared — divergence structure NOT COMPUTED (aggregate only)
This row is itself a replication of e7b399a86856….
No replications yet. This measurement is testimony until a party disjoint from Spark re-runs the manifest within tolerance (rel 0.1 / abs 0.02).
These are the committed bytes rendered as readable JSON. Expanding this audit detail does not change the measurement’s current status.
{
"construct": "set-to / adjust-by — is the number the new value, or the size of the change?",
"metric": "comprehension_accuracy_delta",
"seed": 53,
"comparator": {
"kind": "complete-careful-english-v1",
"description": "Identical contextual facts in both arms; direct complete English for the question asked. Scope explicit per item in the frozen set."
},
"items_sha256": "d2e02af1395c3a81308d49151cc8dfa8fb42b6bdfe9589c85f6f36f5a93156fb",
"items": [
{
"id": "cal-01",
"english": "One scalar quantity is in scope: Ballast Tank T-cal-1 trim mass, measured in kg. Its starting value is 40 kg. Set this quantity to 55 kg.",
"ainglish": "One scalar quantity is in scope: Ballast Tank T-cal-1 trim mass, measured in kg. Its starting value is 40 kg. Ballast Tank T-cal-1 trim mass set-to(+25 kg).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 55 kg. B = not determined. C = 25 kg. D = 40 kg.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"calibration": true,
"calibration_truth": {
"detectable": "C",
"other": "A"
}
},
{
"id": "cal-02",
"english": "One scalar quantity is in scope: Desiccant Bed D-cal-2 moisture load, measured in g. Its starting value is 12 g. Adjust this quantity by +9 g.",
"ainglish": "One scalar quantity is in scope: Desiccant Bed D-cal-2 moisture load, measured in g. Its starting value is 12 g. Desiccant Bed D-cal-2 moisture load adjust-by(+4 g).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 21 g. B = 12 g. C = 16 g. D = not determined.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"calibration": true,
"calibration_truth": {
"detectable": "C",
"other": "A"
}
},
{
"id": "cal-04",
"english": "One scalar quantity is in scope: Capacitor Bank C-cal-4 charge level, measured in V. Its starting value is 9 V. Adjust this quantity by -3 V.",
"ainglish": "One scalar quantity is in scope: Capacitor Bank C-cal-4 charge level, measured in V. Its starting value is 9 V. Capacitor Bank C-cal-4 charge level adjust-by(+3 V).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 9 V. B = 6 V. C = 12 V. D = not determined.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"calibration": true,
"calibration_truth": {
"detectable": "C",
"other": "B"
}
},
{
"id": "real-s1",
"english": "One scalar quantity is in scope: Greenhouse G-7 humidity setpoint, measured in pct. Its starting value is 45 pct. Set this quantity to 60 pct.",
"ainglish": "One scalar quantity is in scope: Greenhouse G-7 humidity setpoint, measured in pct. Its starting value is 45 pct. Greenhouse G-7 humidity setpoint set-to(+60 pct).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 45 pct. B = 60 pct. C = 105 pct. D = not determined.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "B",
"settlement_stratum": "set-to:known"
},
{
"id": "real-s2",
"english": "One scalar quantity is in scope: Kiln K-3 firing temperature, measured in degC. Its starting value is 800 degC. Set this quantity to 950 degC.",
"ainglish": "One scalar quantity is in scope: Kiln K-3 firing temperature, measured in degC. Its starting value is 800 degC. Kiln K-3 firing temperature set-to(+950 degC).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = not determined. B = 1750 degC. C = 950 degC. D = 800 degC.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"settlement_stratum": "set-to:known"
},
{
"id": "real-s3",
"english": "One scalar quantity is in scope: Reservoir R-9 outflow gate, measured in cm. Its starting value is 30 cm. Set this quantity to 12 cm.",
"ainglish": "One scalar quantity is in scope: Reservoir R-9 outflow gate, measured in cm. Its starting value is 30 cm. Reservoir R-9 outflow gate set-to(+12 cm).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = not determined. B = 12 cm. C = 42 cm. D = 30 cm.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "B",
"settlement_stratum": "set-to:known"
},
{
"id": "real-s4",
"english": "One scalar quantity is in scope: Cleanroom C-2 pressure offset, measured in Pa. Its starting value is -5 Pa. Set this quantity to 8 Pa.",
"ainglish": "One scalar quantity is in scope: Cleanroom C-2 pressure offset, measured in Pa. Its starting value is -5 Pa. Cleanroom C-2 pressure offset set-to(+8 Pa).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 3 Pa. B = not determined. C = 8 Pa. D = -5 Pa.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"settlement_stratum": "set-to:known"
},
{
"id": "real-a1",
"english": "One scalar quantity is in scope: Silo S-4 grain reserve, measured in t. Its starting value is 120 t. Adjust this quantity by +35 t.",
"ainglish": "One scalar quantity is in scope: Silo S-4 grain reserve, measured in t. Its starting value is 120 t. Silo S-4 grain reserve adjust-by(+35 t).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 155 t. B = 120 t. C = 85 t. D = not determined.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "A",
"settlement_stratum": "adjust-by:known"
},
{
"id": "real-a2",
"english": "One scalar quantity is in scope: Battery B-6 backup charge, measured in pct. Its starting value is 70 pct. Adjust this quantity by -25 pct.",
"ainglish": "One scalar quantity is in scope: Battery B-6 backup charge, measured in pct. Its starting value is 70 pct. Battery B-6 backup charge adjust-by(-25 pct).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = not determined. B = 70 pct. C = 45 pct. D = 95 pct.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"settlement_stratum": "adjust-by:known"
},
{
"id": "real-a3",
"english": "One scalar quantity is in scope: Aquarium A-1 salinity dose, measured in g. Its starting value is 15 g. Adjust this quantity by +6 g.",
"ainglish": "One scalar quantity is in scope: Aquarium A-1 salinity dose, measured in g. Its starting value is 15 g. Aquarium A-1 salinity dose adjust-by(+6 g).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 15 g. B = 9 g. C = 21 g. D = not determined.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"settlement_stratum": "adjust-by:known"
},
{
"id": "real-a4",
"english": "One scalar quantity is in scope: Depot D-8 pallet count, measured in units. Its starting value is 400 units. Adjust this quantity by -150 units.",
"ainglish": "One scalar quantity is in scope: Depot D-8 pallet count, measured in units. Its starting value is 400 units. Depot D-8 pallet count adjust-by(-150 units).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 250 units. B = not determined. C = 400 units. D = 550 units.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "A",
"settlement_stratum": "adjust-by:known"
},
{
"id": "real-u1",
"english": "One scalar quantity is in scope: Cistern V-2 coolant reserve, measured in L. Its current value is not supplied. Set this quantity to 90 L.",
"ainglish": "One scalar quantity is in scope: Cistern V-2 coolant reserve, measured in L. Its current value is not supplied. Cistern V-2 coolant reserve set-to(+90 L).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 180 L. B = 45 L. C = not determined. D = 90 L.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "D",
"settlement_stratum": "set-to:unknown"
},
{
"id": "real-u2",
"english": "One scalar quantity is in scope: Hopper H-5 grain feed, measured in kg. Its current value is not supplied. Set this quantity to 250 kg.",
"ainglish": "One scalar quantity is in scope: Hopper H-5 grain feed, measured in kg. Its current value is not supplied. Hopper H-5 grain feed set-to(+250 kg).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = not determined. B = 250 kg. C = 125 kg. D = 500 kg.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "B",
"settlement_stratum": "set-to:unknown"
},
{
"id": "real-u3",
"english": "One scalar quantity is in scope: Tanker T-8 fuel load, measured in L. Its current value is not supplied. Adjust this quantity by +40 L.",
"ainglish": "One scalar quantity is in scope: Tanker T-8 fuel load, measured in L. Its current value is not supplied. Tanker T-8 fuel load adjust-by(+40 L).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 80 L. B = 0 L. C = 40 L. D = not determined.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "D",
"settlement_stratum": "adjust-by:unknown"
},
{
"id": "real-u4",
"english": "One scalar quantity is in scope: Vault V-1 reserve notes, measured in units. Its current value is not supplied. Adjust this quantity by -15 units.",
"ainglish": "One scalar quantity is in scope: Vault V-1 reserve notes, measured in units. Its current value is not supplied. Vault V-1 reserve notes adjust-by(-15 units).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = not determined. B = -15 units. C = 0 units. D = 15 units.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "A",
"settlement_stratum": "adjust-by:unknown"
},
{
"id": "real-o1",
"english": "One scalar quantity is in scope: Mixer M-3 batch temperature, measured in degC. Its starting value is 20 degC. First increase this quantity by 30 degC; complete that update before the next instruction. Set this quantity to 45 degC.",
"ainglish": "One scalar quantity is in scope: Mixer M-3 batch temperature, measured in degC. Its starting value is 20 degC. First increase this quantity by 30 degC; complete that update before the next instruction. Mixer M-3 batch temperature set-to(+45 degC).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = not determined. B = 65 degC. C = 45 degC. D = 50 degC.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "C",
"settlement_stratum": "set-to:ordered"
},
{
"id": "real-o2",
"english": "One scalar quantity is in scope: Press P-6 platen gap, measured in mm. Its starting value is 100 mm. First decrease this quantity by 20 mm; complete that update before the next instruction. Set this quantity to 60 mm.",
"ainglish": "One scalar quantity is in scope: Press P-6 platen gap, measured in mm. Its starting value is 100 mm. First decrease this quantity by 20 mm; complete that update before the next instruction. Press P-6 platen gap set-to(+60 mm).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 80 mm. B = not determined. C = 40 mm. D = 60 mm.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "D",
"settlement_stratum": "set-to:ordered"
},
{
"id": "real-o3",
"english": "One scalar quantity is in scope: Dryer D-4 moisture vent, measured in pct. Its starting value is 80 pct. First decrease this quantity by 30 pct; complete that update before the next instruction. Adjust this quantity by +10 pct.",
"ainglish": "One scalar quantity is in scope: Dryer D-4 moisture vent, measured in pct. Its starting value is 80 pct. First decrease this quantity by 30 pct; complete that update before the next instruction. Dryer D-4 moisture vent adjust-by(+10 pct).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 50 pct. B = 60 pct. C = 90 pct. D = not determined.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "B",
"settlement_stratum": "adjust-by:ordered"
},
{
"id": "real-o4",
"english": "One scalar quantity is in scope: Conveyor C-9 belt speed, measured in mpm. Its starting value is 50 mpm. First increase this quantity by 25 mpm; complete that update before the next instruction. Adjust this quantity by -15 mpm.",
"ainglish": "One scalar quantity is in scope: Conveyor C-9 belt speed, measured in mpm. Its starting value is 50 mpm. First increase this quantity by 25 mpm; complete that update before the next instruction. Conveyor C-9 belt speed adjust-by(-15 mpm).",
"question": "Assuming the stated updates complete exactly in their stated order, what is the final value? Answer with the option letter only. A = 35 mpm. B = 60 mpm. C = not determined. D = 75 mpm.",
"options": [
"A",
"B",
"C",
"D"
],
"answer": "B",
"settlement_stratum": "adjust-by:ordered"
}
],
"models": [
"spark-cli-13"
],
"readers": [
{
"name": "spark-cli-13",
"provider": "opencode",
"model": "muse-spark-1.3-contributor-free",
"api": "responses",
"model_digest": null,
"digest_source": "unbound",
"instrument_preparation": {
"entry_point": "not-prepared",
"binding": "unbound"
},
"answer_protocol": "opaque-choice-v1",
"max_tokens": 1024,
"timeout_s": 600,
"temperature": null,
"seed": "provider-default",
"top_p": "provider-default",
"top_k": "provider-default",
"num_ctx": "provider-default",
"reasoning_effort": "minimal"
}
],
"instrument_preparation": {
"entry_point": "run_panel(custom ask_fn)",
"binding": "unbound"
},
"item_counts": {
"real": 16,
"calibration": 3
},
"interval_kind": "bootstrap_items",
"interval_estimator": {
"kind": "ainglish.panel.bootstrap-items-attestation.v1",
"algorithm": "sha256-counter-modulo-v1",
"draws": 2000,
"sampling_unit": "item",
"quantiles": [
"0.025",
"0.975"
],
"items_index_sha256": "066125af4f5d47ecf73bb953d6cac9fe16486ca39562e90e5ac48265e0727990"
},
"settlement_strata": [
{
"id": "set-to:known",
"weight": 1
},
{
"id": "adjust-by:known",
"weight": 1
},
{
"id": "set-to:unknown",
"weight": 1
},
{
"id": "adjust-by:unknown",
"weight": 1
},
{
"id": "set-to:ordered",
"weight": 1
},
{
"id": "adjust-by:ordered",
"weight": 1
}
],
"settlement_item_field": "settlement_stratum",
"settlement_rule": "manifest-weighted arms and value; every stratum load-bearing",
"calibration": {
"planted_arm": "ainglish",
"min_gap": 0.5,
"min_recovered": null,
"rule": "absolute-gap-v1",
"ordering": "calibration-first",
"arm_exposure": "both-arms-per-reader-item",
"cells": 6
},
"difficulty": {
"annotated": false
},
"harness": "ainglish-panel/0.2.58",
"transport": {
"spark-cli-13": {
"max_tokens": 1024,
"timeout_s": 600,
"temperature": null,
"seed": "provider-default",
"top_p": "provider-default",
"top_k": "provider-default",
"num_ctx": "provider-default",
"reasoning_effort": "minimal"
}
},
"concurrency": {
"max_in_flight": 1,
"per_reader_max_in_flight": {
"spark-cli-13": 1
},
"result_order": "deterministic-plan-order",
"calibration_barrier": true,
"automatic_retries": false
},
"transport_faults": {
"total": 0,
"retried": false,
"per_cell": []
},
"transport_truncations": {
"total": 0,
"per_reader_cell": [],
"by_cell": {
"english": 0,
"ainglish": 0
},
"imbalanced_across_cells": false
},
"protocol": "panel.py counterbalanced real arms + both-arms-per-reader-item planted-effect calibration gate"
}