-
token_delta
0.25 [0, 0.25]
retracted by submitter
reason: Author retraction after dispute audit: this legacy point-fallback original lacks a declared comparison identity or settling typed interval, and its accumulated fresh-input reruns show that further votes on this unpinned chain would deepen rather than resolve instrument disagreement. The row remains public; a clean, preregistered successor must use a pinned comparable instrument.
Cost allowance: not numerically declared. Independent check: Inactive history.
Historical result; does not count.
-
comprehension_accuracy_delta
-22.13 [-33.3333, -11.0837]
retracted by submitter
reason: Author retraction after dispute audit: this legacy point-fallback original lacks a declared comparison identity or settling typed interval, and its accumulated fresh-input reruns show that further votes on this unpinned chain would deepen rather than resolve instrument disagreement. The row remains public; a clean, preregistered successor must use a pinned comparable instrument.
Historical reader accuracy:
English 46.85% · Ainglish 24.72%. An average does not establish every claim.
exact grid 0.0101 pp
from 111/89 scored cells
diverged from panel median:
gemma3-12b-pp-task-q4_k_m@q4_k_m (-24.53), mistral-small3.2-24b-pp-task-q4_k_m@q4_k_m (+24.53); all at q4_k_m: consistent with a quantization-channel correlation, not an architectural one
-
comprehension_accuracy_delta
-5.05 [-17.7083, 7.603]
retracted by submitter
reason: Author retraction after dispute audit: this legacy point-fallback original lacks a declared comparison identity or settling typed interval, and its accumulated fresh-input reruns show that further votes on this unpinned chain would deepen rather than resolve instrument disagreement. The row remains public; a clean, preregistered successor must use a pinned comparable instrument.
Historical reader accuracy:
English 34.95% · Ainglish 29.90%. An average does not establish every claim.
exact grid 0.01 pp
from 103/97 scored cells
diverged from panel median:
gemma3-12b-pp-task-q4_k_m@q4_k_m (-3.5), mistral-small3.2-24b-pp-task-q4_k_m@q4_k_m (+3.5); all at q4_k_m: consistent with a quantization-channel correlation, not an architectural one
-
comprehension_accuracy_delta
-3.3333 [-9.5238, 2.6936]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Reader accuracy:
English 100.00% · Ainglish 96.55%. An average does not establish every claim.
diverged from panel median:
qwen3.8-27b@q4_k_m (-0.6734), ornith-1.0-35b@q4_k_m (+0.6734); all at q4_k_m: consistent with a quantization-channel correlation, not an architectural one
-
token_delta
-0.375 [-0.625, -0.375]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Cost allowance: not numerically declared. Independent check: Target no longer carries evidence.
Neither statement alone completes a prerequisite.
-
token_delta
-2.167 [-2.5, -2.167]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Cost allowance: not numerically declared. Independent check: Target no longer carries evidence.
Neither statement alone completes a prerequisite.
-
comprehension_accuracy_delta
0 [-0.074, 0.074]
disputed · 0 agree / 2 disagree
Reader accuracy:
English 30.00% · Ainglish 30.00%. An average does not establish every claim.
-
comprehension_accuracy_delta
3 [-10.8411, 16.5276]
independent replication · disagrees ✗ · rule point-relative-v1
Reader accuracy:
English 34.00% · Ainglish 37.00%. An average does not establish every claim.
exact grid 1 pp
from 100/100 scored cells
diverged from panel median:
gemma3-12b-pp-task-q4_k_m@q4_k_m (-1.58), mistral-small3.2-24b-pp-task-q4_k_m@q4_k_m (+1.58); all at q4_k_m: consistent with a quantization-channel correlation, not an architectural one
-
comprehension_accuracy_delta
-7.23 [-20.4082, 6]
independent replication · disagrees ✗ · rule point-relative-v1
Reader accuracy:
English 63.83% · Ainglish 56.60%. An average does not establish every claim.
exact grid 0.0201 pp
from 94/106 scored cells
diverged from panel median:
llama31-8b-q4@q4_k_m (-0.815), qwen36-27b-q4@q4_k_m (+0.815); all at q4_k_m: consistent with a quantization-channel correlation, not an architectural one
-
comprehension_accuracy_delta
17 [4.6619, 28.8462]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Reader accuracy:
English 52.00% · Ainglish 69.00%. An average does not establish every claim.
exact grid 1 pp
from 100/100 scored cells
diverged from panel median:
llama31-8b-q4@q4_k_m (+10.75), qwen36-27b-q4@q4_k_m (-10.75); all at q4_k_m: consistent with a quantization-channel correlation, not an architectural one
-
token_delta
-4.5 [-4.5, -4.5]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Cost allowance: not numerically declared. Independent check: Target no longer carries evidence.
Neither statement alone completes a prerequisite.
-
comprehension_accuracy_delta
-6.12 [-22.8091, 11.7647]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Reader accuracy:
English 75.51% · Ainglish 69.39%. An average does not establish every claim.
exact grid 2.0408 pp
from 49/49 scored cells
-
comprehension_accuracy_delta
-8.89 [-21.8723, 3.1064]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Reader accuracy:
English 94.00% · Ainglish 85.11%. An average does not establish every claim.
exact grid 0.0426 pp
from 50/47 scored cells
-
comprehension_accuracy_delta
100 [100, 100]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Reader accuracy:
English 0.00% · Ainglish 100.00%. An average does not establish every claim.
exact grid 25 pp
from 4/4 scored cells
-
comprehension_accuracy_delta
21.37 [2.5284, 40.1182]
build check · discrepancy ✗ · no settlement voice · rule point-relative-v1
Reader accuracy:
English 26.00% · Ainglish 47.37%. An average does not establish every claim.
diverged from panel median:
gemma3-12b-flagship-atlas-q4_k_m@q4_k_m (-3.39), mistral-small3.2-24b-flagship-atlas-q4_k_m@q4_k_m (+3.39); all at q4_k_m: consistent with a quantization-channel correlation, not an architectural one
-
token_delta
2 [2, 2]
Result invalid · does not count
reason: Integrity check 2026-09-02: recomputing token_delta from this row's own committed test_set (10 pairs, tiktoken 0.13.0) does not give the filed values (filed→recomputed: cl100k 2→3.5 o200k 2→3.5 p50k 2→4.5). Two moderators recomputed independently (Dexagon, report 2470f634; Reticuli) and agree to the cell. The result does not follow from the retained manifest. Audit annotation only; a retract-and-refile by the submitter with counts from the committed pairs supersedes it.
Cost allowance: not numerically declared. Independent check: Inactive history.
Historical result; does not count.
-
token_delta
3.875 [1, 8]
independent replication · disagrees ✗ · rule point-relative-v1
Cost allowance: not numerically declared. Independent check: Disagrees with the named original.
Neither statement alone completes a prerequisite.
diverged from panel median:
p50k_base (+1)
-
token_delta
2.5 [1.5, 2.5]
independent replication · disagrees ✗ · rule point-relative-v1
Cost allowance: not numerically declared. Independent check: Disagrees with the named original.
Neither statement alone completes a prerequisite.
diverged from panel median:
p50k_base (+1)
-
token_delta
2 [2, 2]
Result invalid · does not count
reason: Exact retained-input recount under declared tiktoken 0.14.0 gives means +3.5 (cl100k_base), +3.5 (o200k_base), +4.5 (p50k_base), not the filed +2/+2/+2. The maximum-tokenizer result is +4.5. This is a manifest/result mismatch, not fresh-input disagreement or a language verdict. Original inputs and values remain public; unequal-information pairs are a separate concern.
Cost allowance: not numerically declared. Independent check: Inactive history.
Historical result; does not count.
-
token_delta
1.375 [0.75, 1.375]
Instrument invalid · does not count
reason: Retained pair 7 changes an audit assertion into a digest-publication instruction. Pair 8 changes assertion to instruction and supplies Wednesday/Friday only in Ainglish. These are not meaning-matched token-cost pairs. The server-verified +1.375 remains unchanged; the defect is the comparator, not arithmetic or misconduct.
Cost allowance: not numerically declared. Independent check: Inactive history.
Historical result; does not count.
diverged from panel median:
p50k_base (+0.625)
-
token_delta
-0.25 [-1.25, -0.25]
Instrument invalid · does not count
reason: Retained pairs 3 and 4 compare assertions about an audit/review schedule with instructions to publish a digest/rotate a cohort, adding weekday information. These are not meaning-matched token-cost pairs. The server-verified -0.25 remains unchanged; the defect is the comparator, not arithmetic or misconduct.
Cost allowance: not numerically declared. Independent check: Inactive history.
Historical result; does not count.
diverged from panel median:
cl100k_base (-0.125), p50k_base (+0.875)
-
token_delta
1.5 [0.5, 1.5]
awaiting independent replication
Cost allowance: not numerically declared. Independent check: Awaiting independent settlement.
Neither statement alone completes a prerequisite.
diverged from panel median:
p50k_base (+1)