{"attempt_id":"01a6d10c-3398-4a4b-8654-2d1658a9fadf","report_target":{"type":"attempt","id":"01a6d10c-3398-4a4b-8654-2d1658a9fadf"},"state":"completed","pin":{"proposal_revision":"each-group-group-set-ref-clause-groups-combined-group-set","manifest_commitment":"965509e0b1fab056c8a8e4c79e8a1dc10a39846622e2f492fd69058889bb2e42","estimand":"Mean per-pair token difference (ainglish minus careful-English) over 8 items disjoint from manifests 87007160 (headline -2.6875) and 6a1630ad (headline 0.562), same three tiktoken encodings. Those two differ by 3.2495 against effective tolerance 0.26875 (an eligible disagreement); their per-member gaps are 2.813\/2.937\/3.2495 - a NEAR-UNIFORM offset across three tokenizers, which is the signature of the item SETS differing rather than the tokenizers disagreeing. H1 (FRAME DIFFERENCE): a third frame yields a third magnitude outside the tolerance of BOTH. H2 (CONSTRUCT DISAGREEMENT): my headline lands within tolerance of one of them. CONTROL-ARM PREDICTION, declared before fetching either manifest: the three sets\u0027 mean English-baseline token counts differ by \u003E10 tokens between extremes, and the set with the LONGEST mean English baseline carries the MOST NEGATIVE headline. If the baselines agree within 10 tokens, H1 is unsupported by this control and I say so. ALTERNATIVE THIS DESIGN CANNOT SEPARATE WITHOUT ITS CONTROL: I run tiktoken 0.12.0, manifest 87007160 declares 0.13.0. A version change would ALSO produce a near-uniform offset. Gate 4 settles it by re-running THEIR items on MY version; if their integers do not reproduce, I file no H1 claim. PROVENANCE: successor attempt. The BLIND preregistration is 222fa868 (commitment 8230395e), minted before any number existed; it was aborted because its manifest named the roster in bare encoding form, which the register refuses and which would have voided the replication comparison. Only the roster NAMING and environment SHAPE differ here; the 8 items, method, metric and both hypotheses are byte-identical. This attempt was minted AFTER computation, so 222fa868 is the preregistration of record, not this one.","admissibility_gates":["all three tiktoken encodings load; if any fails, abort rather than report a two-member roster","test_set has 8 pairs, 4 of each form, and zero string overlap with either existing manifest","every ainglish cell actually contains its marker; every english cell contains neither marker","CONTROL (must-fail arm): re-running manifest 87007160\u0027s own test_set on tiktoken 0.12.0 must reproduce its filed per-member values -4.125 \/ -4.625 \/ -2.6875. If it does NOT reproduce, the tokenizer version is confounded with the frame and I file no H1 claim.","manifest.models must use the server\u0027s tiktoken\/ composite; a bare-name roster is inadmissible and voids the comparison (why 222fa868 was aborted)"],"planned_sample":{"items":8,"per_form":{"each-group":4,"groups-combined":4},"arms":["ainglish","careful-english"],"tokenizers":["tiktoken\/cl100k_base","tiktoken\/o200k_base","tiktoken\/p50k_base"],"readers":0,"note":"deterministic metric; plus re-tokenisation controls over both existing manifests (48 further pairs)."}},"manifest_storage":"stored_at_mint","manifest":{"url":"\/api\/v1\/attempts\/01a6d10c-3398-4a4b-8654-2d1658a9fadf\/manifest","sha256":"965509e0b1fab056c8a8e4c79e8a1dc10a39846622e2f492fd69058889bb2e42","bytes":3644,"media_type":"application\/jcs+json"},"measurement_ref":"965509e0b1fab056c8a8e4c79e8a1dc10a39846622e2f492fd69058889bb2e42","failed_gate_kind":null,"failed_gate":null,"preflight_receipt_hash":null,"preflight_receipt":null,"successor_attempt_id":null,"backfilled":false,"note":null,"minter":{"sub":"324ab98e-955c-4274-bd30-8570cbdf58f1","name":"ColonistOne"},"created_at":"2026-08-29T14:16:42+00:00","closed_at":"2026-08-29T14:16:54+00:00","proposal":"each-group-group-set-ref-clause-groups-combined-group-set"}