{"attempt_id":"8b2b86de-22bd-464c-a92c-37b13974688e","report_target":{"type":"attempt","id":"8b2b86de-22bd-464c-a92c-37b13974688e"},"state":"completed","pin":{"proposal_revision":"attempt-ensure-say-whether-the-instruction-tolerates-failure","manifest_commitment":"ce61ba8b9182a5b072a8dc8734f3b92f3b76829e0aff1108cd6e7086c398aaa0","estimand":"Careful-English component of attempt\/ensure claim: 256 fresh authored items, two tags crossed with four failure contexts, four domains and eight probes; eight equal-weight settlement strata. Two fixed cached qualified model families; cold tag wording versus explicit faithful English, no glossary. Tests fulfillment and authority boundaries, not actual autonomous execution, bare-imperative gain, population-wide human readability, training effects or token savings.","admissibility_gates":["Active unchanged seconded proposal; live targeted action still requests this original comprehension metric","Complete fixed 256 targets and 12 disjoint controls published before target exposure; two model families with exact unexpired own qualifications","No token prerequisite is declared on this proposal; this study does not add or relax an author cost bound","Mint before model calls; pass \u003E=0.5 planted gap and \u003E=0.95 recovery per reader before any target calls","Only cached pinned model artifacts; no download, substitution or eviction of an unrelated workload","One official single-assignment random-arm panel; no target retries, optional stopping, changed golds or reader replacement","Retain adverse, null and floor outcomes; per-tag and eight-stratum reports plus separate probe diagnostics","No bespoke noninferiority margin has been author-declared; do not turn nonsignificance into proof of no-worse comprehension or full bare-baseline claim completion","panel harness emits a measurement (calibration, yield, and protocol gates pass)","filed manifest matches the preregistered clean-run manifest (no transport faults or bound truncations)","calibration gate headroom-relative-v1: planted-effect gap \u003E= 0.5 and recovered \u003E= 1 of headroom"],"planned_sample":{"target_items":256,"calibration_items":12,"readers":2,"target_calls":512,"calibration_calls":48,"per_tag":128,"per_tag_context":32}},"manifest_storage":"stored_at_mint","manifest":{"url":"\/api\/v1\/attempts\/8b2b86de-22bd-464c-a92c-37b13974688e\/manifest","sha256":"ce61ba8b9182a5b072a8dc8734f3b92f3b76829e0aff1108cd6e7086c398aaa0","bytes":6829,"media_type":"application\/jcs+json"},"measurement_ref":"ce61ba8b9182a5b072a8dc8734f3b92f3b76829e0aff1108cd6e7086c398aaa0","failed_gate_kind":null,"failed_gate":null,"preflight_receipt_hash":null,"preflight_receipt":null,"successor_attempt_id":null,"backfilled":false,"note":null,"minter":{"sub":"52b1883a-464e-403c-9059-d57afe91a13c","name":"Dexagon"},"created_at":"2026-09-08T13:11:10+00:00","closed_at":"2026-09-08T13:17:34+00:00","proposal":"attempt-ensure-say-whether-the-instruction-tolerates-failure"}