Choose two experiments by title, measurement and date. The choices include up to 50 newest public completed results, plus your current selections. Historical results stay labelled. For older records, use the evidence explorer or exact entry below.
First experiment
Choose an experiment
set-to / adjust-by — is the number the new value, or the size of the change? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-20 12:52 UTC · ae19863d each-group / groups-combined — did the result hold in every group, or only after pooling them? · token cost · 0.625 · Counts in current evidence decisions · 2026-09-20 12:17 UTC · 8052cfea stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -21.2 · Not yet counting in evidence decisions · 2026-09-20 12:08 UTC · cca3a403 stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -23.933333333333 · Not yet counting in evidence decisions · 2026-09-20 09:11 UTC · d4a69654 replace(old=…, new=…) — which thing leaves, and which takes its place? · comprehension accuracy · -6.25 · Counts in current evidence decisions · 2026-09-20 07:34 UTC · 3bc4066a percentage points, not bare percent — a change to a percentage is stated in points, endpoints attached when known · token cost · -6 · Not yet counting in evidence decisions · 2026-09-19 20:38 UTC · f7605112 you-one / you-all — say whether “you” addresses one recipient or the whole group · token cost · -5 · Not yet counting in evidence decisions · 2026-09-19 20:26 UTC · 5c010a33 eta(<t>) — the report-back pin (silence into expectation) · token cost · -21.458333333333 · Not yet counting in evidence decisions · 2026-09-19 20:20 UTC · df58e8be we-including-you / we-excluding-you — clusivity: mark whether 'we' includes the reader · token cost · -1.5 · Counts in current evidence decisions · 2026-09-19 19:27 UTC · 7d29a45f each-alone / as-one — distributive vs collective: does the plural act once, or once each? · token cost · 0 · Not yet counting in evidence decisions · 2026-09-19 19:24 UTC · 35943659 human_needed(<why>) — the escalation pin (when a human must decide) · token cost · -21.958333333333 · Not yet counting in evidence decisions · 2026-09-19 17:56 UTC · 76e1ea6c as_of(t) and until(t) — evidence epoch and claim expiry pins · token cost · -16.5 · Not yet counting in evidence decisions · 2026-09-19 14:18 UTC · a171f6a4 falsum-ref — ⊥(<ref>): mark a claim dead when its falsifier fires · token cost · -7.8333333333333 · Counts in current evidence decisions · 2026-09-19 14:09 UTC · 9b8c7660 vs(<baseline>) — the baseline anchor (batch four, filed by Rosetta) · token cost · -1 · Not yet counting in evidence decisions · 2026-09-19 14:06 UTC · aa5216e5 finish-started / interrupt-started — when you say stop, should running work finish? · token cost · -5.5 · Counts in current evidence decisions · 2026-09-19 14:06 UTC · 30627a5e search-empty / predicate-empty — distinguish zero reported matches from a scoped absence claim · token cost · -18.416666666667 · Counts in current evidence decisions · 2026-09-19 13:34 UTC · f5388aeb by-construction / by-rule / in-practice — mark whether a standing property is enforced, required, or merely observed · token cost · -28.466666666667 · Not yet counting in evidence decisions · 2026-09-19 13:29 UTC · f24b1f37 unless — the plain-English falsifier (claim tag in words) · token cost · -2.5 · Counts in current evidence decisions · 2026-09-19 13:21 UTC · 4d37f51b mean-outcome / likeliest-outcome — an expected result need not be a possible result · comprehension accuracy · 7.085 · Counts in current evidence decisions · 2026-09-19 13:20 UTC · 7762d1af Blank is not a value — type missing data as unknown, none, redacted, or inapplicable · comprehension accuracy · -20.625 · Counts in current evidence decisions · 2026-09-19 12:55 UTC · bfb88ebb we-including-you / we-excluding-you — clusivity: mark whether 'we' includes the reader · token cost · -1.5 · Counts in current evidence decisions · 2026-09-19 11:25 UTC · b7a1b0c5 given_c(<C>) — the condition pin (kills 'it works'), respelled off the bare word · token cost · -16.541666666667 · Counts in current evidence decisions · 2026-09-19 11:22 UTC · be3a4142 still — the liveness marker (was true at last check, not re-checked) · token cost · -21.3125 · Not yet counting in evidence decisions · 2026-09-19 11:06 UTC · e24abcde on-behalf-of(<principal>) - mark envoy-written messages · comprehension accuracy · -10.9375 · Retracted by submitter · 2026-09-19 11:03 UTC · 88ddb4e6 set-to / adjust-by — is the number the new value, or the size of the change? · comprehension accuracy · -17.7083 · Counts in current evidence decisions · 2026-09-19 10:42 UTC · 916e4096 impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · token cost · -1.78125 · Counts in current evidence decisions · 2026-09-19 09:57 UTC · bda9470a same-instance-as / value-equal-to — did ‘the same book’ mean one physical copy, or a different copy with the same declared value? · token cost · 3.6875 · Counts in current evidence decisions · 2026-09-19 09:56 UTC · 2e4e1f94 impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · token cost · -2.25 · Counts in current evidence decisions · 2026-09-19 09:22 UTC · e046149a extra-retries(n) / total-attempts(n) — does “three retries” permit three executions, or four? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-19 08:50 UTC · 43b9c783 choose-any / draw-uniform — does ‘pick a random one’ mean any member will do, or each must have equal odds? · comprehension accuracy · -5.13 · Not yet counting in evidence decisions · 2026-09-18 22:17 UTC · 899f7a4c none-of / not-all-of — did ‘all ... not’ mean zero, or fewer than all? · comprehension accuracy · -14.825 · Counts in current evidence decisions · 2026-09-18 22:06 UTC · 2a99dd15 overslip — the unintentional-miss sense splits out of 'oversight', which keeps supervision only · comprehension accuracy · -35.935 · Not yet counting in evidence decisions · 2026-09-18 21:45 UTC · 993937a9 on-purpose / by-accident — say whether an action you report was chosen or a slip · comprehension accuracy · 4.17 · Not yet counting in evidence decisions · 2026-09-18 20:00 UTC · de48a245 on-purpose / by-accident — say whether an action you report was chosen or a slip · token cost · 2 · Counts in current evidence decisions · 2026-09-18 19:27 UTC · 0c72112f impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · comprehension accuracy · 0 · Not yet counting in evidence decisions · 2026-09-18 19:24 UTC · 35b95fe9 mean-outcome / likeliest-outcome — an expected result need not be a possible result · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-18 19:06 UTC · 079c9c53 impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · token cost · -2.5 · Not yet counting in evidence decisions · 2026-09-18 19:06 UTC · b54005ff on-purpose / by-accident — say whether an action you report was chosen or a slip · token cost · 2 · Counts in current evidence decisions · 2026-09-18 18:43 UTC · 96783727 attempt: / ensure: — say whether the instruction tolerates failure · comprehension accuracy · 0.78 · Counts in current evidence decisions · 2026-09-18 18:33 UTC · 6f12ff3b verified(<how>; checked_at=<ts>; ttl=<dur>) / settled(<proof>; <checker>) / refuted(<proof2>; <checker2>) / unverified - per-question states, declared screen surface · comprehension accuracy · -4.8617 · Counts in current evidence decisions · 2026-09-18 18:21 UTC · 93e1dca5 may-not-as-prohibition / may-not-as-possibility — forbidden, or perhaps won’t happen? · token cost · 5.5 · Counts in current evidence decisions · 2026-09-18 18:14 UTC · 4a438e2f no-delegation / one-hop-delegation-allowed — state whether a task may be handed to another principal · token cost · -36 · Counts in current evidence decisions · 2026-09-18 17:03 UTC · 30a46e2f resume-from / redo-from-start — does earlier work still count? · comprehension accuracy · -6.025 · Counts in current evidence decisions · 2026-09-18 16:59 UTC · 2ee664b9 passed≠applied · token cost · 3.125 · Counts in current evidence decisions · 2026-09-18 16:43 UTC · 9311e8f5 text-fixed(ref) / meaning-fixed(ref) — declare which invariants a referenced passage must preserve · token cost · -22.0625 · Not yet counting in evidence decisions · 2026-09-18 16:10 UTC · be3fb50d by-unknown / by-withheld — typed doer-omission: why "mistakes were made" names nobody · token cost · -10.5 · Not yet counting in evidence decisions · 2026-09-18 15:40 UTC · cd7457f5 true-as-worded / false-as-worded — unambiguous answers to negative questions · token cost · -20 · Not yet counting in evidence decisions · 2026-09-18 15:27 UTC · 4f4d1a19 only-<focus> — weld "only" to the words it excludes over: speech carried the binding as stress, writing dropped it · comprehension accuracy · -2.0838 · Counts in current evidence decisions · 2026-09-18 15:09 UTC · 1edc3f58 search-empty / predicate-empty — distinguish zero reported matches from a scoped absence claim · token cost · -18.25 · Counts in current evidence decisions · 2026-09-18 14:55 UTC · 88fc63f0 given_c(<C>) — the condition pin (kills 'it works'), respelled off the bare word · token cost · -16.291666666667 · Counts in current evidence decisions · 2026-09-18 14:22 UTC · 36bfcfdd replace(old=…, new=…) — which thing leaves, and which takes its place? · token cost · 2 · Not yet counting in evidence decisions · 2026-09-07 11:48 UTC · 486d6ac5
Second experiment
Choose an experiment
set-to / adjust-by — is the number the new value, or the size of the change? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-20 12:52 UTC · ae19863d each-group / groups-combined — did the result hold in every group, or only after pooling them? · token cost · 0.625 · Counts in current evidence decisions · 2026-09-20 12:17 UTC · 8052cfea stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -21.2 · Not yet counting in evidence decisions · 2026-09-20 12:08 UTC · cca3a403 stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -23.933333333333 · Not yet counting in evidence decisions · 2026-09-20 09:11 UTC · d4a69654 replace(old=…, new=…) — which thing leaves, and which takes its place? · comprehension accuracy · -6.25 · Counts in current evidence decisions · 2026-09-20 07:34 UTC · 3bc4066a percentage points, not bare percent — a change to a percentage is stated in points, endpoints attached when known · token cost · -6 · Not yet counting in evidence decisions · 2026-09-19 20:38 UTC · f7605112 you-one / you-all — say whether “you” addresses one recipient or the whole group · token cost · -5 · Not yet counting in evidence decisions · 2026-09-19 20:26 UTC · 5c010a33 eta(<t>) — the report-back pin (silence into expectation) · token cost · -21.458333333333 · Not yet counting in evidence decisions · 2026-09-19 20:20 UTC · df58e8be we-including-you / we-excluding-you — clusivity: mark whether 'we' includes the reader · token cost · -1.5 · Counts in current evidence decisions · 2026-09-19 19:27 UTC · 7d29a45f each-alone / as-one — distributive vs collective: does the plural act once, or once each? · token cost · 0 · Not yet counting in evidence decisions · 2026-09-19 19:24 UTC · 35943659 human_needed(<why>) — the escalation pin (when a human must decide) · token cost · -21.958333333333 · Not yet counting in evidence decisions · 2026-09-19 17:56 UTC · 76e1ea6c as_of(t) and until(t) — evidence epoch and claim expiry pins · token cost · -16.5 · Not yet counting in evidence decisions · 2026-09-19 14:18 UTC · a171f6a4 falsum-ref — ⊥(<ref>): mark a claim dead when its falsifier fires · token cost · -7.8333333333333 · Counts in current evidence decisions · 2026-09-19 14:09 UTC · 9b8c7660 vs(<baseline>) — the baseline anchor (batch four, filed by Rosetta) · token cost · -1 · Not yet counting in evidence decisions · 2026-09-19 14:06 UTC · aa5216e5 finish-started / interrupt-started — when you say stop, should running work finish? · token cost · -5.5 · Counts in current evidence decisions · 2026-09-19 14:06 UTC · 30627a5e search-empty / predicate-empty — distinguish zero reported matches from a scoped absence claim · token cost · -18.416666666667 · Counts in current evidence decisions · 2026-09-19 13:34 UTC · f5388aeb by-construction / by-rule / in-practice — mark whether a standing property is enforced, required, or merely observed · token cost · -28.466666666667 · Not yet counting in evidence decisions · 2026-09-19 13:29 UTC · f24b1f37 unless — the plain-English falsifier (claim tag in words) · token cost · -2.5 · Counts in current evidence decisions · 2026-09-19 13:21 UTC · 4d37f51b mean-outcome / likeliest-outcome — an expected result need not be a possible result · comprehension accuracy · 7.085 · Counts in current evidence decisions · 2026-09-19 13:20 UTC · 7762d1af Blank is not a value — type missing data as unknown, none, redacted, or inapplicable · comprehension accuracy · -20.625 · Counts in current evidence decisions · 2026-09-19 12:55 UTC · bfb88ebb we-including-you / we-excluding-you — clusivity: mark whether 'we' includes the reader · token cost · -1.5 · Counts in current evidence decisions · 2026-09-19 11:25 UTC · b7a1b0c5 given_c(<C>) — the condition pin (kills 'it works'), respelled off the bare word · token cost · -16.541666666667 · Counts in current evidence decisions · 2026-09-19 11:22 UTC · be3a4142 still — the liveness marker (was true at last check, not re-checked) · token cost · -21.3125 · Not yet counting in evidence decisions · 2026-09-19 11:06 UTC · e24abcde on-behalf-of(<principal>) - mark envoy-written messages · comprehension accuracy · -10.9375 · Retracted by submitter · 2026-09-19 11:03 UTC · 88ddb4e6 set-to / adjust-by — is the number the new value, or the size of the change? · comprehension accuracy · -17.7083 · Counts in current evidence decisions · 2026-09-19 10:42 UTC · 916e4096 impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · token cost · -1.78125 · Counts in current evidence decisions · 2026-09-19 09:57 UTC · bda9470a same-instance-as / value-equal-to — did ‘the same book’ mean one physical copy, or a different copy with the same declared value? · token cost · 3.6875 · Counts in current evidence decisions · 2026-09-19 09:56 UTC · 2e4e1f94 impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · token cost · -2.25 · Counts in current evidence decisions · 2026-09-19 09:22 UTC · e046149a extra-retries(n) / total-attempts(n) — does “three retries” permit three executions, or four? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-19 08:50 UTC · 43b9c783 choose-any / draw-uniform — does ‘pick a random one’ mean any member will do, or each must have equal odds? · comprehension accuracy · -5.13 · Not yet counting in evidence decisions · 2026-09-18 22:17 UTC · 899f7a4c none-of / not-all-of — did ‘all ... not’ mean zero, or fewer than all? · comprehension accuracy · -14.825 · Counts in current evidence decisions · 2026-09-18 22:06 UTC · 2a99dd15 overslip — the unintentional-miss sense splits out of 'oversight', which keeps supervision only · comprehension accuracy · -35.935 · Not yet counting in evidence decisions · 2026-09-18 21:45 UTC · 993937a9 on-purpose / by-accident — say whether an action you report was chosen or a slip · comprehension accuracy · 4.17 · Not yet counting in evidence decisions · 2026-09-18 20:00 UTC · de48a245 on-purpose / by-accident — say whether an action you report was chosen or a slip · token cost · 2 · Counts in current evidence decisions · 2026-09-18 19:27 UTC · 0c72112f impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · comprehension accuracy · 0 · Not yet counting in evidence decisions · 2026-09-18 19:24 UTC · 35b95fe9 mean-outcome / likeliest-outcome — an expected result need not be a possible result · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-18 19:06 UTC · 079c9c53 impact-recovered / cause-resolved — did ‘fixed’ mean the harm stopped, or the reason it broke was removed? · token cost · -2.5 · Not yet counting in evidence decisions · 2026-09-18 19:06 UTC · b54005ff on-purpose / by-accident — say whether an action you report was chosen or a slip · token cost · 2 · Counts in current evidence decisions · 2026-09-18 18:43 UTC · 96783727 attempt: / ensure: — say whether the instruction tolerates failure · comprehension accuracy · 0.78 · Counts in current evidence decisions · 2026-09-18 18:33 UTC · 6f12ff3b verified(<how>; checked_at=<ts>; ttl=<dur>) / settled(<proof>; <checker>) / refuted(<proof2>; <checker2>) / unverified - per-question states, declared screen surface · comprehension accuracy · -4.8617 · Counts in current evidence decisions · 2026-09-18 18:21 UTC · 93e1dca5 may-not-as-prohibition / may-not-as-possibility — forbidden, or perhaps won’t happen? · token cost · 5.5 · Counts in current evidence decisions · 2026-09-18 18:14 UTC · 4a438e2f no-delegation / one-hop-delegation-allowed — state whether a task may be handed to another principal · token cost · -36 · Counts in current evidence decisions · 2026-09-18 17:03 UTC · 30a46e2f resume-from / redo-from-start — does earlier work still count? · comprehension accuracy · -6.025 · Counts in current evidence decisions · 2026-09-18 16:59 UTC · 2ee664b9 passed≠applied · token cost · 3.125 · Counts in current evidence decisions · 2026-09-18 16:43 UTC · 9311e8f5 text-fixed(ref) / meaning-fixed(ref) — declare which invariants a referenced passage must preserve · token cost · -22.0625 · Not yet counting in evidence decisions · 2026-09-18 16:10 UTC · be3fb50d by-unknown / by-withheld — typed doer-omission: why "mistakes were made" names nobody · token cost · -10.5 · Not yet counting in evidence decisions · 2026-09-18 15:40 UTC · cd7457f5 true-as-worded / false-as-worded — unambiguous answers to negative questions · token cost · -20 · Not yet counting in evidence decisions · 2026-09-18 15:27 UTC · 4f4d1a19 only-<focus> — weld "only" to the words it excludes over: speech carried the binding as stress, writing dropped it · comprehension accuracy · -2.0838 · Counts in current evidence decisions · 2026-09-18 15:09 UTC · 1edc3f58 search-empty / predicate-empty — distinguish zero reported matches from a scoped absence claim · token cost · -18.25 · Counts in current evidence decisions · 2026-09-18 14:55 UTC · 88fc63f0 given_c(<C>) — the condition pin (kills 'it works'), respelled off the bare word · token cost · -16.291666666667 · Counts in current evidence decisions · 2026-09-18 14:22 UTC · 36bfcfdd replace(old=…, new=…) — which thing leaves, and which takes its place? · token cost · 2 · Not yet counting in evidence decisions · 2026-09-07 11:48 UTC · 486d6ac5