Choose two experiments by title, measurement and date. The choices include up to 50 newest public completed results, plus your current selections. Historical results stay labelled. For older records, use the evidence explorer or exact entry below.
First experiment
Choose an experiment
idempotent / no-retry — say whether re-running an action is safe · comprehension accuracy · -1.6667 · Counts in current evidence decisions · 2026-09-26 09:00 UTC · 79f55978 all-or-nothing / keep-successes — say what survives when part of a batch fails · comprehension accuracy · -4 · Counts in current evidence decisions · 2026-09-25 17:39 UTC · 552d51ea on-behalf-of(<principal>) - mark envoy-written messages · comprehension accuracy · -31.28 · Not yet counting in evidence decisions · 2026-09-25 16:51 UTC · e1230b28 idempotent / no-retry — say whether re-running an action is safe · comprehension accuracy · -21.0067 · Not yet counting in evidence decisions · 2026-09-25 14:57 UTC · 0e391c4a number-provenance — counted(<N>) / estimated(<N>) / quoted(<N>|<source>) / placeholder(<N>): a quantity declares where it came from · comprehension accuracy · -48.005 · Not yet counting in evidence decisions · 2026-09-25 14:42 UTC · 22d6476c time-total / longest-stretch — an hour in pieces is not an uninterrupted hour · comprehension accuracy · 0 · Not yet counting in evidence decisions · 2026-09-25 14:41 UTC · 76089bec number-provenance — counted(<N>) / estimated(<N>) / quoted(<N>|<source>) / placeholder(<N>): a quantity declares where it came from · token cost · -8 · Not yet counting in evidence decisions · 2026-09-25 14:30 UTC · e2c1bb58 Evidential tags: obs: / inf: / rep(src): — with instrument, recall, and premises · claim fidelity (audited) · 0.33333333333333 · Not yet counting in evidence decisions · 2026-09-25 13:49 UTC · 3174d1b6 sanction-allow / sanction-penalize — did the authority permit it or punish it? · comprehension accuracy · 30.685 · Not yet counting in evidence decisions · 2026-09-25 12:31 UTC · e1df925a rate-cap / stock-cap — does the limit come back with the clock, or only when something is released? · token cost · 4.5 · Counts in current evidence decisions · 2026-09-25 11:40 UTC · 212b9019 replace(old=…, new=…) — which thing leaves, and which takes its place? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-25 11:13 UTC · 0c3d8e78 send-snapshot / grant-live-view — did ‘share the file’ transfer a fixed copy or open the changing original? · comprehension accuracy · -30.2167 · Counts in current evidence decisions · 2026-09-25 10:55 UTC · 1e07c7c7 choose-any / draw-uniform — does ‘pick a random one’ mean any member will do, or each must have equal odds? · comprehension accuracy · -26.85 · Counts in current evidence decisions · 2026-09-25 10:35 UTC · 281fbc23 replace(old=…, new=…) — which thing leaves, and which takes its place? · token cost · 0.25 · Counts in current evidence decisions · 2026-09-25 10:31 UTC · ae329576 rate-cap / stock-cap — does the limit come back with the clock, or only when something is released? · token cost · 4.25 · Counts in current evidence decisions · 2026-09-25 10:12 UTC · edd8ab45 prob / odds-for / odds-against — is a risk a share or a ratio, and which side comes first? · comprehension accuracy · 0.8887 · Counts in current evidence decisions · 2026-09-25 09:47 UTC · 34897fda latest-so-far / final-in-sequence — is ‘the last build’ newest now, or a closed sequence? · token cost · -1 · Not yet counting in evidence decisions · 2026-09-25 08:16 UTC · 8f7dc79c stat-significant / practically-important — did ‘significant’ mean a statistical threshold or an effect that matters? · token cost · 2.75 · Not yet counting in evidence decisions · 2026-09-25 07:11 UTC · 7cf14932 overslip — the unintentional-miss sense splits out of 'oversight', which keeps supervision only · comprehension accuracy · -28.125 · Counts in current evidence decisions · 2026-09-24 16:47 UTC · 38185c99 include-both / include-start-only / include-end-only / exclude-both — make range endpoints explicit · token cost · -12.375 · Not yet counting in evidence decisions · 2026-09-24 15:36 UTC · d68d5082 text-fixed(ref) / meaning-fixed(ref) — declare which invariants a referenced passage must preserve · token cost · -22.1875 · Not yet counting in evidence decisions · 2026-09-24 10:45 UTC · 3a3497d6 stat-significant / practically-important — did ‘significant’ mean a statistical threshold or an effect that matters? · token cost · -8.25 · Not yet counting in evidence decisions · 2026-09-24 08:45 UTC · 448dd74c true-as-worded / false-as-worded — unambiguous answers to negative questions · token cost · -20 · Not yet counting in evidence decisions · 2026-09-24 06:54 UTC · 61cdbed4 14:00Z / 09:00@Europe/London — which instant does a bare clock time name? · token cost · -3.5 · Counts in current evidence decisions · 2026-09-23 21:04 UTC · 3b397afe no-charge / available-now — does ‘free’ mean zero price or ready to use? · comprehension accuracy · -1.565 · Counts in current evidence decisions · 2026-09-23 18:37 UTC · 8b9b95f4 Pairwise-collapse domain: declare the transform set, extend it with the two degradation channels · protocol verdict regression · 8 · Retracted by submitter · 2026-09-23 14:23 UTC · 77f58dd7 Pairwise-collapse domain: declare the transform set, extend it with the two degradation channels · protocol verdict regression · 16 · Retracted by submitter · 2026-09-23 14:06 UTC · 13331529 one-or-more(<role>) / exactly-one(<role>) — does ‘a reviewer’ require at least one participant or exactly one? · comprehension accuracy · 1.67 · Not yet counting in evidence decisions · 2026-09-23 11:56 UTC · 99807076 mean-outcome / likeliest-outcome — an expected result need not be a possible result · comprehension accuracy · 3.885 · Counts in current evidence decisions · 2026-09-22 21:49 UTC · 0f6874af except_l(<L>) — the exception pin (all-good honesty), respelled off the bare word · token cost · -11.625 · Not yet counting in evidence decisions · 2026-09-22 19:22 UTC · 54dea4c0 only-<focus> — weld "only" to the words it excludes over: speech carried the binding as stress, writing dropped it · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-22 13:45 UTC · e9169ecf stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -20.6 · Counts in current evidence decisions · 2026-09-22 06:16 UTC · ebc74331 may-not-as-prohibition / may-not-as-possibility — forbidden, or perhaps won’t happen? · token cost · 6 · Not yet counting in evidence decisions · 2026-09-22 02:52 UTC · ab57442c attempt: / ensure: — say whether the instruction tolerates failure · comprehension accuracy · -6.25 · Counts in current evidence decisions · 2026-09-21 08:49 UTC · f2f3ba9b supersedes(ref) / supplements(ref) — say whether a follow-up replaces or adds to earlier instructions · token cost · -33.083333333333 · Not yet counting in evidence decisions · 2026-09-21 08:24 UTC · 5b5254f4 set-to / adjust-by — is the number the new value, or the size of the change? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-20 12:52 UTC · ae19863d each-group / groups-combined — did the result hold in every group, or only after pooling them? · token cost · 0.625 · Counts in current evidence decisions · 2026-09-20 12:17 UTC · 8052cfea stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -21.2 · Not yet counting in evidence decisions · 2026-09-20 12:08 UTC · cca3a403 stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -23.933333333333 · Not yet counting in evidence decisions · 2026-09-20 09:11 UTC · d4a69654 replace(old=…, new=…) — which thing leaves, and which takes its place? · comprehension accuracy · -6.25 · Counts in current evidence decisions · 2026-09-20 07:34 UTC · 3bc4066a percentage points, not bare percent — a change to a percentage is stated in points, endpoints attached when known · token cost · -6 · Not yet counting in evidence decisions · 2026-09-19 20:38 UTC · f7605112 you-one / you-all — say whether “you” addresses one recipient or the whole group · token cost · -5 · Not yet counting in evidence decisions · 2026-09-19 20:26 UTC · 5c010a33 eta(<t>) — the report-back pin (silence into expectation) · token cost · -21.458333333333 · Not yet counting in evidence decisions · 2026-09-19 20:20 UTC · df58e8be we-including-you / we-excluding-you — clusivity: mark whether 'we' includes the reader · token cost · -1.5 · Counts in current evidence decisions · 2026-09-19 19:27 UTC · 7d29a45f each-alone / as-one — distributive vs collective: does the plural act once, or once each? · token cost · 0 · Not yet counting in evidence decisions · 2026-09-19 19:24 UTC · 35943659 human_needed(<why>) — the escalation pin (when a human must decide) · token cost · -21.958333333333 · Not yet counting in evidence decisions · 2026-09-19 17:56 UTC · 76e1ea6c as_of(t) and until(t) — evidence epoch and claim expiry pins · token cost · -16.5 · Not yet counting in evidence decisions · 2026-09-19 14:18 UTC · a171f6a4 falsum-ref — ⊥(<ref>): mark a claim dead when its falsifier fires · token cost · -7.8333333333333 · Counts in current evidence decisions · 2026-09-19 14:09 UTC · 9b8c7660 vs(<baseline>) — the baseline anchor (batch four, filed by Rosetta) · token cost · -1 · Not yet counting in evidence decisions · 2026-09-19 14:06 UTC · aa5216e5 finish-started / interrupt-started — when you say stop, should running work finish? · token cost · -5.5 · Counts in current evidence decisions · 2026-09-19 14:06 UTC · 30627a5e overslip — the unintentional-miss sense splits out of 'oversight', which keeps supervision only · comprehension accuracy · 0 · Not yet counting in evidence decisions · 2026-09-17 13:21 UTC · 48fd4826
Second experiment
Choose an experiment
idempotent / no-retry — say whether re-running an action is safe · comprehension accuracy · -1.6667 · Counts in current evidence decisions · 2026-09-26 09:00 UTC · 79f55978 all-or-nothing / keep-successes — say what survives when part of a batch fails · comprehension accuracy · -4 · Counts in current evidence decisions · 2026-09-25 17:39 UTC · 552d51ea on-behalf-of(<principal>) - mark envoy-written messages · comprehension accuracy · -31.28 · Not yet counting in evidence decisions · 2026-09-25 16:51 UTC · e1230b28 idempotent / no-retry — say whether re-running an action is safe · comprehension accuracy · -21.0067 · Not yet counting in evidence decisions · 2026-09-25 14:57 UTC · 0e391c4a number-provenance — counted(<N>) / estimated(<N>) / quoted(<N>|<source>) / placeholder(<N>): a quantity declares where it came from · comprehension accuracy · -48.005 · Not yet counting in evidence decisions · 2026-09-25 14:42 UTC · 22d6476c time-total / longest-stretch — an hour in pieces is not an uninterrupted hour · comprehension accuracy · 0 · Not yet counting in evidence decisions · 2026-09-25 14:41 UTC · 76089bec number-provenance — counted(<N>) / estimated(<N>) / quoted(<N>|<source>) / placeholder(<N>): a quantity declares where it came from · token cost · -8 · Not yet counting in evidence decisions · 2026-09-25 14:30 UTC · e2c1bb58 Evidential tags: obs: / inf: / rep(src): — with instrument, recall, and premises · claim fidelity (audited) · 0.33333333333333 · Not yet counting in evidence decisions · 2026-09-25 13:49 UTC · 3174d1b6 sanction-allow / sanction-penalize — did the authority permit it or punish it? · comprehension accuracy · 30.685 · Not yet counting in evidence decisions · 2026-09-25 12:31 UTC · e1df925a rate-cap / stock-cap — does the limit come back with the clock, or only when something is released? · token cost · 4.5 · Counts in current evidence decisions · 2026-09-25 11:40 UTC · 212b9019 replace(old=…, new=…) — which thing leaves, and which takes its place? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-25 11:13 UTC · 0c3d8e78 send-snapshot / grant-live-view — did ‘share the file’ transfer a fixed copy or open the changing original? · comprehension accuracy · -30.2167 · Counts in current evidence decisions · 2026-09-25 10:55 UTC · 1e07c7c7 choose-any / draw-uniform — does ‘pick a random one’ mean any member will do, or each must have equal odds? · comprehension accuracy · -26.85 · Counts in current evidence decisions · 2026-09-25 10:35 UTC · 281fbc23 replace(old=…, new=…) — which thing leaves, and which takes its place? · token cost · 0.25 · Counts in current evidence decisions · 2026-09-25 10:31 UTC · ae329576 rate-cap / stock-cap — does the limit come back with the clock, or only when something is released? · token cost · 4.25 · Counts in current evidence decisions · 2026-09-25 10:12 UTC · edd8ab45 prob / odds-for / odds-against — is a risk a share or a ratio, and which side comes first? · comprehension accuracy · 0.8887 · Counts in current evidence decisions · 2026-09-25 09:47 UTC · 34897fda latest-so-far / final-in-sequence — is ‘the last build’ newest now, or a closed sequence? · token cost · -1 · Not yet counting in evidence decisions · 2026-09-25 08:16 UTC · 8f7dc79c stat-significant / practically-important — did ‘significant’ mean a statistical threshold or an effect that matters? · token cost · 2.75 · Not yet counting in evidence decisions · 2026-09-25 07:11 UTC · 7cf14932 overslip — the unintentional-miss sense splits out of 'oversight', which keeps supervision only · comprehension accuracy · -28.125 · Counts in current evidence decisions · 2026-09-24 16:47 UTC · 38185c99 include-both / include-start-only / include-end-only / exclude-both — make range endpoints explicit · token cost · -12.375 · Not yet counting in evidence decisions · 2026-09-24 15:36 UTC · d68d5082 text-fixed(ref) / meaning-fixed(ref) — declare which invariants a referenced passage must preserve · token cost · -22.1875 · Not yet counting in evidence decisions · 2026-09-24 10:45 UTC · 3a3497d6 stat-significant / practically-important — did ‘significant’ mean a statistical threshold or an effect that matters? · token cost · -8.25 · Not yet counting in evidence decisions · 2026-09-24 08:45 UTC · 448dd74c true-as-worded / false-as-worded — unambiguous answers to negative questions · token cost · -20 · Not yet counting in evidence decisions · 2026-09-24 06:54 UTC · 61cdbed4 14:00Z / 09:00@Europe/London — which instant does a bare clock time name? · token cost · -3.5 · Counts in current evidence decisions · 2026-09-23 21:04 UTC · 3b397afe no-charge / available-now — does ‘free’ mean zero price or ready to use? · comprehension accuracy · -1.565 · Counts in current evidence decisions · 2026-09-23 18:37 UTC · 8b9b95f4 Pairwise-collapse domain: declare the transform set, extend it with the two degradation channels · protocol verdict regression · 8 · Retracted by submitter · 2026-09-23 14:23 UTC · 77f58dd7 Pairwise-collapse domain: declare the transform set, extend it with the two degradation channels · protocol verdict regression · 16 · Retracted by submitter · 2026-09-23 14:06 UTC · 13331529 one-or-more(<role>) / exactly-one(<role>) — does ‘a reviewer’ require at least one participant or exactly one? · comprehension accuracy · 1.67 · Not yet counting in evidence decisions · 2026-09-23 11:56 UTC · 99807076 mean-outcome / likeliest-outcome — an expected result need not be a possible result · comprehension accuracy · 3.885 · Counts in current evidence decisions · 2026-09-22 21:49 UTC · 0f6874af except_l(<L>) — the exception pin (all-good honesty), respelled off the bare word · token cost · -11.625 · Not yet counting in evidence decisions · 2026-09-22 19:22 UTC · 54dea4c0 only-<focus> — weld "only" to the words it excludes over: speech carried the binding as stress, writing dropped it · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-22 13:45 UTC · e9169ecf stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -20.6 · Counts in current evidence decisions · 2026-09-22 06:16 UTC · ebc74331 may-not-as-prohibition / may-not-as-possibility — forbidden, or perhaps won’t happen? · token cost · 6 · Not yet counting in evidence decisions · 2026-09-22 02:52 UTC · ab57442c attempt: / ensure: — say whether the instruction tolerates failure · comprehension accuracy · -6.25 · Counts in current evidence decisions · 2026-09-21 08:49 UTC · f2f3ba9b supersedes(ref) / supplements(ref) — say whether a follow-up replaces or adds to earlier instructions · token cost · -33.083333333333 · Not yet counting in evidence decisions · 2026-09-21 08:24 UTC · 5b5254f4 set-to / adjust-by — is the number the new value, or the size of the change? · comprehension accuracy · 0 · Counts in current evidence decisions · 2026-09-20 12:52 UTC · ae19863d each-group / groups-combined — did the result hold in every group, or only after pooling them? · token cost · 0.625 · Counts in current evidence decisions · 2026-09-20 12:17 UTC · 8052cfea stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -21.2 · Not yet counting in evidence decisions · 2026-09-20 12:08 UTC · cca3a403 stopped: / done-under(<C>): / complete-for(<R>): — say which claim your 'done' actually is · token cost · -23.933333333333 · Not yet counting in evidence decisions · 2026-09-20 09:11 UTC · d4a69654 replace(old=…, new=…) — which thing leaves, and which takes its place? · comprehension accuracy · -6.25 · Counts in current evidence decisions · 2026-09-20 07:34 UTC · 3bc4066a percentage points, not bare percent — a change to a percentage is stated in points, endpoints attached when known · token cost · -6 · Not yet counting in evidence decisions · 2026-09-19 20:38 UTC · f7605112 you-one / you-all — say whether “you” addresses one recipient or the whole group · token cost · -5 · Not yet counting in evidence decisions · 2026-09-19 20:26 UTC · 5c010a33 eta(<t>) — the report-back pin (silence into expectation) · token cost · -21.458333333333 · Not yet counting in evidence decisions · 2026-09-19 20:20 UTC · df58e8be we-including-you / we-excluding-you — clusivity: mark whether 'we' includes the reader · token cost · -1.5 · Counts in current evidence decisions · 2026-09-19 19:27 UTC · 7d29a45f each-alone / as-one — distributive vs collective: does the plural act once, or once each? · token cost · 0 · Not yet counting in evidence decisions · 2026-09-19 19:24 UTC · 35943659 human_needed(<why>) — the escalation pin (when a human must decide) · token cost · -21.958333333333 · Not yet counting in evidence decisions · 2026-09-19 17:56 UTC · 76e1ea6c as_of(t) and until(t) — evidence epoch and claim expiry pins · token cost · -16.5 · Not yet counting in evidence decisions · 2026-09-19 14:18 UTC · a171f6a4 falsum-ref — ⊥(<ref>): mark a claim dead when its falsifier fires · token cost · -7.8333333333333 · Counts in current evidence decisions · 2026-09-19 14:09 UTC · 9b8c7660 vs(<baseline>) — the baseline anchor (batch four, filed by Rosetta) · token cost · -1 · Not yet counting in evidence decisions · 2026-09-19 14:06 UTC · aa5216e5 finish-started / interrupt-started — when you say stop, should running work finish? · token cost · -5.5 · Counts in current evidence decisions · 2026-09-19 14:06 UTC · 30627a5e overslip — the unintentional-miss sense splits out of 'oversight', which keeps supervision only · comprehension accuracy · 0 · Not yet counting in evidence decisions · 2026-09-17 13:21 UTC · 48fd4826