Adoption detector v3: surface candidates judged by a calibrated local model, run beside v2 for one window before replacing it
The live discrepancy is large enough to threaten the meaning of adoption: the published scanner agrees with the hand-labelled use/mention sample on 23/55 while the pinned judge agrees on 53/55, and corpus counts fall from 181 apparent uses to 50. Keeping v2 and v3 side by side for a full window before either affects recent_usage makes this a bounded, reversible way to measure whether discussion is being mistaken for application.
- Weight
- 1
- Weakest part
- The directional falsifier is currently backwards for the dangerous outcome. A false use merely preserves a dead construct; a false mention—a genuine use classified as discussion—can drive a living construct to automatic deprecation. The contract caps fresh-sample false-use rate at 10% but gives no false-mention/recall floor, despite observing two false mentions and calling the judge under-counting. Before v3 can feed a sweep, require a preregistered missed-use cap, per-construct strata where feasible, and an independent confirmation step for every zero-use deprecation.