Finding 1 · 2026-08-16 (later)
The knockout's own control was inverted: three unrelated failure modes all certified PASS
A knockout control meant to tell three failure modes apart — real pocket-dependence, a broken run, and pure ligand memorisation — instead passed all three, because it checked whether the mutant-arm AUC fell under a hardcoded 0.70 rather than comparing it to chance (0.5).
What we expected
A pocket-mutation knockout control should distinguish “the model's score depends on the protein” (mutant AUC collapses toward chance) from “the model is memorising the ligand” (mutant AUC survives above chance).
What happened
pocket_knockout_verdict tested only whether the mutant-arm AUC fell below 0.70, and never compared it to chance (0.5). Three completely different outcomes all passed:
| mutant arm | what it actually means | old verdict |
|---|---|---|
| AUC 0.586, CI [0.529, 0.640] | residual signal survived — partial memorisation | PASS |
| every score identical at 0.5 | failed fold / clamped output — measures nothing | PASS |
| AUC 0.000, perfectly inverted | ranking fully retained, sign flipped | PASS |
Why it happened
The check was written against a hardcoded threshold (0.70) that “looks like it is testing the thing” — pocket-dependence — but the threshold isn't the thing; a control needs to test against the null (chance), not an arbitrary cutoff.
What changed
Collapse is now defined as the after-arm's confidence interval containing 0.5. A second bug was introduced fixing the first — a zero-width CI was used to infer “degenerate,” which misclassified a perfect-memorisation AUC of 1.000 as a “failed fold” — and was caught only by testing the fix rather than trusting it. Commit 23ce351.
“A control whose degenerate cases all resolve to PASS is worse than no control, because it launders a broken run as a successful one.”
Lab notebook, line 6070.
This is the standard we would hold your data to.
If that sounds like what your result needs, thirty minutes is enough to find out.