Findings
Twelve things our own controls caught us getting wrong
A small team running Boltz-2 structure-based screening against a longevity/cancer target panel, repeatedly building controls to catch their own errors, and reporting what those controls found — including several results that overturned the team's own prior conclusions. Ordered by how compelling each finding would be to a skeptical computational chemist, most compelling first.
| Date | Finding |
|---|---|
| 2026-08-16 | Finding 1 — The knockout's own control was inverted: three unrelated failure modes all certified PASS A pocket-knockout control checked only whether the mutant-arm AUC fell below a hardcoded 0.70, never against chance (0.5) — so residual signal, a failed fold, and a perfectly inverted ranking all passed. |
| 2026-08-16 | Finding 2 — The queue silently launched four concurrent copies of one GPU job A hand-maintained list of “GPU job in progress” script names fell out of sync with the queue's actual steps, so the health check read the GPU as idle and started a second, then a third, then a fourth copy of the same job. |
| 2026-08-14 | Finding 3 — The shim silently dropped every pocket constraint it was ever sent, for four months The Boltz-2 shim never translated pocket constraints into the boltz YAML it sent to the model, while its own log line described the constraint being applied. 2,444 real 7AAD scores existed at the moment of the fix, every one unconstrained. |
| 2026-07-03 | Finding 4 — Three calibration reference sets were built from the wrong protein's ligands 8D0M/CD38, 4CFE/AMPK and 4L7B/KEAP1 all pulled ChEMBL actives for a different protein entirely, under a comment that falsely claimed the IDs were “Verified” — explaining two calibration results that had looked like modelling failures. |
| 2026-08-16 | Finding 5 — A PASS was retracted as contamination four minutes after it was reported An extended knockout panel reported PASS, then checking the class labels found a wrong-target reference, natural products of unknown activity, unidentifiable IDs, and one alkaloid counted twelve times. Verified-only re-analysis: INCONCLUSIVE. |
| 2026-08-16 | Finding 6 — Scaffold-disjoint is not similarity-disjoint, and ChEMBL is analog-dense enough to make it expensive to fix A scaffold-disjoint 7AAD split shared zero Bemis-Murcko scaffolds, yet 6 of 20 test compounds sat within Tanimoto 0.80 of a train compound. Enforcing similarity disjointness evicts 62–87% of the test set, and eviction rises with scale. |
| 2026-08-16 | Finding 7 — D3: the powered holdout cleared 12/12 on both targets, and 2YXJ's FAIL flipped to PASS A properly-powered n=12 blinded holdout passed 12/12 on both 7AAD and 2YXJ; 2YXJ's earlier FAIL turned out to be n=3 noise — the third gate this week whose verdict inverted once properly powered. |
| 2026-08-16 | Finding 8 — D4: the PARP1 natural-product hit list came back empty — as pre-registered Re-scoring 60 flagged PARP1 compounds at the newly constrained threshold (0.897) left zero credible natural-product hits; the two candidates that cleared the bar failed on replicate noise, affinity-head disagreement, and a flagged genotoxic scaffold. |
| 2026-08-16 (later) | Finding 9 — Metrics that ranked ties by input order — a bug our own data happened never to trigger EF, BEDROC and PPV broke ties by input position rather than invariantly; a constructed test showed EF@10% swinging from 10.0 to 0.0 on identical data depending on compound order. Boltz-2's own scores happened to be tie-free. |
| 2026-08-16 (later) | Finding 10 — The five audit controls existed but were never called by either driver Five written controls — run_cpu_controls, margin_over_baseline, n_for_robust_rho, similarity_split — had zero callers from either evaluation driver. Once wired in, the instrument condemned the split the driver had been using. |
| 2026-08-16 | Finding 11 — An undocumented hard dependency on a free external API brought down every GPU prediction mid-campaign Every Boltz-2 call fetches a fresh MSA from api.colabfold.com; when that server went unreachable, every prediction failed at a constant 64.6s behind a misleading GPU-architecture error, because the shim never logged boltz's real stderr. |
| 2026-08-16 (later) | Finding 12 — The team had 4CFE exactly backwards, and deprioritised the one target where a win could actually be attributed to structure 4CFE/AMPK's near-zero ligand-only floor had been read as “nothing there” and deprioritised; panel-wide leakage measurement shows 4CFE is actually the lowest-leakage target (16%, vs. 60–82% elsewhere) — the one place a win could be attributed to structure. |
| 2026-09-02 | Finding 13 — The GNINA top-1000 screen came back empty at the validated thresholds, as pre-registered Boltz-2 scoring of the GNINA-ranked top-1000 natural products against both eligible targets (2YXJ, 7AAD) at thresholds 0.658 / 0.784 produced 0 hits above the bar across 465 fresh-scored compounds — the pipeline demonstrably scored, the chemistry did not match the pockets, and the D4 gate held. |
Source: the campaign lab notebook. Every number above is copied verbatim from the source document; nothing has been rounded or reconstructed.
This is the standard we would hold your data to.
If that sounds like what your result needs, thirty minutes is enough to find out.