Finding 6 · 2026-08-16
Scaffold-disjoint is not similarity-disjoint, and ChEMBL is analog-dense enough to make it expensive to fix
A Bemis-Murcko scaffold-disjoint 7AAD split shared zero scaffolds between train and test, yet 6 of 20 test compounds sat within Tanimoto 0.80 of a training compound (max 0.936). Enforcing true similarity disjointness evicts most of the test set — and the eviction rate gets worse, not better, as the benchmark scales up.
What we expected
A Bemis-Murcko scaffold-disjoint train/test split should mean the test set is not chemically similar to training data — the standard basis for claiming a model “generalizes.”
What happened
controls.scaffold_leakage_check failed the team's own 7AAD split: “zero shared Bemis-Murcko scaffolds, yet 6 of 20 test compounds sat within Tanimoto 0.80 of a train compound (max 0.936), and 70% within 0.60.” A similarity_split() was added (scaffold-disjoint and similarity-capped), and enforcing it evicts most of the test set:
| target | scaffold-only n_test | + sim ≤0.60 n_test | evicted |
|---|---|---|---|
| 8D0M | 53 | 7 | 87% |
| 4L7B | 52 | 20 | 62% |
| 7AAD | 52 | 13 | 75% |
Scaling the benchmark up does not fix it — eviction rises with size, because a larger training set gives every test compound more chances at a near-neighbour:
- 7AAD: 12 test (73% evicted) @ n=150 → 35 test (85% evicted) @ n=800
- 4JSV: 14 test (69% evicted) @ n=150 → 26 test (89% evicted) @ n=800
Why it happened
“ChEMBL is built from analog series; a target's compounds cluster densely by construction.” This is described as the measured, target-specific version of the general literature finding that scaffold splits overestimate generalization (Guo et al. 2024): “85–89% of a scaffold-disjoint test set is still within Tanimoto 0.60 of training data.”
What changed
Similarity leakage was measured to inflate the ligand-only performance floor by roughly 0.03–0.07 AUC — real but modest, and this cuts against the project's own interest (it makes the memorisation bar structure models must clear look higher than it actually is). Policy: quote both splits, and never quote scaffold-only alone.
“Enforcing similarity disjointness is brutally expensive on this data... This is a property of medicinal-chemistry data, not of our pipeline.”
Lab notebook, lines 5678, 5684.
This is the standard we would hold your data to.
If that sounds like what your result needs, thirty minutes is enough to find out.