Finding 4 · 2026-07-03

Three calibration reference sets were built from the wrong protein's ligands

Three of seven *_actives.csv calibration files were pulling ChEMBL ligands for the wrong protein entirely — hardcoded under a comment that falsely claimed the IDs were “Verified” — which turned out to explain two calibration runs that had looked like modelling failures.

What we expected

The chembl_id in each *_actives.csv calibration file corresponds to the target the PDB structure represents.

What happened, resolved via the live ChEMBL API

  • 8D0M/CD38 → CHEMBL3227 = GRM5 (metabotropic glutamate receptor 5). Correct CD38 = CHEMBL4660.
  • 4CFE/AMPK → CHEMBL4630 = Chk1 (+ CHEMBL2111425 integrin). Correct AMPK = CHEMBL4045 + CHEMBL2116.
  • 4L7B/KEAP1 → CHEMBL5619 = APEX1 (DNA-repair nuclease). Correct KEAP1 = CHEMBL2069156.

7AAD/PARP1, 4I5I/SIRT1, 2YXJ/BCL-XL and 4JSV/mTOR were checked and confirmed correct.

Why it happened

Hardcoded wrong ChEMBL IDs in build_control_sets.py:TARGETS, under a comment that falsely claimed “Verified.”

What changed

This “explains the mystery calibration failures (4L7B ρ=−0.39 ‘anti-predictive,’ 4CFE ρ=0.34) — wrong actives.” The previously reported “CD38 CLEARED / 9 qualifying leads” result was retracted the same day it was reported.

A new tools/target_provenance.py now maps PDB→protein→UniProt→ChEMBL and verifies against live ChEMBL and SIFTS before any refset is built, with a hard gate. The three wrong-target refsets were quarantined, correct actives re-pulled (CD38→345, AMPK→233, KEAP1→210), and the affected target cards marked INVALID_PENDING_RECALIBRATION.

“Root cause: hardcoded wrong IDs in build_control_sets.py:TARGETS (under a comment falsely claiming ‘Verified’).”

Lab notebook, line 3070.

This is the standard we would hold your data to.

If that sounds like what your result needs, thirty minutes is enough to find out.

Book a 30-min call