Evidence
Every result here is one of ours that did not hold up.
Any vendor can show you a benchmark they won. The useful question is what happens when their own result fails — whether anyone notices, and whether they tell you. This is our answer, dated and unedited.
The ledger
Results we retracted, and what we changed.
| Date | What happened | Status |
|---|---|---|
| 2026-07-02 | 19 of 23 “novel” chemotypes were PAINS A re-derivation of our own lead list found that the great majority of what we had been calling novel chemistry consisted of pan-assay interference compounds — molecules that look active against everything because of how they behave in an assay, not because they bind. The lead re-ranker now flags them and deduplicates scaffolds before anything is called a hit. | Fixed |
| 2026-07-07 | CYP inhibition AUROC 0.94–0.96 → roughly chance Our headline CYP numbers were measured in-distribution. On genuinely novel chemistry — the only setting that matters for a discovery programme — performance fell to approximately 0.47–0.54, which is to say no better than guessing. We corrected the claim everywhere it appeared rather than leaving it on the pages that were converting. | Retracted |
| 2026-08-06 | Every “ensemble = 3” was silently a single sample A configuration path meant that runs recorded as three-sample ensembles had in fact been scored once. Every uncertainty estimate derived from that setting was meaningless — not wrong in a way that showed up as an error, which is what made it dangerous. Fixed, and the affected runs re-scored. | Fixed |
| 2026-08-16 | A pooled knockout PASS at n=89, retracted One of our strongest validation results turned out to rest on a contaminated panel: class labels had been assigned from a naming convention rather than from measured evidence. We retracted the PASS and changed the rule so a compound’s class can never again be inferred from what it is called. | Retracted |
| 2026-08-21 | This page published a calibration number we had already retracted For six days this page showed PARP1 calibration as PASS, ρ = 0.88, n = 9, and the chart below claimed two targets cleared the ρ ≥ 0.70 gate. Our own rerun on 2026-08-14 had already measured that same correlation at ρ = 0.407, CI [0.048, 0.660], n = 24 — a fail against the same bar. BCL-XL’s 0.717 was likewise superseded by 0.627. The card generator preferred a promoted snapshot over the later, better-powered measurement, so a stale number outlived its own correction. The precedence rule now takes the most recent measurement, the cards and chart are regenerated, and no target clears the gate. | Retracted |
| Today | 0 of our 7 targets clear our own gates Not one target on our panel is currently validated enough to emit ranked leads, and none is currently usable even as a binary hit-gate. We could change that number this afternoon by loosening the gates. The gates are the product, so we have not. | Standing |
This ledger is maintained, not curated. Entries are added when we find them and are never removed once published.
A real target card
PARP1 (PDB 7AAD) — verdict BLOCKED
This is not a mock-up. It is the card our harness generated on 2026-08-21, reproduced in full — 5 of 9 gates passing, and the target consequently not cleared to emit leads.
| Gate | Verdict | Evidence |
|---|---|---|
| Receptor provenanceIs the structure actually this protein? | PASS | pocket verified |
| CalibrationDo predicted values track measured ones? | FAIL | ρ = 0.407 · n = 24 |
| DiscriminationDoes it beat a fingerprint baseline? | PASS | AUC = 0.9886 · vs ECFP4 baseline 0.8685 · actives = 24, decoys = 22 |
| Blinded holdoutHeld-out actives against a pre-fixed threshold | PASS | holdout validated |
| Negative controlFalse-positive rate on known non-binders | UNKNOWN | FPR not yet measured |
| EnrichmentAre real actives concentrated at the top? | PASS | AUC = 0.87 |
| ReplicationDoes the result survive a re-run? | PASS | replication pack bundled |
| Calibration (affinity head)Continuous potency, not just binary | UNKNOWN | ρ = 0.6009 · n = 24 · the 95% interval does not separate this correlation from the 0.7 bar |
| Negative control (affinity head)Same, for the affinity head | UNKNOWN | FPR not yet measured |
Look at the discrimination row. An AUC of 1.0 is the kind of number that gets put on a slide — and it is marked UNKNOWN, because it was computed against a single decoy. A gate that reports insufficient power instead of a perfect score is the whole reason to have gates.
The results, plotted
Every figure is drawn from the harness output.
These are rendered directly from data/reports/target_cards/*.json by a
script, not drawn by hand, so a figure cannot quietly disagree with the data behind
it. Regression tests compare the numbers on each chart against that JSON.
Case study
PARP1: a strong result, reported with its weak control.
This is the closest thing we have to a clean win, and it is more useful to you as an illustration of how we report one.
Discrimination AUC (95% CI 0.962–1.000, n=46, 2,000-sample bootstrap), against a plain ECFP4 fingerprint baseline of 0.8685 — so the model is genuinely adding something beyond chemical similarity.
Held-out potent actives clearing a threshold that was fixed before they were scored. A blinded test, not a retrospective fit.
The pocket-knockout control. Discrimination dropped from 0.9886 to 0.6420 when the binding pocket was disrupted — suggestive, but underpowered at this n, so the confidence interval still spans both collapse and survival.
What we did not claim
Two natural products cleared the recalibrated threshold in that campaign, and both failed on margin — the gap sat inside the replicate standard deviation. That is not a lead, so we did not report one. The honest output of a well-run campaign is often “nothing here yet.”
The target itself remains BLOCKED on our current panel: the discrimination and negative-control gates read UNKNOWN pending a re-run against the present panel configuration, which is why it is not emitting leads despite the numbers above. Measurements and gate status are different things, and we do not let the better one stand in for the other.
The seven gates
Receptor provenance · calibration · discrimination against a fingerprint baseline · negative-control false-positive rate · enrichment · blinded holdout · replication.
A target must clear all seven before it may emit a single ranked lead. All of it is public — the bars, the null models, the abstention cut-offs, the pocket-selection procedure. The detail is on the methods page. We hold nothing back, because the method was never the moat: the hard part is running these against your own result and publishing the answer when it comes back inconclusive.
This is the standard we would hold your data to.
If that sounds like what your result needs, thirty minutes is enough to find out.