Evidence

Every result here is one of ours that did not hold up.

Any vendor can show you a benchmark they won. The useful question is what happens when their own result fails — whether anyone notices, and whether they tell you. This is our answer, dated and unedited.

The ledger

Results we retracted, and what we changed.

Date What happened Status
2026-07-02 19 of 23 “novel” chemotypes were PAINS A re-derivation of our own lead list found that the great majority of what we had been calling novel chemistry consisted of pan-assay interference compounds — molecules that look active against everything because of how they behave in an assay, not because they bind. The lead re-ranker now flags them and deduplicates scaffolds before anything is called a hit. Fixed
2026-07-07 CYP inhibition AUROC 0.94–0.96 → roughly chance Our headline CYP numbers were measured in-distribution. On genuinely novel chemistry — the only setting that matters for a discovery programme — performance fell to approximately 0.47–0.54, which is to say no better than guessing. We corrected the claim everywhere it appeared rather than leaving it on the pages that were converting. Retracted
2026-08-06 Every “ensemble = 3” was silently a single sample A configuration path meant that runs recorded as three-sample ensembles had in fact been scored once. Every uncertainty estimate derived from that setting was meaningless — not wrong in a way that showed up as an error, which is what made it dangerous. Fixed, and the affected runs re-scored. Fixed
2026-08-16 A pooled knockout PASS at n=89, retracted One of our strongest validation results turned out to rest on a contaminated panel: class labels had been assigned from a naming convention rather than from measured evidence. We retracted the PASS and changed the rule so a compound’s class can never again be inferred from what it is called. Retracted
2026-08-21 This page published a calibration number we had already retracted For six days this page showed PARP1 calibration as PASS, ρ = 0.88, n = 9, and the chart below claimed two targets cleared the ρ ≥ 0.70 gate. Our own rerun on 2026-08-14 had already measured that same correlation at ρ = 0.407, CI [0.048, 0.660], n = 24 — a fail against the same bar. BCL-XL’s 0.717 was likewise superseded by 0.627. The card generator preferred a promoted snapshot over the later, better-powered measurement, so a stale number outlived its own correction. The precedence rule now takes the most recent measurement, the cards and chart are regenerated, and no target clears the gate. Retracted
Today 0 of our 7 targets clear our own gates Not one target on our panel is currently validated enough to emit ranked leads, and none is currently usable even as a binary hit-gate. We could change that number this afternoon by loosening the gates. The gates are the product, so we have not. Standing

This ledger is maintained, not curated. Entries are added when we find them and are never removed once published.

A real target card

PARP1 (PDB 7AAD) — verdict BLOCKED

This is not a mock-up. It is the card our harness generated on 2026-08-21, reproduced in full — 5 of 9 gates passing, and the target consequently not cleared to emit leads.

GateVerdictEvidence
Receptor provenanceIs the structure actually this protein? PASS pocket verified
CalibrationDo predicted values track measured ones? FAIL ρ = 0.407 · n = 24
DiscriminationDoes it beat a fingerprint baseline? PASS AUC = 0.9886 · vs ECFP4 baseline 0.8685 · actives = 24, decoys = 22
Blinded holdoutHeld-out actives against a pre-fixed threshold PASS holdout validated
Negative controlFalse-positive rate on known non-binders UNKNOWN FPR not yet measured
EnrichmentAre real actives concentrated at the top? PASS AUC = 0.87
ReplicationDoes the result survive a re-run? PASS replication pack bundled
Calibration (affinity head)Continuous potency, not just binary UNKNOWN ρ = 0.6009 · n = 24 · the 95% interval does not separate this correlation from the 0.7 bar
Negative control (affinity head)Same, for the affinity head UNKNOWN FPR not yet measured

Look at the discrimination row. An AUC of 1.0 is the kind of number that gets put on a slide — and it is marked UNKNOWN, because it was computed against a single decoy. A gate that reports insufficient power instead of a perfect score is the whole reason to have gates.

The results, plotted

Every figure is drawn from the harness output.

These are rendered directly from data/reports/target_cards/*.json by a script, not drawn by hand, so a figure cannot quietly disagree with the data behind it. Regression tests compare the numbers on each chart against that JSON.

The whole panel, honestly
Matrix of 7 targets against 8 gates. Every target's overall verdict is BLOCKED. Many individual gates read PASS, several read UNKNOWN, and the summary line states 0 of 7 GREEN and 0 of 7 screening-eligible.
Seven targets, eight gates. Nothing is green. Several targets pass most of their gates and are still blocked, because a single UNKNOWN is enough — an unmeasured negative control is not a passed one.
Does the model beat a fingerprint?
Bar chart comparing model AUC-ROC against an ECFP4 fingerprint baseline for four targets. 7AAD shows AUC 1.000 hatched and labelled not meaningful because it rests on 8 actives and 1 decoy. 4I5I shows the model at 0.786 below its baseline of 0.829.
The most expensive model in the world is worth nothing if a chemical fingerprint keeps pace. 4I5I loses to its own baseline — 0.786 against 0.829. And 7AAD’s perfect 1.000 is hatched out, because it was computed against a single decoy and means nothing at all.
The knockout, and why it settles nothing yet
Before and after chart of PARP1 pocket-knockout. AUC falls from 0.9886 to 0.6420 with a drop 95% confidence interval of 0.1905 to 0.5189 at n=46. The derived after-arm range straddles the chance line at 0.5 and the verdict is marked inconclusive.
Disrupt the binding pocket and the score should collapse toward chance. It fell — but at n=46 the interval still straddles the 0.5 line, so this is inconclusive, not a result. The obvious move would be to report the drop and stop talking about the interval.
Calibration against the gate
Spearman rho per target with sample sizes, against a dashed reference line at the rho equals 0.70 gate threshold. Values range from -0.371 for 4CFE to 0.627 for 2YXJ. No target clears the threshold. Two targets are drawn as dashed placeholders labelled unknown, not measured.
Not one target clears ρ ≥ 0.70. 4CFE is negative at −0.371 — its predictions run backwards against measured potency. Two targets are drawn as unmeasured rather than as zero, because a measurement nobody took is not a score of nought.

Case study

PARP1: a strong result, reported with its weak control.

This is the closest thing we have to a clean win, and it is more useful to you as an illustration of how we report one.

0.9886

Discrimination AUC (95% CI 0.962–1.000, n=46, 2,000-sample bootstrap), against a plain ECFP4 fingerprint baseline of 0.8685 — so the model is genuinely adding something beyond chemical similarity.

12 / 12

Held-out potent actives clearing a threshold that was fixed before they were scored. A blinded test, not a retrospective fit.

Inconclusive

The pocket-knockout control. Discrimination dropped from 0.9886 to 0.6420 when the binding pocket was disrupted — suggestive, but underpowered at this n, so the confidence interval still spans both collapse and survival.

What we did not claim

Two natural products cleared the recalibrated threshold in that campaign, and both failed on margin — the gap sat inside the replicate standard deviation. That is not a lead, so we did not report one. The honest output of a well-run campaign is often “nothing here yet.”

The target itself remains BLOCKED on our current panel: the discrimination and negative-control gates read UNKNOWN pending a re-run against the present panel configuration, which is why it is not emitting leads despite the numbers above. Measurements and gate status are different things, and we do not let the better one stand in for the other.

The seven gates

Receptor provenance · calibration · discrimination against a fingerprint baseline · negative-control false-positive rate · enrichment · blinded holdout · replication.

A target must clear all seven before it may emit a single ranked lead. All of it is public — the bars, the null models, the abstention cut-offs, the pocket-selection procedure. The detail is on the methods page. We hold nothing back, because the method was never the moat: the hard part is running these against your own result and publishing the answer when it comes back inconclusive.

This is the standard we would hold your data to.

If that sounds like what your result needs, thirty minutes is enough to find out.

Book a 30-min call