This is the tool, not the offer. We built ToxScreen for our own screening work and left it free to use. The consulting practice it came out of is here.

Calibrated ADMET and protein-binding diagnostics, with the uncertainty published.

ToxScreen applies an adversarial, multi-gate validation discipline to computational drug-discovery — a target must pass seven independent checks before any lead is emitted. Every number carries its confidence interval and its failure mode.

4 calibrated ADMET endpoints 7 adversarial gates per target Failures published, not hidden
Built on open science, honest about uncertainty CYP AUROC 0.94–0.95 in-distribution Applicability-domain abstention Built on ADMET-AI Zero mock data

Benchmarks

Measured on novel chemistry — and published in full

Every headline here survives an independent, de-leaked test. Where the model can’t stand behind a number, we say so — and the product abstains instead of guessing.

What calibration buys you — hERG on de-leaked novel chemistry
0.5 0.7 0.8 screen 0.9 0.71 0.84 +0.13 AUROC ADMET-AI ToxScreen

On de-leaked novel chemistry — a Karim cross-dataset holdout with every InChIKey overlap removed — the calibrated ToxScreen pipeline (applicability-domain re-weighting + isotonic calibration) lifts hERG discrimination from the raw ADMET-AI baseline 0.71 to 0.84 AUROC, clearing the 0.8 screening bar. Just as important, it fixes the model’s confidence: expected calibration error drops from 0.32 to 0.08.

3-seed mean ± 0.02 SD. hERG is the one endpoint with a genuine novel-chemistry number; the full per-endpoint de-leak — including where we fail — is detailed below.

+0.13 AUROC
hERG lift on novel chemistry
ADMET-AI 0.71 → ToxScreen 0.84, de-leaked (Karim cross-dataset, InChIKey-filtered).
0.32 → 0.08
Calibration error (ECE)
Raw ADMET-AI → calibrated. A stated “0.80” now means roughly a real 0.80.
0.47–0.54
CYP AUROC on novel scaffolds
Independent ChEMBL de-leak (n=701 / 237) — near chance, so ToxScreen abstains rather than guess.
4 / 4
Endpoints de-leak-tested & published
We publish the failing and inconclusive numbers, not only the flattering ones.

Calibration — published in full, including where it fails

hERG (7CN1)
0.69 AUROC
FAIL (random split)
~0.84 calibrated (0.71 baseline) on novel chemistry
CYP3A4 (6MA7)
0.95 AUROC
PASS
in-dist. only — novel-chem AUROC 0.47 (n=701), ~chance
CYP2D6 (4WNT)
0.94 AUROC
PASS
in-dist. only — novel-chem test inconclusive (99% single-class)
CYP2C9 (1OG5)
0.94 AUROC
PASS
in-dist. only — novel-chem AUROC 0.54 (n=237), ~chance
It tells you when it doesn't know. CYP inhibition is AUROC 0.94–0.95 in-distribution (random-split calibration holdout) — but on an independent cross-dataset test (ChEMBL, InChIKey-filtered against ADMET-AI's training data), CYP3A4 and CYP2C9 AUROC drops to 0.47–0.54, statistically indistinguishable from chance; treat the headline CYP number as unvalidated for novel chemistry. hERG is the one target with an honest novel-chemistry number: 0.69 on a random split, and the calibrated ToxScreen pipeline (ChemBERTa applicability domain + isotonic re-weighting) reaches ~0.84 AUROC on de-leaked novel chemistry (ADMET-AI alone: 0.71). We calibrate conservatively and abstain when a compound falls outside the model's applicability domain instead of guessing. We publish the failing/unvalidated numbers, not just the flattering ones — see the Methodology page for the full de-leak writeup.
In-distribution vs. novel chemistry — every endpoint, side by side

The headline AUROCs (blue) are a random-split holdout of the same dataset the model trained on. The honest test is independent novel chemistry (ChEMBL, InChIKey-de-leaked). Where the two dots sit far apart, the in-distribution number does not transfer.

hERG blockade 7CN1 novel-validated
0.4.5.81.0 0.69 0.67

Random split 0.69; independent novel chemistry 0.67 (n=8,886). Stable, not inflated — the one endpoint with a real novel-chemistry number.

CYP3A4 inhibition 6MA7 in-dist only
0.4.5.81.0 0.95 0.47

In-distribution 0.95 collapses to 0.47 on novel chemistry (n=701) — a coin flip. Unvalidated for new scaffolds.

CYP2D6 inhibition 4WNT inconclusive OOD
0.4.5.81.0 0.94 novel: inconclusive

In-distribution 0.94. The novel-chemistry holdout was 99% single-class, so no honest AUROC can be reported.

CYP2C9 inhibition 1OG5 in-dist only
0.4.5.81.0 0.94 0.54

In-distribution 0.94 → 0.54 on novel chemistry (n=237) — near chance.

In-distribution — random-split calibration holdout Independent novel chemistry — ChEMBL, InChIKey-de-leaked chance (0.5) · screening bar (0.8)
Three of four CYP endpoints look excellent in-distribution (AUROC 0.94–0.95) but two of them fall to chance on independent novel chemistry, and the third can’t be tested honestly. We publish this gap rather than hide it — and the product abstains on out-of-domain compounds instead of reporting a number it can’t stand behind.
What calibration buys you — hERG on novel chemistry
0.5 0.7 0.8 bar 0.9 0.71 0.84 +0.13 AUROC ADMET-AI ToxScreen

On de-leaked novel chemistry, the calibrated ToxScreen pipeline (ChemBERTa applicability-domain re-weighting + isotonic calibration) lifts hERG discrimination from the raw ADMET-AI baseline 0.71 to 0.84 AUROC — and, just as important, fixes the model’s confidence:

0.32 → 0.08
Expected calibration error
Raw ADMET-AI → calibrated. Lower is better — a stated 80% is closer to a real 80%.
0.844 ±0.02
Novel-chem AUROC
3-seed mean ± SD, Karim cross-dataset with InChIKey de-leaking.
Calibration is the product. Raw model scores over-state their own confidence (ECE 0.32); after calibration a “0.80” means roughly a real 0.80 (ECE 0.08), and out-of-domain compounds are flagged for abstention rather than guessed.

Random-split benchmark, 3 seeds, n≈1,800 per target (ADMET-AI on ChEMBL-derived actives/inactives). See the Methodology page for full calibration status and known limitations.

Workflow

How it works

1

Submit a SMILES

Paste your compound's SMILES string. We validate with RDKit and canonicalize automatically — no account setup friction.

2

Calibrated ADMET scoring

ADMET-AI (a chemprop graph neural network) scores hERG, CYP3A4, CYP2D6 and CYP2C9, with isotonic-calibrated probabilities and Tanimoto applicability-domain abstention when a compound is outside the model's domain. Deterministic CPU inference, ~1–7 s per compound.

3

Report ready

Download a print-ready HTML, PDF, Excel or JSON report with traffic-light risk classification, the full ADMET profile, and the methodology behind every number.

Capabilities

Exactly what ToxScreen does

No mystery box. Here is the full surface area — every target, endpoint, format and integration — and, just as plainly, what it is not for.

4
CTI safety targets
hERG + CYP3A4 / 2D6 / 2C9, combined into one weighted Composite Tox Index.
15+
ADMET endpoints / compound
Absorption, distribution, metabolism, clearance and safety flags in every report.
1–7 s
Per-compound, on CPU
Deterministic — the same SMILES always returns the same result. No GPU, no queue.
5
Export formats
HTML, PDF, Excel, CSV, JSON — plus REST API, Python SDK and CLI.
CapabilityWhat it doesAvailable in
Screening engine
4-target CTI safety panel hERG blockade + CYP3A4 / 2D6 / 2C9 inhibition, weighted into a Composite Tox Index (0–1) across five risk bands. All plans
Extended ADMET panel 15+ supplementary endpoints — Caco-2, P-gp, BBB, PPB, Vd, clearance, half-life, Ames, DILI, carcinogenicity, LD50 and more. All plans
Applicability-domain abstention A Tanimoto domain check flags out-of-domain chemistry and abstains rather than reporting a number it can’t stand behind. All plans
Calibrated probabilities Isotonic + conformal calibration with intervals — a “0.8” is close to a real 0.8, not a raw model score. All plans
RDKit pre-flight Structure validation, canonicalization and PAINS / physicochemical sanity checks before any compound is scored. All plans
Reports & export
Multi-format reports Print-ready HTML, PDF, Excel, CSV and JSON — traffic-light verdict, full ADMET profile, and the methodology behind every number. All plans
Sample report gallery Real pipeline output for reference compounds, so you can inspect a full report before signing up. Public
Integration & workflow
REST API, Python SDK & CLI Documented endpoints (OpenAPI), a typed Python client, and a command-line tool for scripting screens into your pipeline. Starter +
Batch CSV upload Score up to 100 compounds per job from a single CSV, with per-row status and a combined report. Pro +
Signed webhooks HMAC-SHA256-signed completion callbacks so your systems react the moment a screen finishes. Pro +
Projects, CRO worklist & SAR tracking Group compounds into projects, export a CRO assay worklist, and track structure–activity-relationship trends over time. Pro +
Science & integrity
Published calibration — incl. failures Full per-endpoint status, the novel-chemistry de-leak test, and every known limitation — the flattering and the unflattering. Public
Zero mock data Every score traces to a real computation or a measured value. No placeholders, ever. Always

What ToxScreen is for

  • Early triage — flag hERG and CYP liabilities before you commit to synthesis or a wet-lab assay.
  • Prioritising a list — rank a set of candidates by calibrated computational risk.
  • Catching over-reach — know when a compound is too novel for the model to score, so you don’t trust a bad number.
  • A reproducible first pass — a cheap, deterministic filter that feeds a shortlist into the assays that matter.
  • An auditable rationale — a documented, calibrated risk story per compound you can hand to a colleague.

What ToxScreen is not

  • Not a lab replacement — it does not substitute for in vitro or in vivo toxicology.
  • Not a regulatory tool — not GLP, not IND-enabling, not an FDA / EMA / PMDA submission.
  • Not medical advice — no diagnosis, treatment, or clinical guidance; “low risk” is not “safe”.
  • Not a potency predictor — it estimates liability risk, not on-target activity or efficacy.
  • Not validated on novel CYP chemistry — where it can’t stand behind a number, it abstains instead of guessing.

The validation scoreboard

Every target must clear an adversarial scorecard before it can emit a single lead.

The same discipline governs every screen we run: validated receptors only, calibrated thresholds, blinded holdouts, and published failures. Of seven candidate targets, one clears every gate. We publish the whole board — passes and failures alike — because the failures are the evidence that the passes mean something.

The validation scoreboard — every target, every gate

Each target must clear an adversarial scorecard before it can emit a single lead. We publish the whole board, passes and failures alike. Of seven candidates, one clears every gate.

Target Receptor Discrimination Neg-control Blinded holdout Verdict
PARP1 7AAD ✓ ✓ ✓ ✓ Green · emit leads
CD38 8D0M ✓ ✗ ✓ ✗ Blocked
AMPK 4CFE ✓ ✗ ✗ – Blocked
SIRT1 4I5I ✓ ✗ ✗ ✗ Blocked
KEAP1 4L7B ✓ ✗ – ✓ Blocked
BCL-XL 2YXJ ✓ – ✓ – Blocked
mTOR 4JSV ✗ – – – Blocked
✓  gate passed ✗  gate failed –  not established (insufficient data)
The authoritative target-card panel, generated 2026-07-09. 1 of 7 targets is GREEN — PARP1, the only one to clear receptor fit, binder-vs-decoy discrimination, a negative-control false-positive gate, and a 3-of-3 blinded holdout. Every other candidate fails at least one gate and is held back. We would rather ship one honest target than seven hopeful ones.
Why discrimination alone isn’t enough
−0.5 0.0 0.5 1.0 0.80 screening bar PARP1 1.00 / 0.33 ✓ validated AMPK 0.99 / 0.34 ✗ neg-control BCL-XL 0.94 / 0.66 ✗ unvalidated SIRT1 0.82 / 0.48 ✗ neg-control KEAP1 0.66 / −0.39 ✗ discrimination
Discrimination — retrospective ROC-AUC (binders vs decoys) Potency ranking — Spearman ρ (vs measured affinity)
The same targets, on raw retrospective signal. Four of five separate binders from decoys respectably (AUC 0.82–0.99) — yet on the scoreboard above only PARP1 survives. High discrimination is necessary but not sufficient: AMPK and SIRT1 also light up against inert negative-control metabolites (false positives), BCL-XL can’t be validated on the few measured actives available, and KEAP1 fails discrimination outright. Ranking potency (ρ) is harder still and remains the open frontier. Boltz-2 retrospective benchmark vs measured ChEMBL bioactivity, n = 17–42 per target. CD38, previously charted here, was withdrawn after its reference set was traced to the wrong protein; it is Blocked on the scoreboard above.
How a natural product becomes a lead
1

Validate the target

Before a target can emit a single lead it must clear an adversarial scorecard — discrimination, negative-control false-positive rate, and a blinded holdout. Most candidate targets fail. We publish which, and why.

2

Screen natural-product space

Structure-based co-folding scores each compound against the validated pocket. Novel scaffolds are welcome — a physics-informed model generalises across chemotypes, so unprecedented natural products are candidates, not noise.

3

Kill the artifacts

Each hit is re-scored against a panel of unrelated proteins. Bind the target but not the decoys → a real, selective hit. Bind everything → a frequent-hitter artifact, discarded. PAINS and aggregation filters run alongside.

4

Rank, and caveat honestly

What survives is a ranked list of target-selective binder hypotheses — with an explicit confidence flag and the plain statement that these are computational leads for follow-up, not potency claims.

PARP1 — the one target that has cleared every gate
PARP1 catalytic domain (PDB 7AAD) with olaparib bound in the pocket
The NAD⁺-consuming enzyme PARP1 (catalytic domain, PDB 7AAD) with the inhibitor olaparib bound in the pocket (green sticks). It is the only target in our panel that passes every gate — discrimination, negative-control, and a 3-of-3 blinded holdout. Explore it in 3D on RCSB →
The aspiration: turn natural-product space into an auditable pipeline of ranked, target-selective longevity hit hypotheses — validated against real bioactivity, honest about every number, no over-claiming.

Computational hit hypotheses for research — not validated therapeutics. Experimental confirmation required.

Live sample

See a real report

Every report below is generated offline by the real pipeline — pick a compound.

Generating report…

Computational due diligence

A second opinion on the numbers a deal rests on.

We run the screens, we audit the claims, and we publish the failures. That discipline is available as a paid engagement: target de-risking before you wet-lab, and independent technical review of a lead series or deal before you write a check. Delivered with ProjectAlpha compute, in a versioned report you can reproduce.

Target de-risking

Is the target ligandable and is the receptor the right one? A target card with calibration, discrimination, and neg-control evidence — before wet-lab spend.

Lead-series review

Are the "hits" real binders or PAINS artifacts? Triage with the same gates that caught 19 of 23 over-claimed chemotypes in our own pipeline.

Investor-grade DD

An independent read on validation stats, leakage, and underpowering — the patterns that cost real money when missed. Fixed-scope, opinion letter, evidence appendix.

Request a diagnostic

The tool

Free to use

We built this for our own screening work and left it open. We are not going to talk you into a subscription — if you need more than the tool does, that is a conversation, not a plan upgrade. Metered tiers still exist inside the signed-in app from when this was sold as a product; they are not part of what we sell now.

Free
$0
3 lifetime screenings
  • 4-target CTI safety panel
  • Extended ADMET panel (15+ endpoints)
  • HTML + PDF + Excel + JSON reports
  • Abstains outside the applicability domain
  • No credit card
Start free
Need more than this?
Talk to us
Screening at volume, on your target
  • Validation and calibration on your data
  • A screening stack stood up in-house
  • Campaigns run on our GPU node
  • A written read you can reproduce
Book a call