Calibrated ADMET predictions that flag when your chemistry is too novel to trust — built on ADMET-AI, honest about uncertainty.
Benchmarks
Every headline here survives an independent, de-leaked test. Where the model can’t stand behind a number, we say so — and the product abstains instead of guessing.
On de-leaked novel chemistry — a Karim cross-dataset holdout with every InChIKey overlap removed — the calibrated ToxScreen pipeline (applicability-domain re-weighting + isotonic calibration) lifts hERG discrimination from the raw ADMET-AI baseline 0.71 to 0.84 AUROC, clearing the 0.8 screening bar. Just as important, it fixes the model’s confidence: expected calibration error drops from 0.32 to 0.08.
3-seed mean ± 0.02 SD. hERG is the one endpoint with a genuine novel-chemistry number; the full per-endpoint de-leak — including where we fail — is detailed below.
The headline AUROCs (blue) are a random-split holdout of the same dataset the model trained on. The honest test is independent novel chemistry (ChEMBL, InChIKey-de-leaked). Where the two dots sit far apart, the in-distribution number does not transfer.
Random split 0.69; independent novel chemistry 0.67 (n=8,886). Stable, not inflated — the one endpoint with a real novel-chemistry number.
In-distribution 0.95 collapses to 0.47 on novel chemistry (n=701) — a coin flip. Unvalidated for new scaffolds.
In-distribution 0.94. The novel-chemistry holdout was 99% single-class, so no honest AUROC can be reported.
In-distribution 0.94 → 0.54 on novel chemistry (n=237) — near chance.
On de-leaked novel chemistry, the calibrated ToxScreen pipeline (ChemBERTa applicability-domain re-weighting + isotonic calibration) lifts hERG discrimination from the raw ADMET-AI baseline 0.71 to 0.84 AUROC — and, just as important, fixes the model’s confidence:
Random-split benchmark, 3 seeds, n≈1,800 per target (ADMET-AI on ChEMBL-derived actives/inactives). See the Methodology page for full calibration status and known limitations.
Workflow
Paste your compound's SMILES string. We validate with RDKit and canonicalize automatically — no account setup friction.
ADMET-AI (a chemprop graph neural network) scores hERG, CYP3A4, CYP2D6 and CYP2C9, with isotonic-calibrated probabilities and Tanimoto applicability-domain abstention when a compound is outside the model's domain. Deterministic CPU inference, ~1–7 s per compound.
Download a print-ready HTML, PDF, Excel or JSON report with traffic-light risk classification, the full ADMET profile, and the methodology behind every number.
Capabilities
No mystery box. Here is the full surface area — every target, endpoint, format and integration — and, just as plainly, what it is not for.
| Capability | What it does | Available in |
|---|---|---|
| Screening engine | ||
| 4-target CTI safety panel | hERG blockade + CYP3A4 / 2D6 / 2C9 inhibition, weighted into a Composite Tox Index (0–1) across five risk bands. | All plans |
| Extended ADMET panel | 15+ supplementary endpoints — Caco-2, P-gp, BBB, PPB, Vd, clearance, half-life, Ames, DILI, carcinogenicity, LD50 and more. | All plans |
| Applicability-domain abstention | A Tanimoto domain check flags out-of-domain chemistry and abstains rather than reporting a number it can’t stand behind. | All plans |
| Calibrated probabilities | Isotonic + conformal calibration with intervals — a “0.8” is close to a real 0.8, not a raw model score. | All plans |
| RDKit pre-flight | Structure validation, canonicalization and PAINS / physicochemical sanity checks before any compound is scored. | All plans |
| Reports & export | ||
| Multi-format reports | Print-ready HTML, PDF, Excel, CSV and JSON — traffic-light verdict, full ADMET profile, and the methodology behind every number. | All plans |
| Sample report gallery | Real pipeline output for reference compounds, so you can inspect a full report before signing up. | Public |
| Integration & workflow | ||
| REST API, Python SDK & CLI | Documented endpoints (OpenAPI), a typed Python client, and a command-line tool for scripting screens into your pipeline. | Starter + |
| Batch CSV upload | Score up to 100 compounds per job from a single CSV, with per-row status and a combined report. | Pro + |
| Signed webhooks | HMAC-SHA256-signed completion callbacks so your systems react the moment a screen finishes. | Pro + |
| Projects, CRO worklist & SAR tracking | Group compounds into projects, export a CRO assay worklist, and track structure–activity-relationship trends over time. | Pro + |
| Science & integrity | ||
| Published calibration — incl. failures | Full per-endpoint status, the novel-chemistry de-leak test, and every known limitation — the flattering and the unflattering. | Public |
| Zero mock data | Every score traces to a real computation or a measured value. No placeholders, ever. | Always |
Where we're headed
ToxScreen catches your liabilities. The engine behind it now screens natural-product space for compounds that bind validated longevity targets — under the same rules that govern the tox panel: validated targets only, publish where it fails, abstain when the model doesn't know.
Each target must clear an adversarial scorecard before it can emit a single lead. We publish the whole board, passes and failures alike. Of seven candidates, one clears every gate.
| Target | Receptor | Discrimination | Neg-control | Blinded holdout | Verdict |
|---|---|---|---|---|---|
| PARP1 7AAD | ✓ | ✓ | ✓ | ✓ | Green · emit leads |
| CD38 8D0M | ✓ | ✗ | ✓ | ✗ | Blocked |
| AMPK 4CFE | ✓ | ✗ | ✗ | – | Blocked |
| SIRT1 4I5I | ✓ | ✗ | ✗ | ✗ | Blocked |
| KEAP1 4L7B | ✓ | ✗ | – | ✓ | Blocked |
| BCL-XL 2YXJ | ✓ | – | ✓ | – | Blocked |
| mTOR 4JSV | ✗ | – | – | – | Blocked |
Before a target can emit a single lead it must clear an adversarial scorecard — discrimination, negative-control false-positive rate, and a blinded holdout. Most candidate targets fail. We publish which, and why.
Structure-based co-folding scores each compound against the validated pocket. Novel scaffolds are welcome — a physics-informed model generalises across chemotypes, so unprecedented natural products are candidates, not noise.
Each hit is re-scored against a panel of unrelated proteins. Bind the target but not the decoys → a real, selective hit. Bind everything → a frequent-hitter artifact, discarded. PAINS and aggregation filters run alongside.
What survives is a ranked list of target-selective binder hypotheses — with an explicit confidence flag and the plain statement that these are computational leads for follow-up, not potency claims.

Computational hit hypotheses for research — not validated therapeutics. Experimental confirmation required.
Live sample
Every report below is generated offline by the real pipeline — pick a compound.
Pricing
Start free — no credit card. Upgrade when the screenings pay for themselves.
Your SMILES are stored in an access-controlled database, associated only with your account, never shared with third parties for training, marketing, or analytics, and deleted on request. Predictions run on our own secured servers (CPU inference) — no third-party AI provider sees your structures. Email info@toxscreen.ai to delete all your submitted SMILES and reports.
Research use only. ToxScreen is a computational pre-screening tool. Predictions are not a substitute for in vitro or in vivo toxicology studies. Decisions affecting human or animal exposure must be supported by experimentally validated assays.
Not a regulatory submission tool. ToxScreen output cannot replace IND-enabling studies, GLP toxicology packages, or any FDA / EMA / PMDA-mandated assay. The reports are intended for early-stage triage prior to wet-lab work.
No medical advice. ToxScreen does not provide diagnosis, treatment recommendations, or clinical guidance. Compounds flagged "low risk" may still be unsafe.
Confidence intervals. Computational predictions carry uncertainty. Per-target AUC posted in the validation strip above quantifies model performance on the calibration benchmark; performance on out-of-distribution chemistry will differ.
Read the full methodology → · Pipeline, calibration data, known limitations, references, privacy, full legal terms.