Calibrated ADMET and protein-binding diagnostics, with the uncertainty published.
ToxScreen applies an adversarial, multi-gate validation discipline to computational drug-discovery — a target must pass seven independent checks before any lead is emitted. Every number carries its confidence interval and its failure mode.
Benchmarks
Measured on novel chemistry — and published in full
Every headline here survives an independent, de-leaked test. Where the model can’t stand behind a number, we say so — and the product abstains instead of guessing.
On de-leaked novel chemistry — a Karim cross-dataset holdout with every InChIKey overlap removed — the calibrated ToxScreen pipeline (applicability-domain re-weighting + isotonic calibration) lifts hERG discrimination from the raw ADMET-AI baseline 0.71 to 0.84 AUROC, clearing the 0.8 screening bar. Just as important, it fixes the model’s confidence: expected calibration error drops from 0.32 to 0.08.
3-seed mean ± 0.02 SD. hERG is the one endpoint with a genuine novel-chemistry number; the full per-endpoint de-leak — including where we fail — is detailed below.
Calibration — published in full, including where it fails
The headline AUROCs (blue) are a random-split holdout of the same dataset the model trained on. The honest test is independent novel chemistry (ChEMBL, InChIKey-de-leaked). Where the two dots sit far apart, the in-distribution number does not transfer.
Random split 0.69; independent novel chemistry 0.67 (n=8,886). Stable, not inflated — the one endpoint with a real novel-chemistry number.
In-distribution 0.95 collapses to 0.47 on novel chemistry (n=701) — a coin flip. Unvalidated for new scaffolds.
In-distribution 0.94. The novel-chemistry holdout was 99% single-class, so no honest AUROC can be reported.
In-distribution 0.94 → 0.54 on novel chemistry (n=237) — near chance.
On de-leaked novel chemistry, the calibrated ToxScreen pipeline (ChemBERTa applicability-domain re-weighting + isotonic calibration) lifts hERG discrimination from the raw ADMET-AI baseline 0.71 to 0.84 AUROC — and, just as important, fixes the model’s confidence:
Random-split benchmark, 3 seeds, n≈1,800 per target (ADMET-AI on ChEMBL-derived actives/inactives). See the Methodology page for full calibration status and known limitations.
Workflow
How it works
Submit a SMILES
Paste your compound's SMILES string. We validate with RDKit and canonicalize automatically — no account setup friction.
Calibrated ADMET scoring
ADMET-AI (a chemprop graph neural network) scores hERG, CYP3A4, CYP2D6 and CYP2C9, with isotonic-calibrated probabilities and Tanimoto applicability-domain abstention when a compound is outside the model's domain. Deterministic CPU inference, ~1–7 s per compound.
Report ready
Download a print-ready HTML, PDF, Excel or JSON report with traffic-light risk classification, the full ADMET profile, and the methodology behind every number.
Capabilities
Exactly what ToxScreen does
No mystery box. Here is the full surface area — every target, endpoint, format and integration — and, just as plainly, what it is not for.
| Capability | What it does | Available in |
|---|---|---|
| Screening engine | ||
| 4-target CTI safety panel | hERG blockade + CYP3A4 / 2D6 / 2C9 inhibition, weighted into a Composite Tox Index (0–1) across five risk bands. | All plans |
| Extended ADMET panel | 15+ supplementary endpoints — Caco-2, P-gp, BBB, PPB, Vd, clearance, half-life, Ames, DILI, carcinogenicity, LD50 and more. | All plans |
| Applicability-domain abstention | A Tanimoto domain check flags out-of-domain chemistry and abstains rather than reporting a number it can’t stand behind. | All plans |
| Calibrated probabilities | Isotonic + conformal calibration with intervals — a “0.8” is close to a real 0.8, not a raw model score. | All plans |
| RDKit pre-flight | Structure validation, canonicalization and PAINS / physicochemical sanity checks before any compound is scored. | All plans |
| Reports & export | ||
| Multi-format reports | Print-ready HTML, PDF, Excel, CSV and JSON — traffic-light verdict, full ADMET profile, and the methodology behind every number. | All plans |
| Sample report gallery | Real pipeline output for reference compounds, so you can inspect a full report before signing up. | Public |
| Integration & workflow | ||
| REST API, Python SDK & CLI | Documented endpoints (OpenAPI), a typed Python client, and a command-line tool for scripting screens into your pipeline. | Starter + |
| Batch CSV upload | Score up to 100 compounds per job from a single CSV, with per-row status and a combined report. | Pro + |
| Signed webhooks | HMAC-SHA256-signed completion callbacks so your systems react the moment a screen finishes. | Pro + |
| Projects, CRO worklist & SAR tracking | Group compounds into projects, export a CRO assay worklist, and track structure–activity-relationship trends over time. | Pro + |
| Science & integrity | ||
| Published calibration — incl. failures | Full per-endpoint status, the novel-chemistry de-leak test, and every known limitation — the flattering and the unflattering. | Public |
| Zero mock data | Every score traces to a real computation or a measured value. No placeholders, ever. | Always |
What ToxScreen is for
- Early triage — flag hERG and CYP liabilities before you commit to synthesis or a wet-lab assay.
- Prioritising a list — rank a set of candidates by calibrated computational risk.
- Catching over-reach — know when a compound is too novel for the model to score, so you don’t trust a bad number.
- A reproducible first pass — a cheap, deterministic filter that feeds a shortlist into the assays that matter.
- An auditable rationale — a documented, calibrated risk story per compound you can hand to a colleague.
What ToxScreen is not
- Not a lab replacement — it does not substitute for in vitro or in vivo toxicology.
- Not a regulatory tool — not GLP, not IND-enabling, not an FDA / EMA / PMDA submission.
- Not medical advice — no diagnosis, treatment, or clinical guidance; “low risk” is not “safe”.
- Not a potency predictor — it estimates liability risk, not on-target activity or efficacy.
- Not validated on novel CYP chemistry — where it can’t stand behind a number, it abstains instead of guessing.
The validation scoreboard
Every target must clear an adversarial scorecard before it can emit a single lead.
The same discipline governs every screen we run: validated receptors only, calibrated thresholds, blinded holdouts, and published failures. Of seven candidate targets, one clears every gate. We publish the whole board — passes and failures alike — because the failures are the evidence that the passes mean something.
Each target must clear an adversarial scorecard before it can emit a single lead. We publish the whole board, passes and failures alike. Of seven candidates, one clears every gate.
| Target | Receptor | Discrimination | Neg-control | Blinded holdout | Verdict |
|---|---|---|---|---|---|
| PARP1 7AAD | ✓ | ✓ | ✓ | ✓ | Green · emit leads |
| CD38 8D0M | ✓ | ✗ | ✓ | ✗ | Blocked |
| AMPK 4CFE | ✓ | ✗ | ✗ | – | Blocked |
| SIRT1 4I5I | ✓ | ✗ | ✗ | ✗ | Blocked |
| KEAP1 4L7B | ✓ | ✗ | – | ✓ | Blocked |
| BCL-XL 2YXJ | ✓ | – | ✓ | – | Blocked |
| mTOR 4JSV | ✗ | – | – | – | Blocked |
Validate the target
Before a target can emit a single lead it must clear an adversarial scorecard — discrimination, negative-control false-positive rate, and a blinded holdout. Most candidate targets fail. We publish which, and why.
Screen natural-product space
Structure-based co-folding scores each compound against the validated pocket. Novel scaffolds are welcome — a physics-informed model generalises across chemotypes, so unprecedented natural products are candidates, not noise.
Kill the artifacts
Each hit is re-scored against a panel of unrelated proteins. Bind the target but not the decoys → a real, selective hit. Bind everything → a frequent-hitter artifact, discarded. PAINS and aggregation filters run alongside.
Rank, and caveat honestly
What survives is a ranked list of target-selective binder hypotheses — with an explicit confidence flag and the plain statement that these are computational leads for follow-up, not potency claims.

Computational hit hypotheses for research — not validated therapeutics. Experimental confirmation required.
Live sample
See a real report
Every report below is generated offline by the real pipeline — pick a compound.
Computational due diligence
A second opinion on the numbers a deal rests on.
We run the screens, we audit the claims, and we publish the failures. That discipline is available as a paid engagement: target de-risking before you wet-lab, and independent technical review of a lead series or deal before you write a check. Delivered with ProjectAlpha compute, in a versioned report you can reproduce.
Target de-risking
Is the target ligandable and is the receptor the right one? A target card with calibration, discrimination, and neg-control evidence — before wet-lab spend.
Lead-series review
Are the "hits" real binders or PAINS artifacts? Triage with the same gates that caught 19 of 23 over-claimed chemotypes in our own pipeline.
Investor-grade DD
An independent read on validation stats, leakage, and underpowering — the patterns that cost real money when missed. Fixed-scope, opinion letter, evidence appendix.
The tool
Free to use
We built this for our own screening work and left it open. We are not going to talk you into a subscription — if you need more than the tool does, that is a conversation, not a plan upgrade. Metered tiers still exist inside the signed-in app from when this was sold as a product; they are not part of what we sell now.
- 4-target CTI safety panel
- Extended ADMET panel (15+ endpoints)
- HTML + PDF + Excel + JSON reports
- Abstains outside the applicability domain
- No credit card
- Validation and calibration on your data
- A screening stack stood up in-house
- Campaigns run on our GPU node
- A written read you can reproduce
Privacy & your structures
Your SMILES are stored in an access-controlled database, associated only with your account, never shared with third parties for training, marketing, or analytics, and deleted on request. Predictions run on our own secured servers (CPU inference) — no third-party AI provider sees your structures. Email info@toxscreen.ai to delete all your submitted SMILES and reports.
Important disclaimers
Research use only. ToxScreen is a computational pre-screening tool. Predictions are not a substitute for in vitro or in vivo toxicology studies. Decisions affecting human or animal exposure must be supported by experimentally validated assays.
Not a regulatory submission tool. ToxScreen output cannot replace IND-enabling studies, GLP toxicology packages, or any FDA / EMA / PMDA-mandated assay. The reports are intended for early-stage triage prior to wet-lab work.
No medical advice. ToxScreen does not provide diagnosis, treatment recommendations, or clinical guidance. Compounds flagged "low risk" may still be unsafe.
Confidence intervals. Computational predictions carry uncertainty. Per-target AUC posted in the validation strip above quantifies model performance on the calibration benchmark; performance on out-of-distribution chemistry will differ.
Read the full methodology → · Pipeline, calibration data, known limitations, references, privacy, full legal terms.