Insilico Medicine · 2026-07-30 · major
Insilico DDD Benchmark — Nature-Portfolio-cited yardstick for drug-discovery AI
Insilico Medicine launched the Drug Discovery and Development (DDD) Benchmark as a Service: 300+ decontaminated tasks plus end-to-end candidate-nomination runs, with a public leaderboard at dddbench.insilico.com.

The first standardised, decontaminated leaderboard that scores frontier AI on real drug-discovery work, from disease biology to preclinical candidate nomination.
Quick facts
| Maker | Insilico Medicine |
|---|---|
| Portal | dddbench.insilico.com |
| Suite 1 | Drug Discovery Foundations — 300+ decontaminated evaluations |
| Suite 2 | Drug Candidate Essentials — end-to-end hit → preclinical candidate |
| Coverage | Disease biology, molecular property prediction, retrosynthesis, structure-based design, clinical dev |
| Access | Any org with a chat-completions API; private assessments + verified leaderboard placement |
| Anchored on | 12+ years of Insilico's validated programs (31 preclinical candidates) |
What is it?
The DDD Benchmark is a Benchmark-as-a-Service from Insilico Medicine that scores frontier and foundation models on real drug-discovery tasks. It runs on a two-suite design — Drug Discovery Foundations for component skills and Drug Candidate Essentials for end-to-end programs — with a public leaderboard hosted at dddbench.insilico.com.
How does it work?
Drug Discovery Foundations runs 300+ evaluations across disease biology, molecular property prediction, retrosynthesis, structure-based design, and clinical development, using out-of-distribution test sets and rigorously decontaminated public data. Drug Candidate Essentials assesses whether a model can carry a program from hit identification through preclinical candidate nomination, anchored to Insilico's own 12+ years of validated pipeline work.
Why does it matter?
Existing AI benchmarks for chemistry and biology are frequently contaminated, so an LLM can look great and still fail the real preclinical decisions a pharma team cares about. Insilico's DDD Benchmark uses out-of-distribution test sets pulled from actual industry programs — including its 31 nominated preclinical candidates and the Phase-III Rentosertib — giving buyers of foundation-model drug-discovery tools a like-for-like scorecard.
Who is it for?
pharma R&D leaders, foundation-model teams targeting biology, and evaluators comparing drug-discovery AI
Frequently asked questions
- Why does drug-discovery AI need its own benchmark?
- Insilico argues that most public AI benchmarks are contaminated: frontier models score highly by memorising test questions that leaked into training data, then fail on real preclinical candidate decisions. The DDD Benchmark uses out-of-distribution test sets and decontaminated public data drawn from Insilico's own 12+ years of validated programs, so contamination cannot pump the score.
- What exactly does the DDD Benchmark measure?
- DDD ships two suites. Drug Discovery Foundations runs 300+ evaluations across disease biology, molecular property prediction and optimisation, retrosynthesis, structure-based design, and clinical development. Drug Candidate Essentials scores an end-to-end program from hit identification through preclinical candidate nomination, anchored to Insilico's real pipeline (which has nominated 31 preclinical candidates and put Rentosertib/ISM001-055 into Phase III).
- Who can submit to the DDD leaderboard?
- The DDD Benchmark is a Benchmark-as-a-Service — any organization with a model accessible through a standard chat-completions API can request an assessment. Insilico offers private evaluations, verified score reports, and optional placement on the public leaderboard at dddbench.insilico.com.
Try it
dddbench.insilico.com