AI/TLDR

Insilico Medicine · 2026-07-30 · major

Insilico DDD Benchmark — Nature-Portfolio-cited yardstick for drug-discovery AI

Insilico Medicine launched the Drug Discovery and Development (DDD) Benchmark as a Service: 300+ decontaminated tasks plus end-to-end candidate-nomination runs, with a public leaderboard at dddbench.insilico.com.

Insilico Medicine DDD Benchmark press release cover on BioSpace

The first standardised, decontaminated leaderboard that scores frontier AI on real drug-discovery work, from disease biology to preclinical candidate nomination.

Quick facts

MakerInsilico Medicine
Portaldddbench.insilico.com
Suite 1Drug Discovery Foundations — 300+ decontaminated evaluations
Suite 2Drug Candidate Essentials — end-to-end hit → preclinical candidate
CoverageDisease biology, molecular property prediction, retrosynthesis, structure-based design, clinical dev
AccessAny org with a chat-completions API; private assessments + verified leaderboard placement
Anchored on12+ years of Insilico's validated programs (31 preclinical candidates)

What is it?

The DDD Benchmark is a Benchmark-as-a-Service from Insilico Medicine that scores frontier and foundation models on real drug-discovery tasks. It runs on a two-suite design — Drug Discovery Foundations for component skills and Drug Candidate Essentials for end-to-end programs — with a public leaderboard hosted at dddbench.insilico.com.

How does it work?

Drug Discovery Foundations runs 300+ evaluations across disease biology, molecular property prediction, retrosynthesis, structure-based design, and clinical development, using out-of-distribution test sets and rigorously decontaminated public data. Drug Candidate Essentials assesses whether a model can carry a program from hit identification through preclinical candidate nomination, anchored to Insilico's own 12+ years of validated pipeline work.

Why does it matter?

Existing AI benchmarks for chemistry and biology are frequently contaminated, so an LLM can look great and still fail the real preclinical decisions a pharma team cares about. Insilico's DDD Benchmark uses out-of-distribution test sets pulled from actual industry programs — including its 31 nominated preclinical candidates and the Phase-III Rentosertib — giving buyers of foundation-model drug-discovery tools a like-for-like scorecard.

Who is it for?

pharma R&D leaders, foundation-model teams targeting biology, and evaluators comparing drug-discovery AI

Frequently asked questions

Why does drug-discovery AI need its own benchmark?
Insilico argues that most public AI benchmarks are contaminated: frontier models score highly by memorising test questions that leaked into training data, then fail on real preclinical candidate decisions. The DDD Benchmark uses out-of-distribution test sets and decontaminated public data drawn from Insilico's own 12+ years of validated programs, so contamination cannot pump the score.
What exactly does the DDD Benchmark measure?
DDD ships two suites. Drug Discovery Foundations runs 300+ evaluations across disease biology, molecular property prediction and optimisation, retrosynthesis, structure-based design, and clinical development. Drug Candidate Essentials scores an end-to-end program from hit identification through preclinical candidate nomination, anchored to Insilico's real pipeline (which has nominated 31 preclinical candidates and put Rentosertib/ISM001-055 into Phase III).
Who can submit to the DDD leaderboard?
The DDD Benchmark is a Benchmark-as-a-Service — any organization with a model accessible through a standard chat-completions API can request an assessment. Insilico offers private evaluations, verified score reports, and optional placement on the public leaderboard at dddbench.insilico.com.

Try it

dddbench.insilico.com

Sources · 4 outlets

Tags

  • insilico-medicine
  • ddd-benchmark
  • drug-discovery
  • pharma-ai
  • molecular-design
  • computational-chemistry
  • target-identification
  • foundation-model
  • evaluation
  • leaderboard

← All releases · Learn AI