How Generative AI Is Accelerating Drug Discovery for the World's Top Pharma Companies
From lead identification to regulatory submission, generative molecular AI is compressing the drug discovery cycle from years to months. Our approach and the science behind it.
Rina Shah
Director of Life Sciences, Pragmatiq
Key Takeaways
- Generative molecular models can propose novel drug candidates 200x faster than traditional high-throughput screening
- Multi-property optimisation simultaneously targets efficacy, selectivity, ADMET, and synthesisability
- Pragmatiq's platform has contributed to 4 IND-enabling programmes for top-10 pharma clients
- Regulatory-ready documentation generated automatically from model outputs, reducing submission prep by 60%
The Economics of Drug Discovery
The average cost of bringing a new molecular entity from discovery to market approval is $2.6 billion and takes 12 to 15 years. The attrition rate is staggering: fewer than 1 in 10,000 compounds screened in early discovery reaches clinical trials, and fewer than 1 in 10 that enter Phase I ultimately receive approval.
The bottleneck is not science — it is search. The chemical space of drug-like molecules is estimated at 10^60 compounds. Traditional high-throughput screening can test perhaps 10^6 compounds per year. Even with combinatorial chemistry expanding the accessible library, the search is essentially random. You are looking for a key in a space so large that physical exploration is impossible.
What Generative Molecular AI Changes
Generative AI models for drug discovery do not search chemical space randomly. They learn the structure-activity relationships that govern molecular behaviour — which structural features correlate with target binding, metabolic stability, cell permeability, toxicity — and use this learned representation to generate novel molecules that are predicted to satisfy multiple desired properties simultaneously.
Our platform uses a diffusion-based 3D molecular generative model trained on 180 million molecules from public and proprietary databases, including crystallographic binding data from the Protein Data Bank and ADMET profiles from ChEMBL. Given a target protein structure and a set of property constraints — binding affinity above a threshold, hERG toxicity below a threshold, oral bioavailability predicted above 30% — the model generates structurally diverse candidate molecules ranked by predicted multi-property score.
In benchmarking against traditional lead optimisation campaigns, our generative approach identifies candidates meeting all specified constraints in an average of 6 weeks, compared to 12 to 18 months for conventional medicinal chemistry cycles.
Beyond Generation: The Validation Pipeline
Generating novel molecules is only useful if the predictions are accurate. We use a multi-stage validation pipeline that progressively filters generated candidates.
The first filter is computational: molecular dynamics simulations assess binding stability; QSAR models predict ADMET properties; retrosynthesis models score synthesisability to ensure the molecule can be made in a real laboratory. Candidates that pass computational validation are flagged for experimental confirmation.
We partner with three CRO networks in India, the UK, and Singapore for rapid experimental validation. Biochemical binding assays and cellular toxicity screens on shortlisted candidates typically return results within 3 to 4 weeks. These experimental results feed back into the model's training loop, continuously improving prediction accuracy for subsequent campaigns.
The integration of generative design with rapid experimental feedback — what we call the 'Design-Make-Test-Learn' closed loop — is what differentiates our platform from tools that treat molecule generation as a one-shot exercise.
Regulatory Considerations
Regulators have been cautious about AI-generated drug candidates, and rightly so. The FDA's emerging framework for AI/ML in drug development emphasises the need for model explainability, documented training data provenance, and prospective validation.
Our platform generates a complete audit trail for every candidate: the training data used, the property constraints specified, the model version deployed, and the experimental data generated. This documentation package — what we call the AI Drug Discovery File — is structured to directly populate the Chemistry, Manufacturing, and Controls (CMC) sections of an IND application.
Four of our five current IND-enabling programmes have received positive pre-IND feedback from the FDA regarding our AI documentation approach. We believe that regulatory clarity will follow adoption, and that early movers who invest in rigorous documentation now will have a significant competitive advantage when formal guidance is issued.
Rina Shah
Director of Life Sciences, Pragmatiq
A member of Pragmatiq's leadership and research team, writing on AI, venture building, and the industries we serve.
More Articles
Why Clinical Decision Support Systems Are the Next Frontier of AI in Medicine
Gopala Krishna Bhatt · March 2025
Personalised Learning at Scale: How PurpleGene® is Redefining K-12 Education in India
Arjun Kulkarni · February 2025
LAND LORDZ: Building India's First Personalised Agriculture System with Drone AI
Pragmatiq Research Team · January 2025
Want to work with us?
Whether you're building a product, scaling a team, or exploring AI for your industry — let's talk.