What Makes a Good Drug Target: Lessons from Failed Trials
Drug discovery has a well-documented failure rate. Estimates consistently place the probability of a compound entering Phase I eventually reaching approval somewhere around 5-10%, depending on the disease area and the criteria used. The causes of attrition are distributed across the pipeline: pharmacokinetics, toxicology, and inadequate efficacy are all contributors. But when late-stage Phase II and III failures are analysed, a recurring theme emerges that is specifically tied to the biology of the target: the drug reached its intended molecular target, but the target was not actually driving the disease in the human patients enrolled in the trial.
This is not a pharmacology failure. It is a target selection failure. And the question worth asking is: could that failure have been anticipated earlier, with different data?
The Classic Anatomy of a Target Selection Failure
A typical target selection failure follows a recognisable pattern. A target gene or protein is identified from bulk tissue data in a disease model: either a mouse model, a human cell line, or bulk RNA-seq from patient tissue. The target shows clear differential expression in disease conditions. Functional studies in cell lines or mouse models show that perturbation of the target modulates the disease phenotype. The target is considered validated and a drug development program begins.
The program proceeds through lead identification, optimisation, and eventually into clinical trials. Interim analysis or final readout shows no significant efficacy, or efficacy only in a subset of patients that was not pre-specified. Post-hoc analyses attempt to identify why the drug did not work in the enrolled population.
In a meaningful proportion of these cases, the retrospective analysis reveals a mismatch between the model system where the target was identified and the human disease: the cell type responsible for driving the disease in human patients is not well-represented in the model. The target may be expressed in the right cell type in the mouse, but that cell type contributes a different fraction of the pathological mechanism in human disease. Or the target was identified from bulk human data but was actually expressed primarily in a cell population that is present in the tissue but not the dominant driver of pathology.
What the Literature Shows About Failure Modes
The published literature on late-stage clinical failures is extensive, and examining the publicly documented cases reveals several recurring patterns in target selection that are relevant to how cell-type data changes the calculus.
Several programs in inflammatory diseases identified targets from bulk tissue analysis and showed strong preclinical evidence in rodent models, then failed at Phase II or Phase III because the human disease involves a more complex mixture of cell types contributing to pathology than the model captured. The target was real and expressed where it was thought to be expressed, but the cell type was a contributor to, rather than the driver of, the human disease phenotype. Interventions that modified one contributor without addressing the others produced insufficient efficacy.
In neurodegenerative diseases, the target selection challenge is compounded by the fact that human brain tissue is difficult to access at the relevant disease stages, and mouse models have different cell-type proportions and activation states from human disease. Targets that perform well in mouse AD models have a poor track record in human trials, partly because the microglial biology, which is central to human disease pathology, is substantially different in mice.
Oncology presents the problem of tumour heterogeneity: a target that is expressed in the bulk of a tumour cell population may be absent or low in the stem-like or drug-resistant subpopulation that eventually repopulates the tumour after initial response. Identifying this population and its specific molecular programme is a single-cell-resolution problem.
The Role of Cell-Type Evidence in Target Confidence
The question is not whether human genetics or pathway biology is wrong as a starting point for target identification. Those are important inputs. The question is what additional evidence raises confidence that a genetically associated or pathway-relevant gene is actually mechanistically important in the right cell population in human tissue.
Single-cell expression evidence is one component of that additional evidence. Specifically, the combination of: (a) a target gene being differentially expressed in a specific cell type in disease tissue relative to the same cell type in healthy tissue, (b) that cell type being compositionally expanded or transcriptionally shifted in disease, and (c) the gene's expression being concentrated in that cell type rather than broadly distributed across the tissue, constitutes a higher-confidence hypothesis than association data alone.
We are not claiming this evidence is sufficient. A gene can satisfy all three of the above criteria and still fail in the clinic for other reasons: its expression in the target cell is downstream of the driver rather than causal, the animal models used for in vivo validation do not recapitulate the relevant cell biology, or the drug's pharmacological profile prevents adequate target engagement in the relevant tissue compartment. Cell-type expression evidence raises confidence but does not guarantee success.
What Good Target Evidence Looks Like in Practice
From reviewing publicly documented examples of programs that succeeded in the clinic alongside the failures, a rough picture of what distinguishes higher-quality target evidence emerges.
Targets with human genetic support (GWAS associations, Mendelian disease genes, rare variant enrichment in disease patients) have a better track record than those identified from expression data alone. The reason is mechanistic: a gene with a loss-of-function variant that protects against a disease is more likely to be causally involved than one that is merely correlated with disease state in expression data. Cell-type expression evidence then tells you where in the tissue that causal gene is acting.
Targets with consistent evidence across multiple independent studies and multiple model systems are more likely to be real. A target that shows up clearly in bulk data from one cohort, in single-cell data from a second cohort, and in functional perturbation screens in relevant cell lines is stronger than one supported by a single dataset. The power of large-scale single-cell reference datasets is that they enable this cross-study consistency check in ways that were not previously possible.
Targets with a clear mechanistic rationale for how perturbation could interrupt the disease process are stronger than those identified by association alone. If you can explain why inhibiting this gene in this cell type would reduce the pathological output of that cell type, and the explanation is grounded in known biology, the hypothesis is more likely to hold up as you learn more. If the only rationale is correlation, fragility is high.
The Honest Limits of This Framework
Analysing failures retrospectively is easier than predicting them prospectively. The factors that distinguish a successful target hypothesis from a failed one are only clearly visible in hindsight, and the selection biases in what gets written up and published mean that the apparent pattern in published analyses may not fully reflect the underlying reality.
There is also selection bias in which programs reach clinical trials at all. Programs with weak cell-type evidence and uncertain human biology may have been stopped in preclinical development, and those successes are invisible in analyses of clinical trial outcomes. The failure rate we observe at Phase II and III reflects the entire selection process up to that point, not just the last step.
What we can say is that cell-type resolved evidence from human tissue changes the shape of the target hypothesis in ways that have historically been associated with better outcomes. It does not make target selection easy. It makes it more systematic and more anchored in the biology that actually matters for clinical success, which is a meaningful improvement over starting from bulk data and model systems that may not capture the relevant human cellular context.