The Role of Cell Atlases in Preclinical Target Validation
Before a target hypothesis reaches animal work, there is usually a period of in silico triage: literature searches, protein structure lookups, pathway databases. These are useful filters. But they do not answer the most basic question about whether a target makes biological sense in the disease context: which cells express it, and are those the cells that drive the pathology?
Cell atlases have changed what is possible during this triage phase. The Human Cell Atlas project, CELLxGENE, and related reference datasets have made single-cell expression profiles for tens of millions of cells across dozens of organs and tissue types publicly available. A target hypothesis can now be interrogated against human biology at cell-type resolution before a single in vitro experiment runs, and the data needed to do that is open.
What Cell Atlases Actually Provide
A cell atlas, in the single-cell context, is an annotated reference collection of cells with known types, subtypes, tissue origins, and often donor metadata such as age and disease status. The Human Cell Atlas has catalogued over 50 million cells across more than 33 organs. CELLxGENE centralizes data from hundreds of published studies in a queryable format.
For a target validator, the practical utility is straightforward. You can ask: in healthy human tissue, which annotated cell types express this gene, and at what level? Which tissues show expression? What is the distribution across donors? When disease-state samples are available in the atlas, you can begin to ask how expression changes in the diseased context relative to healthy controls.
These are not answers. They are constraints. A target that shows ubiquitous high expression across dozens of cell types in multiple organs is a different risk profile than one with narrow, cell-type-specific expression in the tissue of interest. Atlas data cannot tell you whether modulating a target will produce a therapeutic effect. It can tell you whether the target expression pattern is consistent with the therapeutic hypothesis.
Specificity as a First-Order Filter
Cell-type specificity is one of the most useful filters atlas data provides at the preclinical stage. A target that is expressed in the disease-relevant cell population but not in surrounding populations is mechanistically tractable in a way that a broadly expressed target is not. Broad expression does not disqualify a target, but it raises the question of on-target effects in non-disease tissues that need to be accounted for in the program design.
The specificity question has a precision and a recall component. Precision: is expression in disease-relevant cells high relative to off-target cells? Recall: is expression present in a sufficient fraction of disease-relevant cells to expect a pharmacological effect? A target can fail on either dimension. Expression restricted to a rare subtype that is only transiently present may not be tractable even if the biology is otherwise compelling.
Atlas data gives you the first estimates of both. Specificity scores calculated from atlas cell-type expression profiles are not equivalent to the complex pharmacological picture you eventually get from in vivo work, but they are a fast, low-cost filter that can eliminate clearly problematic targets early.
Using Atlas Data to Contextualize Disease Hypotheses
The more nuanced use of atlas data is in building the biological rationale for a target hypothesis, not just filtering against it. Consider a hypothesis that a particular cell population drives inflammatory tissue damage in a chronic disease. Atlas data from both healthy and diseased tissue samples can show whether that cell population expands in disease, whether it acquires distinctive expression programs in the diseased state, and which genes are differentially expressed in the disease context relative to the same cell type in healthy tissue.
That differential expression profile is the basis for identifying candidate targets. Genes that are upregulated specifically in the disease-associated cell population and not in healthy controls of the same cell type are candidates that could, in principle, selectively modulate the disease-driving biology. Genes that are also upregulated in healthy versions of the same cell type may not provide the therapeutic window needed to achieve efficacy without toxicity.
This framing depends on having access to disease-state single-cell data, not just healthy tissue references. Atlas collections are increasingly including disease samples, but coverage is uneven. Common inflammatory diseases, certain cancers, and conditions with active research communities are reasonably well covered. Rare diseases may have only a handful of single-cell datasets, or none at all. For those indications, healthy reference atlas data is still useful for characterizing the biology of the relevant cell types, but the disease-specific differential expression analysis is limited by what is in the literature.
The Transferability Question
Atlas data is almost always from humans, which is why it is valuable for target validation. The catch is that preclinical development still relies heavily on animal models. A target hypothesis that is well-supported by human atlas data may not have clean animal model equivalents, either because the cell-type equivalents have different gene expression programs or because the relevant human cell population does not have a direct murine counterpart.
This is not an argument against using atlas data. It is an argument for being explicit about the translation step when moving from computational validation to in vivo work. When the human atlas supports a cell-type-specific expression hypothesis but the mouse model uses a cell type with different marker gene expression, the connection needs to be made explicitly rather than assumed. The growing availability of mouse cell atlases through resources like the Mouse Cell Atlas and Allen Brain Cell Atlas means this comparison is now more tractable than it was a few years ago.
Integration with Computational Target Prioritization
Atlas data is most powerful when it is integrated into a broader computational prioritization workflow rather than used as a standalone query. At Relation Therapeutics, atlas reference data is part of how the platform contextualizes the disease cell populations it identifies. When the model surfaces a cell population associated with a disease state, atlas references help anchor that population's identity, characterize its gene expression specificity, and compare it to equivalent cell types in healthy tissue.
The output is a target hypothesis that comes with a built-in biological rationale: this gene is differentially expressed in this specific cell population in disease relative to both healthy cells of the same type and to neighboring cell populations. That framing is more informative than a ranked gene list without context, because it tells discovery teams what question the target hypothesis is answering and what additional evidence would strengthen or weaken the case.
Atlas data does not replace in vitro or in vivo validation. Target hypotheses generated from atlas-informed computational analysis still need experimental follow-up. The value of atlas integration at the preclinical stage is in narrowing the hypothesis space before expensive experimental work begins, so that the experiments that do run are answering questions the computational analysis could not resolve rather than recapitulating things it already showed.