How do 41 scattered GWAS loci turn out to be one pathway?
Genome-wide association studies give amazing insights into human biology — but it is a huge challenge to interpret the functions of all the noncoding variants discovered, even for a single locus, let alone the many thousands of signals now known.
We developed an approach to accelerate this. Pick a cell type enriched for heritability; build a Variant-to-Gene map with the Activity-by-Contact model; build a Gene-to-Program map with Perturb-seq; trace the path from variant to gene to program (V2G2P); then test for convergence on programs.
The first half uses the Activity-by-Contact (ABC) model to build a genome-wide map of enhancer–gene regulatory interactions, then links noncoding variants to their target genes. The second half uses Perturb-seq to connect genes into pathways in an unbiased and systematic fashion: we knock down all genes near all GWAS signals for a disease, then use the data to reconstruct gene regulatory “programs” with unsupervised machine learning.
Together with Rajat Gupta, we applied this approach to study the contribution of endothelial cells to genetic risk for coronary artery disease.
After applying Perturb-seq we found 50 programs matching a wide array of pathways — responses to inflammation, laminar blood flow, EMT, and others. If you knock down the right gene in vitro, you can trigger responses downstream of a wide array of physiological stimuli.
Next we trace the path from variant to gene to program, and test whether the links are enriched in particular programs. For coronary artery disease, we found five. They contain well-known risk genes like NOS3 and PLPP3, along with many new ones such as TLNRD1.
All five programs are branches downstream of the Cerebral Cavernous Malformations (CCM) complex — previously linked to a rare vascular disease in the brain, but not to genetic risk for coronary artery disease. TLNRD1, a poorly studied gene, is a new member of this complex.
This has three implications for how we approach the variant-to-function challenge.
One. It validates the hypothesis that causal genes for a disease are connected in gene regulatory networks — and shows that we can find them with systematic Perturb-seq and V2G2P. Here we linked 41 coronary artery disease GWAS loci to a single pathway in endothelial cells.
Two. Neither variant-to-gene information (from ABC or eQTLs) nor gene-to-pathway information (from Perturb-seq) is sufficiently specific to identify causal genes in a GWAS locus. You need both. In our analysis, the average GWAS locus has 8.5 expressed genes, 2 genes with a variant-to-gene link, 4 genes connected to a gene program — and 1 nominated gene after V2G2P.
This complexity is just how the genome is wired. Enhancers can regulate many genes in a relatively nonspecific fashion, and random genes in a locus will be involved in a wide variety of pathways. So we need comprehensive maps of both variant-to-gene and gene-to-pathway connections to interpret GWAS loci.
Three. The approach generalizes. In revision we applied V2G2P successfully to other vascular traits like blood pressure, and to red blood cell traits using the genome-scale Perturb-seq dataset from Joseph Replogle and the Weissman lab. For blood pressure, V2G2P detects distinct genes and programs in endothelial cells from those linked to coronary artery disease risk.
By applying Perturb-seq and the V2G2P approach across many cell types and states, it should be possible to nominate causal disease genes for a large fraction of GWAS loci and map how they converge onto particular cellular pathways.
Let’s get to work.
Read the paper in Nature.