Understanding gene regulation may be the key to interpreting human disease genetics. Genes are regulated in part by stretches of DNA called enhancers, which define when, where, and how strongly each gene is turned on. Mapping enhancers and how they function in specific cell types is necessary for understanding gene regulation and disease-related genetic variants. But the location and activity of enhancers are highly cell-type-specific, which makes it difficult to accurately predict enhancer-gene interactions.
In recent years, researchers have developed several computational models to predict enhancer–gene regulatory interactions using measurements of chromatin state and three-dimensional contacts. These models have produced enhancer–gene maps spanning hundreds of cells and tissues. However, these methods remain limited, and confirming their accuracy is difficult because the necessary experiments have been done in only a handful of cell types.
In a study recently published in Nature Genetics , researchers at Stanford University, including first author Maya Sheth, and senior author, Jesse Engreitz, PhD , developed single-cell enhancer-to-gene prediction models, scE2G, that predict genome-wide enhancer interactions from either scATAC or multiomic scATAC and scRNA-seq data. The scE2G models use the single-cell data to predict which DNA regions act as enhancers and which genes they control.
The researchers trained the models using CRISPR experiments in which scientists had directly tested more than 10,000 candidate enhancer–gene pairs. Once trained, the models can be applied to data from other cell types. Because single-cell data can sort cell types apart computationally, scE2G can build maps for cell types that are too rare or too difficult to isolate for bulk methods to reach, and can show how gene regulation differs from one type of cell to another.
The models perform well on datasets of varying sizes and sequencing depths, meaning they can be applied to the many single-cell datasets researchers have already collected. The team has already used them on complex tissues to trace disease-associated variants to their target genes, linking two genes, INPP4B and IL15, to the number of lymphocytes in the blood, which is a kind of connection that would have been difficult to make from noncoding DNA alone. As single-cell datasets continue to expand, the models could eventually chart enhancer–gene regulation across thousands of the cell types that make up the human body.
Additional Stanford University investigators include X Rosa Ma, Andreas R Gschwind, Anthony S Tan, James Galante, Dulguun Amgalan, Danilo Dubocanin, Kayla Brand, Lars M Steinmetz, and Anshul Kundaje.