Research

We develop computational and experimental approaches to understand transcriptional regulation, genetic and somatic variation, repetitive DNA, and the mechanisms that connect genome sequence to cellular function and disease.

Diagram of transcriptional regulation in eukaryotes
Transcriptional regulation in eukaryotes. Source.

Our scientific focus

From regulatory sequence to biological function

Understanding gene regulation at the transcriptional level is critical to understanding complex biological systems and human disease. In virtually all organisms, gene regulation is mediated by a regulatory code in which distinct combinations of transcription factors collaborate to regulate the expression of individual genes. This code is complex, is not readily apparent from sequence alone, and involves cis-regulatory modules located both upstream of and within genes.

Variation in these regulatory regions contributes to molecular and phenotypic differences among individuals. Our goal is to identify the mechanisms underlying transcriptional regulation and determine how variation in regulatory and repetitive regions changes genome function.

Current programs

Research projects

Diagram illustrating tandem-repeat characterization

Repetitive variation

Targeted characterization of tandem repeats

Known tandem repeats make up approximately 3% of the human genome and are highly variable across individuals. Tandem repeats are intrinsically unstable, and their expansion is known to cause more than 50 human diseases, including ALS, ataxia, and Huntington’s disease. Characterization and discovery of these loci have proven difficult with short-read sequencing because of their repetitive structure.

Long-read technologies such as Oxford Nanopore sequencing can span repeat elements in their entirety, but their error profiles create additional challenges for accurate copy-number estimation. We combine targeted nanopore sequencing with improved computational methods to characterize tandem repeats at both healthy and pathogenic lengths.

Workflow for long-read characterization of genomic variation

Somatic mosaicism

Long-read sequencing for variant characterization

The human genome varies not only from person to person, but also among cells within an individual. This somatic mosaicism arises after conception and occurs at different rates across tissues. We are working through the SMaHT initiative to investigate somatic variation systematically while improving the sequencing and bioinformatic methods required to detect previously overlooked classes of variation.

We also study somatic mutations and mobile-element insertions in the human brain, including their potential roles in neurodegenerative disease. Long-read sequencing allows us to identify L1HS, Alu, and SVA insertions and determine their genomic locations with substantially greater resolution than conventional approaches.

Diagram connecting genetic variation with regulatory function

Regulatory genomics

The impact of genetic variation on gene regulation

Genome-wide association studies have identified many variants associated with complex traits and disease, but determining their functional consequences remains a major challenge. This is especially true for noncoding variants, which account for most trait-associated loci.

We integrate functional genomic assays and machine-learning models to predict the regulatory activity of noncoding variants in general and in specific tissues or cell types. These approaches underlie tools including SURF, TURF, TLand, and RegulomeDB. We also use functional predictions to improve genetic risk models, particularly in populations that remain underrepresented in genomic studies.

Diagram of transposable-element-derived chromatin loops

Transposable elements and 3D genome structure

Mobile-element-derived chromatin-looping variability

Transposable elements constitute at least 45% of the human genome and have often been dismissed as nonfunctional sequence. A growing body of work instead shows that particular elements can affect gene regulation, disease susceptibility, and genome organization. Our previous research identified transposable elements containing binding motifs for CTCF, a protein central to chromatin-loop formation and three-dimensional genome architecture.

We study whether population-level differences in transposable-element insertions alter CTCF binding and chromatin looping. This work combines computational analysis with targeted capture using guide RNAs and Oxford Nanopore long-read sequencing to resolve specific element families and their regulatory effects.

Diagram of a high-throughput inverted reporter assay

Functional genomics

High-throughput assays for silencers and enhancer blockers

Cis-regulatory elements control the timing, location, and magnitude of gene expression. Promoters and enhancers have been studied extensively, but silencers and enhancer blockers remain comparatively poorly mapped and characterized. A major obstacle is the lack of massively parallel reporter assays designed specifically to test negative regulatory activity.

We are developing high-throughput reporter assays that use dCas9- or LacI-based signal inversion to identify and characterize silencers and enhancer blockers with improved sensitivity, specificity, and lower cell-number and sequencing requirements.