To mitigate reference bias and improve gene expression analysis, researchers at UC Santa Cruz have developed a method for analyzing RNA sequencing data genome-wide using a “pantranscriptome,” which combines a transcriptome and a pangenome—a reference that contains genetic material from a cohort of diverse individuals, rather than just a single linear strand.
“This is pangenome plus transcriptome—that combination has never really been done before until now,” said Jordan Eizenga, co-first author of the paper published in Nature Methods. “This is the first time anyone has attempted to incorporate the pangenome as a standard feature of the RNA sequencing mapping.”
“With this toolkit, we are employing this more diverse data that we can now get from the pangenome to improve the measurement of gene expression data, something that can widely vary between individuals,” explained study leader Benedict Paten. “The aim is to make the impact of this more diverse data felt on studies that are looking at gene expression, resulting in better analysis for cell models, organoid models, and other research applications.”
Search Antibodies Search Now Use our Antibody Search Tool to find the right antibody for your research. Filter
by Type, Application, Reactivity, Host, Clonality, Conjugate/Tag, and Isotype.
Using these new tools, the researchers can take the spliced segments of an individual’s RNA, map where they align on a pangenome, identify which haplotype the data belongs to, and analyze gene expression. First, the areas of the genome the RNA sequencing data comes from are identified, including the splice sites, and those points marked on the pangenome reference. Those marked points are then compared to a pantranscriptome consisting of haplotype-specific transcripts generated from the reference data contained within the pangenome.
Finally, the new workflow generates estimates of levels of gene expression based on this comparison between the mapped data and the transcripts in the pantranscriptome, and identifies which haplotypes the genes come from.
“It's definitely a very forward-looking study in that other genome-wide expression methods are not yet really utilizing pangenomes and haplotype information,” said Jonas Sibbesen, co-first author on the study. “We're now thinking ahead as to what pangenomics might additionally bring to the table in transcriptomic analyses.”