Thanks to next-generation sequencing, whole-genome sequencing has become routine, accessible, and relatively inexpensive. Yet subsets of the whole genome often provide equally reliable results. One such collection of genes, the protein-coding exome, numbers approximately 180,000 genes, or two percent of the genome.
As the blueprint for an organism’s protein-generating apparatus, the exome holds the secrets to many diseases. Thus, investigators often prefer to sequence the exome or even subsets of the exome to save time and money. These subsets, or gene panels, focus on a limited number of genes whose expressions (or lack thereof) are associated with diseases.
Note that the term “whole-genome sequencing” (WGS) is somewhat inaccurate. WGS skips genomic regions that are difficult to sequence, such as regions with high GC content, large repeat sequences repeats, centromeres, and telomeres.
Less can be more
Simple arithmetic tells us that if the exome comprises just 2% of the genome, exome sequencing misses 98% of the genetic message. According to Shawn Baker, Ph.D., adviser and consultant at SanDiegOmics.com, that is not necessarily a serious shortcoming since most of the non-exonic regions are poorly characterized, and their functions not very well understood. “Whether you do whole-genome or whole-exome depends on who you are and what kind of science you’re doing. A researcher might be more interested in the entire genome, while a clinician is more interested in obtaining an answer quickly.”
Hence the popularity of exome subsets, or panels, that look for anywhere from a few exomic genes to several hundred. Baker adds, “This approach gives the biggest bang for the buck, by focusing on genes that are more likely to be involved in the question at hand.” Panels are also popular among researchers who then have the option to run larger panels, whole-exome, or whole-genome experiments if necessary.
Whole-genome sequencing is often conducted in discovery mode, for example in the investigation of rare diseases of unknown causes. The less is known about the scientific issue under investigation, the broader the net is cast. Assuming that much of what is knowable about the state of an organism is reflected somehow in its genome, the choice of protein-coding genes might seem arbitrary. Why not select any 2%, or some other small fraction of genes, arbitrarily?
“The reason is that proteins make you who you are, and determine how your cells work,” Baker explains. “If you were given the choice between experiencing a random mutation you would prefer to have it outside the exome, because there it’s less likely to cause a problem. So it follows that if you’re looking into the cause of a problem you’d more likely find it in the exome.”
One final objection: Why conduct exome sequencing when whole-genome sequencing provides exomic data and everything else?
“Many researchers feel that’s the approach to take,” Baker says. “However, it’s more expensive, and more difficult to analyze the data. That’s why academic researchers with the luxury of time and funding, may choose whole-genome sequencing. But the more applied your research, the faster you need answers, so exome sequencing is the way to go.”
However, Baker concedes that the debate on whole-genome versus exome sequencing continues, even among physicians.
“Individuals who sequence for medical diagnosis are passionate about which sequencing protocol is right. The exome people say, ‘I only have so much money to spend; I can help more patients if I sequence only the exome.’ The whole-genome adherents argue that they can better serve individual patients by looking at the entire genome.”
As sequencing capabilities improve, and costs fall even further, the consensus opinion is that cost differences between WGS and exome sequencing will vanish, and may even favor WGS.
“The cost differential used to be quite large but has fallen to a factor of two or three,” Baker says. “Historically, for any DNA sequencing experiment, generating the sequence accounted for 99% of the total cost. As the sequencing component drops in price, other aspects of the sequencing workflow become more prominent, for example sample preparation or data analysis. For example, sample prep for WGS is significantly lower than for exome sequencing.” The cost of sequencing the exome will always be lower than for WGS, but that difference may eventually become insignificant when the cost of all workflow steps are figured in.
Exome sequencing begins with a pull-down of genes of interest, hybridization to oligos affixed to beads, and washing away everything else. Exome sequencing is also conducted at higher depth, 100x vs. just 30x for whole-genome sequencing.
Well-defined gene panels
The notion of exome vs. whole-genome sequencing is somewhat contrived. The exome may be small compared with the whole genome, but it is still quite large. Physicians and drug companies are more interested in yes/no answers, which is why genomic panels, comprised of anywhere from a dozen to several hundred genes, are where the future of medical sequencing lies.
“A wide range of panels is possible, consisting of genes specifically targeting specific diseases,” says Jeffrey Chu, Ph.D., director of NGS at Applied Biological Materials. “The number of genes varies between panels and may focus on a specific region of the gene, for example only SNP loci. The idea is to perform very deep sequencing on a small number of genes, with much lower cost than whole-genome sequencing.”
In this model of medical sequencing, gene panel suppliers will compete based on their products’ specificity and sensitivity. “Some panels will have 20 genes, others 25 genes,” Chu says. “The up-front cost of designing and validating such panels is high, but the per-test costs can be quite low in the long run.”
According to Chu, the falling cost of sequencing generally, which will close the cost gap between WGS and exome sequencing, is not a good reason to prefer WGS. “With WGS you will get a lot of extra data, for example intergenic regions and intronic regions, which may or may not offer insight into the answer you’re looking for. We still don’t know the function of much of the genome.”
With the exome, at least we know that the gene codes for a protein, or has some regulatory function, such that “the data you sequence will be the most relevant.”
Prof. Mick Watson of the Roslin Institute, University of Edinburgh, also believes that the WGS vs. exome sequencing conundrum will not be settled on the basis of cost-per-base alone. “The cost of analyzing and interpreting a genome is far higher than for an exome. And, the first thing you do with whole-genome data is to look at the exome, because the impact of variants is way easier to interpret inside the protein-coding exome than outside of it. “
Furthermore, other identifiable genomic categories exist within the genome. But, as Watson notes, “these regions are harder to interpret. There are promoters, enhancers, regions that are bound to or are in contact with one another in 3D space, and others. We know that WGS hits are enriched in regulatory, non-protein-coding parts of the genome. However, interpreting the impact of a variant in a non-coding area of the genome is far harder than the coding regions, because we know the genetic code.”
Watson believes that gene panels have a lot of life left in them, and may play a special role in future sequencing projects. “Panels are cheap and easy to interpret. Whole genomes may become cheap too, but the data storage and analysis is expensive. I think we’ll see analysis of whole genomes used to identify smaller panels, which we will use in much more innovative ways, e.g., sequence them on a smartphone.”