De Novo Sequencing: Assembling a Genome from Scratch

BlueskyReddit
 De Novo Sequencing
Jeffrey Perkel has been a scientific writer and editor since 2000. He holds a PhD in Cell and Molecular Biology from the University of Pennsylvania, and did postdoctoral work at the University of Pennsylvania and at Harvard Medical School.

There’s no denying the power of genome sequencing has entered the public’s imagination, especially in the clinical arena.

growing number of news reports attest to the power of genomic technologies to crack medical mysteries, including the Milwaukee Journal Sentinel’s extraordinary Pulitzer Prize-winning story of a child with a devastating gastrointestinal condition who received life-altering treatment thanks to exome sequencing.

But here’s the thing about most of these anecdotes: They all depend upon comparison of a sample to a reference genome. Think about it. How do you know what mutations or variants are significant? Compare the sequence to a “normal” control. 

Sometimes, though, especially in the case of sparsely studied organisms, there is no reference, no scaffold upon which to pin sequencing reads. In such a situation, genome assemblies must be built “de novo,” that is, from scratch.

Think of it as a jigsaw puzzle. Sequencing projects in which a reference is available are like puzzles for which a box-top picture of the final design is available; researchers need only compare their pieces to the scaffold to see where they fit. In a de novo project there is no such picture, just millions upon millions of pieces. (Though researchers can sometimes circumvent this problem by aligning their reads to a closely related organism.)

“The sequencing itself isn’t difficult,” explains Todd Arnold, head of research and development at 454 Life Sciences (part of Roche Sequencing Solutions). But figuring out how the pieces fit together, especially in the case of repetitive elements, can be enormously challenging. “The challenge with de novo assembly [is that] you don’t have anything to refer back to,” he says.

Size matters

As with a physical puzzle, when it comes to genome assembly the big pieces are easier to assemble than the small ones, and therein lies the problem. The capillary electrophoresis-based sequencers that produced the first human genome produced relatively long reads of 700 to 1,000 bases or so, long enough to span most repetitive elements. Yet today, genome sequencing with such hardware is virtually unheard of. Next-gen instruments typically generate substantially shorter reads—about 400 bases on Life Technologies’ Ion Torrent PGM™ and up to 2 x 150 paired-end bases on Illumina’s HiSeq® 2500—and as a result, many cannot be mapped unambiguously. (Life Technologies' newer Ion Proton™ instrument supports 200 bp reads, but more of them, for more than 10 Gb per run on the current PI™ chip.)

Some next-gen technologies provide longer reads. Pacific Biosciences reads “routinely” exceed 10 kb with the company’s XL chemistry, according to a 2013 poster presentation, for instance. 454 Life Sciences’ reads average 700 bases and can extend to a kilobase each, with the updated software launched at the end of 2012 being "the first offering up to one million reads with read lengths and accuracy that are comparable to traditional Sanger-based methods," Arnold says.

On the other hand, though, 454's GS FLX+ generates far fewer reads than, say, an Illumina sequencer -- one million reads versus up to 3 billion on a HiSeq 2500 -- making it ideal for smaller genomes such as bacteria, viruses, and small eukaryotic genomes, as well as a source of long reads for complex eukaryote de novo sequencing projects.

One popular strategy is to combine platforms to leverage the strengths of both short and long reads. Michael Freitag, associate professor in the Department of Biochemistry and Biophysics at Oregon State University, with colleagues in Germany published in 2010 one of the first de novo assemblies of a eukaryotic genome from short, next-gen sequence reads; it was that of the fungus, Sordaria macrospora [1]. That genome weighed in at just 40 Mb on seven chromosomes and had, it turns out, relatively little repeated sequence. “We were a little lucky,” Freitag says. Still, to complete the project the team had to combine reads from both 454 and Illumina, using the long reads of the former to scaffold the shorter Illumina data. In the end, the team ended up with 4,781 contigs, the largest of which was 2.5 Mb.

De novo libraries

Another approach to effectively increase read length is to generate pairs of relatively short reads with a defined distance between them. For instance, in so-called paired-end sequencing bases are read from both ends of a long fragment, but because the average length of the total fragment is known (libraries are built using size-selected DNA), the distance between those reads is also known, making it easier to build a high-quality assembly.

A complementary strategy is mate-pair sequencing. In this case genomic fragments thousands of bases long are circularized by joining the two ends. Then a region around that ligated junction is cleaved, purified and sequenced from either end, again providing both sequence and scaffolding data.

Mate-pair libraries can be of varying length, and typically more than one is needed. The Broad Institute’s ALLPATHS-LG assembler, for instance, requires at least two libraries, one short, such that the two paired reads overlap (e.g., 100 bp from either end of a 180-bp fragment), and one long (about 3,000 bp).

“The larger you can make the pieces, the easier it is to put together the puzzle,” explains Arnold.

Steven Jones, associate director of the Genome Sciences Center at the British Columbia Cancer Agency, has assembled several genomes de novo. In 2009, he combined capillary-electrophoresis, 454 and Illumina data to assemble the 32.5 Mb filamentous fungus Grosmannia clavigera [2]. More recently, he assembled the 20-Gb genome of the white spruce (Picea glauca) entirely from Illumina reads, thanks both to a new assembler, longer sequences and new mate-pair libraries [3].

“The Illumina Mi-Seq platform is giving us 500 base-pair reads, and there is good technology now to get mate-pair libraries for this platform, which can provide long-range contiguity of the assemblies,” Jones says.

Mate-pair library preparation kits are available from Illumina, Life Technologies and Lucigen.

Another option for long reads is the Moleculo technology Illumina acquired in 2012. Basically, long fragments (~10 kb) are separated, fragmented and tagged so that all the fragments from one long segment carry the same barcode. Then the pieces are combined and sequenced, with the barcodes used to confidently assign reads to longer scaffolds.

According to data presented at the Plants and Animal Genomes Conference earlier this year, reads using Moleculo technology average eight to 10 kb in length. Kits are not yet available; in a January 2013 interview with Bio-IT World, company cofounder Mickey Kertesz said the technology currently is offered only as a service. (Illumina declined to be interviewed for this article).

Assembly tools

The other key tool you’ll need for de novo genome assembly is an assembler, the software tool that pieces the sequencing reads together into contigs. Several are available, including ALLPATHS-LG, Velvet (used by both Freitag and Jones), AbySS (used to assemble the white-spruce genome) and SOAPdenovo (from BGI). Ion Torrent’s sequencing tools include the open-source assembler MIRA3, says Andy Felton, vice president of product management for Ion Torrent at Life Technologies.

According to Felton, bioinformatics analysis is the most difficult step in any next-gen sequencing project, and that’s particularly true for de novo assembly. New users may struggle with metrics such as N50 (a quality score), the number of gaps and contigs and especially with how tweaking program variables can change those values.

“Understanding what the metrics are telling you about the genome and manipulating the data to get the best quality assembly are the two areas that probably take the most to learn,” Felton says.

Ultimately, each project is different, and only you can determine the parameters required to complete a project successfully. Beth Shapiro, associate professor of ecology and evolutionary biology at the University of California, Santa Cruz, and an advisory board member of the Genome 10K project (an attempt to sequence the genomes of 10,000 vertebrate species) has a number of genomes on her plate, including some 20 mammals and the passenger pigeon.

That latter genome is a bit problematic. The passenger pigeon is extinct, and samples are hard to come by. The best Shapiro can hope for in terms of length, she says, is about 100 bp. As a result, she is also sequencing a closely related (and extant) species, the band-tailed pigeon (genome size ~2 to 3 Gb).

“We’re doing de novo assembly of the band-tailed pigeon and using that as a scaffold to align the passenger pigeon,” Shapiro explains.

As for Freitag, he says the world of genome assembly is completely different from when he started the Sordaria project. Between mate-pair libraries and longer reads, he says, “We could do a much better assembly for about 10% of the money.”

References
[1] Nowrousian, M, et al., “De novo assembly of a 40 Mb eukaryotic genome from short sequence reads: Sordaria macrospora, a model organism for fungal morphogenesis,” PLoS Genet, 6[4]:e1000891, 2010.

[2] DiGuistini, S, et al., “De novo genome sequence assembly of a filamentous fungus using Sanger, 454 and Illumina sequence data,” Genome Biol, 10:R94, 2009.

[3] Birol, I, et al., “Assembling the 20 Gb white spruce (Picea glauca) genome from whole-genome shotgun sequencing data,” Bioinformatics, 29:1492-7, 2013.

Image: Life Technologies Ion 318™ chip for the Ion PGM.

  • <<
  • >>

Join the discussion