Sie befinden Sich nicht im Netzwerk der Universität Paderborn. Der Zugriff auf elektronische Ressourcen ist gegebenenfalls nur via VPN oder Shibboleth (DFN-AAI) möglich. mehr Informationen...
Ergebnis 4 von 21

Details

Autor(en) / Beteiligte
Titel
GRASS: a generic algorithm for scaffolding next-generation sequencing assemblies
Ist Teil von
  • Bioinformatics (Oxford, England), 2012-06, Vol.28 (11), p.1429-1437
Ort / Verlag
Oxford: Oxford University Press
Erscheinungsjahr
2012
Link zum Volltext
Quelle
MEDLINE
Beschreibungen/Notizen
  • The increasing availability of second-generation high-throughput sequencing (HTS) technologies has sparked a growing interest in de novo genome sequencing. This in turn has fueled the need for reliable means of obtaining high-quality draft genomes from short-read sequencing data. The millions of reads usually involved in HTS experiments are first assembled into longer fragments called contigs, which are then scaffolded, i.e. ordered and oriented using additional information, to produce even longer sequences called scaffolds. Most existing scaffolders of HTS genome assemblies are not suited for using information other than paired reads to perform scaffolding. They use this limited information to construct scaffolds, often preferring scaffold length over accuracy, when faced with the tradeoff. We present GRASS (GeneRic ASsembly Scaffolder)-a novel algorithm for scaffolding second-generation sequencing assemblies capable of using diverse information sources. GRASS offers a mixed-integer programming formulation of the contig scaffolding problem, which combines contig order, distance and orientation in a single optimization objective. The resulting optimization problem is solved using an expectation-maximization procedure and an unconstrained binary quadratic programming approximation of the original problem. We compared GRASS with existing HTS scaffolders using Illumina paired reads of three bacterial genomes. Our algorithm constructs a comparable number of scaffolds, but makes fewer errors. This result is further improved when additional data, in the form of related genome sequences, are used. GRASS source code is freely available from http://code.google.com/p/tud-scaffolding/. Supplementary data are available at Bioinformatics online.

Weiterführende Literatur

Empfehlungen zum selben Thema automatisch vorgeschlagen von bX