{"doi":"10.1101/267401","title":"Rapid low-cost assembly of the\n                  <i>Drosophila melanogaster</i>\n                  reference genome using low-coverage, long-read sequencing","abstract":"<jats:title>ABSTRACT</jats:title>\n                <jats:p>\n                  Accurate and comprehensive characterization of genetic variation is essential for deciphering the genetic basis of diseases and other phenotypes. A vast amount of genetic variation stems from large-scale sequence changes arising from the duplication, deletion, inversion, and translocation of sequences. In the past 10 years, high-throughput short reads have greatly expanded our ability to assay sequence variation due to single nucleotide polymorphisms. However, a recent\n                  <jats:italic>de novo</jats:italic>\n                  assembly of a second\n                  <jats:italic>Drosophila melanogaster</jats:italic>\n                  reference genome has revealed that short read genotyping methods miss hundreds of structural variants, including those affecting phenotypes. While genomes assembled using high-coverage long reads can achieve high levels of contiguity and completeness, concerns about cost, errors, and low yield have limited widespread adoption of such sequencing approaches. Here we resequenced the reference strain of\n                  <jats:italic>D. melanogaster</jats:italic>\n                  (ISO1) on a single Oxford Nanopore MinION flow cell run for 24 hours. Using only reads longer than 1 kb or with at least 30x coverage, we assembled a highly contiguous\n                  <jats:italic>de novo</jats:italic>\n                  genome. The addition of inexpensive paired reads and subsequent scaffolding using an optical map technology achieved an assembly with completeness and contiguity comparable to the\n                  <jats:italic>D. melanogaster</jats:italic>\n                  reference assembly. Comparison of our assembly to the reference assembly of ISO1 uncovered a number of structural variants (SVs), including novel LTR transposable element insertions and duplications affecting genes with developmental, behavioral, and metabolic functions. Collectively, these SVs provide a snapshot of the dynamics of genome evolution. Furthermore, our assembly and comparison to the\n                  <jats:italic>D. melanogaster</jats:italic>\n                  reference genome demonstrates that high-quality\n                  <jats:italic>de novo</jats:italic>\n                  assembly of reference genomes and comprehensive variant discovery using such assemblies are now possible by a single lab for under $1,000 (USD).\n                </jats:p>","journal":null,"year":null,"id":589179,"datarank":0.46515015740656446,"base_score":2.302585092994046,"endowment":2.302585092994046,"self_citation_contribution":0.3453877639491069,"citation_network_contribution":0.11976239345745755,"self_endowment_contribution":0.3453877639491069,"citer_contribution":0.11976239345745755,"corpus_percentile":null,"corpus_rank":null,"citation_count":9,"citer_count":6,"citers_with_citation_signal":5,"citers_with_endowment":5,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":557278,"name":"Mahul Chakraborty","orcid":"0000-0003-2414-9187","position":1,"is_corresponding":false},{"id":24426,"name":"Danny E. Miller","orcid":"0000-0001-6096-8601","position":2,"is_corresponding":false},{"id":1055521,"name":"Shannon Kalsow","orcid":null,"position":3,"is_corresponding":false},{"id":472219,"name":"Kate Hall","orcid":null,"position":4,"is_corresponding":false},{"id":1507429,"name":"Anoja G. Perera","orcid":null,"position":5,"is_corresponding":false},{"id":1507430,"name":"J.J. Emerson","orcid":null,"position":6,"is_corresponding":false},{"id":202767,"name":"R. Scott Hawley","orcid":null,"position":7,"is_corresponding":false},{"id":762737,"name":"Edwin A. Solares","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Rapid low-cost assembly of the\n                  <i>Drosophila melanogaster</i>\n                  reference genome using low-coverage, long-read sequencing","abstract":"<jats:title>ABSTRACT</jats:title>\n                <jats:p>\n                  Accurate and comprehensive characterization of genetic variation is essential for deciphering the genetic basis of diseases and other phenotypes. A vast amount of genetic variation stems from large-scale sequence changes arising from the duplication, deletion, inversion, and translocation of sequences. In the past 10 years, high-throughput short reads have greatly expanded our ability to assay sequence variation due to single nucleotide polymorphisms. However, a recent\n                  <jats:italic>de novo</jats:italic>\n                  assembly of a second\n                  <jats:italic>Drosophila melanogaster</jats:italic>\n                  reference genome has revealed that short read genotyping methods miss hundreds of structural variants, including those affecting phenotypes. While genomes assembled using high-coverage long reads can achieve high levels of contiguity and completeness, concerns about cost, errors, and low yield have limited widespread adoption of such sequencing approaches. Here we resequenced the reference strain of\n                  <jats:italic>D. melanogaster</jats:italic>\n                  (ISO1) on a single Oxford Nanopore MinION flow cell run for 24 hours. Using only reads longer than 1 kb or with at least 30x coverage, we assembled a highly contiguous\n                  <jats:italic>de novo</jats:italic>\n                  genome. The addition of inexpensive paired reads and subsequent scaffolding using an optical map technology achieved an assembly with completeness and contiguity comparable to the\n                  <jats:italic>D. melanogaster</jats:italic>\n                  reference assembly. Comparison of our assembly to the reference assembly of ISO1 uncovered a number of structural variants (SVs), including novel LTR transposable element insertions and duplications affecting genes with developmental, behavioral, and metabolic functions. Collectively, these SVs provide a snapshot of the dynamics of genome evolution. Furthermore, our assembly and comparison to the\n                  <jats:italic>D. melanogaster</jats:italic>\n                  reference genome demonstrates that high-quality\n                  <jats:italic>de novo</jats:italic>\n                  assembly of reference genomes and comprehensive variant discovery using such assemblies are now possible by a single lab for under $1,000 (USD).\n                </jats:p>","is_dataset_classified":null,"base_score":2.302585092994046,"endowment":2.302585092994046,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"20725694","pmcid":null,"openalex_id":"https://openalex.org/W2949865853","authors":[],"funders":[{"funder_name":"National Institutes of Health","grant_id":"1S10OD010794-01","title":"PacBio RS Single Molecule, Real-Time (SMRT) DNA Sequencer"},{"funder_name":"National Institutes of Health","grant_id":"1S10RR025496-01","title":"High Throughput DNA Sequencer"},{"funder_name":"National Science Foundation","grant_id":"1321846","title":"Graduate Research Fellowship Program (GRFP)"},{"funder_name":"National Institutes of Health","grant_id":"1S10OD021718-01","title":"High-Throughput DNA Sequencer"},{"funder_name":"National Institutes of Health","grant_id":"5P30CA062203-08","title":"CORE--BIOSTATISTICS FACILITY"},{"funder_name":"National Institutes of Health","grant_id":"1R01RR024862-01A1","title":"A resource for the genetic analysis of complex traits"},{"funder_name":"National Institutes of Health","grant_id":"3P30CA062203-27S1","title":"Clinical Trial Advancement and Development for Sarcomas"},{"funder_name":"National Institutes of Health","grant_id":"5R01GM123303-02","title":"Structural variation population size and the evolution of genome complexity"},{"funder_name":"National Science Foundation","grant_id":"1500284","title":"California LSAMP Bridge to the Doctorate Program at the University of California-Irvine (2015-2017)"}],"total_grants":9,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2018,"count":1},{"year":2019,"count":3},{"year":2020,"count":2},{"year":2021,"count":3}],"oa_status":"green","license":"cc-by-nc-nd","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2018/06/09/267401.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2018/06/09/267401.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/267401","host_type":"publisher"},{"url":"https://doi.org/10.1101/267401","host_type":"repository"},{"url":"https://doi.org/10.1534/g3.118.200162","host_type":""},{"url":"https://academic.oup.com/g3journal/article-pdf/8/10/3143/37190314/g3journal3143.pdf","host_type":""},{"url":"https://pubmed.ncbi.nlm.nih.gov/30018084","host_type":""},{"url":"http://dx.doi.org/10.1534/g3.118.200162","host_type":""},{"url":"https://doaj.org/article/824ecb00c58c4590a073329cb92a7ec3","host_type":""},{"url":"https://dx.doi.org/10.1534/g3.118.200162","host_type":""},{"url":"https://dx.doi.org/10.1101/267401","host_type":""},{"url":"https://escholarship.org/content/qt8bp73388/qt8bp73388.pdf","host_type":""},{"url":"https://escholarship.org/uc/item/8bp73388","host_type":""},{"url":"http://dx.doi.org/10.1101/267401","host_type":""},{"url":"https://doi.org/https://doi.org/10.1534/g3.118.200162","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Chromosomal and Genetic Variations","Genetic diversity and population structure","0301 basic medicine","0206 medical engineering","02 engineering and technology","03 medical and health sciences"],"mesh_terms":[],"keywords":["Sequence assembly","Reference genome","Genome","Biology","Structural variation","Drosophila melanogaster","Genetics","Minion","Computational biology","Nanopore sequencing","Transposable element","Copy-number variation","Whole genome sequencing","Gene","Transcriptome","570","1.1 Normal biological development and functioning","Genome, Insect","Bioinformatics and Computational Biology","QH426-470","Investigations","2.1 Biological and endogenous factors","Animals","Gene Library","Human Genome","Statistics","Genetic Variation","High-Throughput Nucleotide Sequencing","Molecular Sequence Annotation","DNA","Genomics","Sequence Analysis, DNA","Biological Sciences","Mitochondrial","Biochemistry and cell biology","Genome, Mitochondrial","genome assembly","Drosophila","Generic health relevance","Insect","Sequence Analysis","Biotechnology"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-23T14:16:04.706439Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}