{"doi":"10.1101/2020.05.04.077180","title":"Highly accurate long-read HiFi sequencing data for five complex genomes","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  The PacBio\n                  <jats:sup>®</jats:sup>\n                  HiFi sequencing method yields highly accurate long-read sequencing datasets with read lengths averaging 10-25 kb and accuracies greater than 99.5%. These accurate long reads can be used to improve results for complex applications such as single nucleotide and structural variant detection, genome assembly, assembly of difficult polyploid or highly repetitive genomes, and assembly of metagenomes. Currently, there is a need for sample data sets to both evaluate the benefits of these long accurate reads as well as for development of bioinformatic tools including genome assemblers, variant callers, and haplotyping algorithms. We present deep coverage HiFi datasets for five complex samples including the two inbred model genomes\n                  <jats:italic>Mus musculus</jats:italic>\n                  and\n                  <jats:italic>Zea mays</jats:italic>\n                  , as well as two complex genomes, octoploid\n                  <jats:italic>Fragaria</jats:italic>\n                  ×\n                  <jats:italic>ananassa</jats:italic>\n                  and the diploid anuran\n                  <jats:italic>Rana muscosa</jats:italic>\n                  . Additionally, we release sequence data from a mock metagenome community. The datasets reported here can be used without restriction to develop new algorithms and explore complex genome structure and evolution. Data were generated on the PacBio Sequel II System.\n                </jats:p>","journal":null,"year":null,"id":596359,"datarank":0.5837730447165941,"base_score":3.8918202981106265,"endowment":3.8918202981106265,"self_citation_contribution":0.5837730447165941,"citation_network_contribution":0.0,"self_endowment_contribution":0.5837730447165941,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":48,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1527331,"name":"Kristin Mars","orcid":null,"position":1,"is_corresponding":false},{"id":1527332,"name":"Greg Young","orcid":null,"position":2,"is_corresponding":false},{"id":350867,"name":"Yu‐Chih Tsai","orcid":"0000-0002-2958-0278","position":3,"is_corresponding":false},{"id":1527333,"name":"Joseph W. Karalius","orcid":"0000-0003-3592-1339","position":4,"is_corresponding":false},{"id":1527334,"name":"Jane M. Landolin","orcid":null,"position":5,"is_corresponding":false},{"id":109456,"name":"Nicholas Maurer","orcid":"0000-0003-0007-7887","position":6,"is_corresponding":false},{"id":615106,"name":"David Kudrna","orcid":"0000-0002-3092-3629","position":7,"is_corresponding":false},{"id":1527335,"name":"Michael A. Hardigan","orcid":null,"position":8,"is_corresponding":false},{"id":51363,"name":"Cynthia Steiner","orcid":"0000-0002-2131-8072","position":9,"is_corresponding":false},{"id":1297,"name":"Steven J. Knapp","orcid":"0000-0001-6498-5409","position":10,"is_corresponding":false},{"id":265747,"name":"Doreen Ware","orcid":"0000-0002-8125-3821","position":11,"is_corresponding":false},{"id":226135,"name":"Beth Shapiro","orcid":"0000-0002-2733-7776","position":12,"is_corresponding":false},{"id":2120,"name":"Paul Peluso","orcid":"0000-0002-9723-5185","position":13,"is_corresponding":false},{"id":60498,"name":"David R. Rank","orcid":"0000-0001-9213-6965","position":14,"is_corresponding":false},{"id":1527330,"name":"Ting Hon","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Highly accurate long-read HiFi sequencing data for five complex genomes","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  The PacBio\n                  <jats:sup>®</jats:sup>\n                  HiFi sequencing method yields highly accurate long-read sequencing datasets with read lengths averaging 10-25 kb and accuracies greater than 99.5%. These accurate long reads can be used to improve results for complex applications such as single nucleotide and structural variant detection, genome assembly, assembly of difficult polyploid or highly repetitive genomes, and assembly of metagenomes. Currently, there is a need for sample data sets to both evaluate the benefits of these long accurate reads as well as for development of bioinformatic tools including genome assemblers, variant callers, and haplotyping algorithms. We present deep coverage HiFi datasets for five complex samples including the two inbred model genomes\n                  <jats:italic>Mus musculus</jats:italic>\n                  and\n                  <jats:italic>Zea mays</jats:italic>\n                  , as well as two complex genomes, octoploid\n                  <jats:italic>Fragaria</jats:italic>\n                  ×\n                  <jats:italic>ananassa</jats:italic>\n                  and the diploid anuran\n                  <jats:italic>Rana muscosa</jats:italic>\n                  . Additionally, we release sequence data from a mock metagenome community. The datasets reported here can be used without restriction to develop new algorithms and explore complex genome structure and evolution. Data were generated on the PacBio Sequel II System.\n                </jats:p>","is_dataset_classified":null,"base_score":3.8918202981106265,"endowment":3.8918202981106265,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"23304386","pmcid":null,"openalex_id":"https://openalex.org/W3021022548","authors":[],"funders":[{"funder_name":"National Science Foundation","grant_id":"1744001","title":"TRANSFORM-PGR: Whole genome assembly of the maize NAM founders"}],"total_grants":1,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2020,"count":1},{"year":2021,"count":7},{"year":2022,"count":13},{"year":2023,"count":9},{"year":2024,"count":12},{"year":2025,"count":5}],"oa_status":"green","license":"cc-by-nc-nd","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2020/05/05/2020.05.04.077180.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2020/05/05/2020.05.04.077180.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/2020.05.04.077180","host_type":"publisher"},{"url":"https://doi.org/10.1101/2020.05.04.077180","host_type":"repository"},{"url":"https://escholarship.org/uc/item/26s4r638","host_type":"repository"},{"url":"https://escholarship.org/content/qt26s4r638/qt26s4r638.pdf?t=qlqg3b","host_type":""},{"url":"https://doi.org/10.1038/s41597-020-00743-4","host_type":""},{"url":"https://www.nature.com/articles/s41597-020-00743-4.pdf","host_type":""},{"url":"https://pubmed.ncbi.nlm.nih.gov/33203859","host_type":""},{"url":"http://dx.doi.org/10.1038/s41597-020-00743-4","host_type":""},{"url":"https://dx.doi.org/10.1101/2020.05.04.077180","host_type":""},{"url":"https://dx.doi.org/10.1038/s41597-020-00743-4","host_type":""},{"url":"http://dx.doi.org/10.1101/2020.05.04.077180","host_type":""},{"url":"https://doi.org/https://doi.org/10.1038/s41597-020-00743-4","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Chromosomal and Genetic Variations","Genetic diversity and population structure","0301 basic medicine","0206 medical engineering","02 engineering and technology","03 medical and health sciences"],"mesh_terms":[],"keywords":["Genome","Metagenomics","Biology","Sequence assembly","Computational biology","DNA sequencing","Polyploid","Computer science","Genetics","Gene","Statistics and Probability","Data Descriptor","Ranidae","Library and Information Sciences","Fragaria","Zea mays","Education","Mice","Animals","Human Genome","High-Throughput Nucleotide Sequencing","Plant","DNA","Sequence Analysis, DNA","Computer Science Applications","Metagenome","Generic health relevance","Statistics, Probability and Uncertainty","Sequence Analysis","Genome, Plant","Information Systems"],"sdg_mappings":[{"sdg_number":3,"sdg_label":"3. Good health"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-28T10:00:20.725491Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}