{"doi":"10.1101/gr.276015.121","title":"Long-read sequencing of 111 rice genomes reveals significantly larger pan-genomes","abstract":"<jats:p>The concept of pan-genome, which is the collection of all genomes from a population, has shown a great potential in genomics study, especially for crop sciences. The rice pan-genome constructed from the second-generation sequencing (SGS) data is about 270 Mb larger than <jats:italic>Nipponbare</jats:italic>, the rice reference genome (NipRG), but it is still disadvantaged by incompleteness and loss of genomic contexts. The third-generation sequencing (TGS) with long reads can help to construct better pan-genomes. In this paper, we report a high-quality rice pan-genome construction method by introducing a series of new steps to deal with the long-read data, including unmapped sequence block filtering, redundancy removing, and sequence block elongating. Compared to NipRG, the long-read sequencing-based pan-genome constructed from 105 rice accessions, which contains 604 Mb novel sequences, is much more comprehensive than the one constructed from ∼3000 rice genomes sequenced with short reads. The repetitive sequences are the main components of novel sequences, which partially explain the differences between the pan-genomes based on TGS and SGS. Adding six wild rice accessions, there are about 879 Mb novel sequences and 19,000 novel genes in the rice pan-genome in total. In addition, we have created high-quality reference genomes for all representative rice populations, including five gapless reference genomes. This study has made significant progress in our understanding of the rice pan-genome, and this pan-genome construction method for long-read data can be applied to accelerate a broad range of genomics studies.</jats:p>","journal":"Genome Research","year":2022,"id":605695,"datarank":0.6681520944380261,"base_score":4.454347296253507,"endowment":4.454347296253507,"self_citation_contribution":0.6681520944380261,"citation_network_contribution":0.0,"self_endowment_contribution":0.6681520944380261,"citer_contribution":0.0,"corpus_percentile":70.1,"corpus_rank":3930,"citation_count":85,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1554573,"name":"Hongzhang Xue","orcid":null,"position":1,"is_corresponding":false},{"id":1554574,"name":"Xiaorui Dong","orcid":null,"position":2,"is_corresponding":false},{"id":896453,"name":"Min Li","orcid":"0000-0001-8803-3162","position":3,"is_corresponding":false},{"id":1554575,"name":"Xiaoming Zheng","orcid":null,"position":4,"is_corresponding":false},{"id":241773,"name":"Zhikang Li","orcid":"0000-0001-7017-0097","position":5,"is_corresponding":false},{"id":1554576,"name":"Jianlong Xu","orcid":null,"position":6,"is_corresponding":false},{"id":816055,"name":"Wensheng Wang","orcid":"0000-0002-8842-3432","position":7,"is_corresponding":false},{"id":1543354,"name":"Chaochun Wei","orcid":"0000-0002-1031-034X","position":8,"is_corresponding":false},{"id":247505,"name":"Fan Zhang","orcid":"0000-0003-2849-1713","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Long-read sequencing of 111 rice genomes reveals significantly larger pan-genomes","abstract":"<jats:p>The concept of pan-genome, which is the collection of all genomes from a population, has shown a great potential in genomics study, especially for crop sciences. The rice pan-genome constructed from the second-generation sequencing (SGS) data is about 270 Mb larger than <jats:italic>Nipponbare</jats:italic>, the rice reference genome (NipRG), but it is still disadvantaged by incompleteness and loss of genomic contexts. The third-generation sequencing (TGS) with long reads can help to construct better pan-genomes. In this paper, we report a high-quality rice pan-genome construction method by introducing a series of new steps to deal with the long-read data, including unmapped sequence block filtering, redundancy removing, and sequence block elongating. Compared to NipRG, the long-read sequencing-based pan-genome constructed from 105 rice accessions, which contains 604 Mb novel sequences, is much more comprehensive than the one constructed from ∼3000 rice genomes sequenced with short reads. The repetitive sequences are the main components of novel sequences, which partially explain the differences between the pan-genomes based on TGS and SGS. Adding six wild rice accessions, there are about 879 Mb novel sequences and 19,000 novel genes in the rice pan-genome in total. In addition, we have created high-quality reference genomes for all representative rice populations, including five gapless reference genomes. This study has made significant progress in our understanding of the rice pan-genome, and this pan-genome construction method for long-read data can be applied to accelerate a broad range of genomics studies.</jats:p>","is_dataset_classified":null,"base_score":4.454347296253507,"endowment":4.454347296253507,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"35396275","pmcid":"PMC9104699","openalex_id":"https://openalex.org/W4223446368","authors":[],"funders":[{"funder_name":"Natural Science Foundation of Shanghai","grant_id":"20ZR1428200","title":null},{"funder_name":"National Natural Science Foundation of China","grant_id":"32170643","title":null},{"funder_name":"National Natural Science Foundation of China","grant_id":"31971928","title":null},{"funder_name":"National Natural Science Foundation of China","grant_id":"U21A20214","title":null},{"funder_name":"the Hainan Provincial Joint Project of Sanya Yazhou Bay Science and Technology City","grant_id":"320LH044","title":null},{"funder_name":"the Hainan Yazhou Bay Seed Lab Project","grant_id":"B21HJ0215","title":null},{"funder_name":"the Hainan Yazhou Bay Seed Lab Project","grant_id":"B21HJ0223","title":null},{"funder_name":"the Hainan Yazhou Bay Seed Lab Project","grant_id":"B21HJ0508","title":null},{"funder_name":"the Agricultural Science and Technology Innovation Program and the Cooperation and Innovation Mission","grant_id":"CAAS-ZDXT202001","title":null},{"funder_name":"SJTU JiRLMDS Joint Research Fund","grant_id":"MDS-JF-2019A07","title":null},{"funder_name":"the National High-level Personnel of Special Support Program","grant_id":"","title":null},{"funder_name":"the","grant_id":"","title":null},{"funder_name":"the CAAS Innovative Team Award","grant_id":"","title":null}],"total_grants":13,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2022,"count":13},{"year":2023,"count":20},{"year":2024,"count":24},{"year":2025,"count":25},{"year":2026,"count":3}],"oa_status":"bronze","license":null,"oa_locations":[{"url":"https://genome.cshlp.org/content/32/5/853.full.pdf","host_type":"journal"},{"url":"https://genome.cshlp.org/content/32/5/853.full.pdf","host_type":"publisher"},{"url":"https://syndication.highwire.org/content/doi/10.1101/gr.276015.121","host_type":"publisher"},{"url":"https://doi.org/10.1101/gr.276015.121","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/35396275","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/9104699","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC9104699","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC9104699?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Genomics and Phylogenetic Studies","Chromosomal and Genetic Variations","Plant Disease Resistance and Genetics"],"mesh_terms":["Oryza","Genome","Sequence Analysis, DNA","Genomics","High-Throughput Nucleotide Sequencing"],"keywords":["Genome","Biology","Reference genome","Genomics","Genetics","DNA sequencing","Whole genome sequencing","Computational biology","Population","Gene"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-30T03:02:21.500305Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}