{"doi":"10.1111/1755-0998.13455","title":"Novel quality metrics allow identifying and generating high‐quality assemblies of piRNA clusters","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:p>\n                    In most animals, it is thought that the proliferation of a transposable element (TE) is stopped when the TE jumps into a piRNA cluster. Despite this central importance, little is known about the composition and the evolutionary dynamics of piRNA clusters. This is largely because piRNA clusters are notoriously difficult to assemble as they are frequently composed of highly repetitive DNA. With long reads, we may finally be able to obtain reliable assemblies of piRNA clusters. Unfortunately, it is unclear how to generate and identify the best assemblies, as many assembly strategies exist and standard quality metrics are ignorant of TEs. To address these problems, we introduce several novel quality metrics that assess: (a) the fraction of completely assembled piRNA clusters, (b) the quality of the assembled clusters and (c) whether an assembly captures the overall TE landscape of an organisms (i.e. the abundance, the number of SNPs and internal deletions of all TE families). The requirements for computing these metrics vary, ranging from annotations of piRNA clusters to consensus sequences of TEs and genomic sequencing data. Using these novel metrics, we evaluate the effect of assembly algorithm, polishing, read length, coverage, residual polymorphisms and finally identify strategies that yield reliable assemblies of piRNA clusters. Based on an optimized approach, we provide assemblies for the two\n                    <jats:italic>Drosophila melanogaster</jats:italic>\n                    strains Canton‐S and Pi2. About 80% of known piRNA clusters were assembled in both strains. Finally, we demonstrate the generality of our approach by extending our metrics to humans and\n                    <jats:italic>Arabidopsis thaliana</jats:italic>\n                    .\n                  </jats:p>","journal":"Molecular Ecology Resources","year":2022,"id":597664,"datarank":0.5101796072493234,"base_score":3.4011973816621555,"endowment":3.4011973816621555,"self_citation_contribution":0.5101796072493234,"citation_network_contribution":0.0,"self_endowment_contribution":0.5101796072493234,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":29,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":2,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":855895,"name":"Florian Schwarz","orcid":"0000-0002-1088-4773","position":1,"is_corresponding":false},{"id":1531186,"name":"Odontsetseg Cannalonga","orcid":null,"position":2,"is_corresponding":false},{"id":997322,"name":"Robert Kofler","orcid":"0000-0001-9960-7248","position":3,"is_corresponding":false},{"id":997321,"name":"Filip Wierzbicki","orcid":"0000-0002-6171-2461","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Novel quality metrics allow identifying and generating high‐quality assemblies of piRNA clusters","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:p>\n                    In most animals, it is thought that the proliferation of a transposable element (TE) is stopped when the TE jumps into a piRNA cluster. Despite this central importance, little is known about the composition and the evolutionary dynamics of piRNA clusters. This is largely because piRNA clusters are notoriously difficult to assemble as they are frequently composed of highly repetitive DNA. With long reads, we may finally be able to obtain reliable assemblies of piRNA clusters. Unfortunately, it is unclear how to generate and identify the best assemblies, as many assembly strategies exist and standard quality metrics are ignorant of TEs. To address these problems, we introduce several novel quality metrics that assess: (a) the fraction of completely assembled piRNA clusters, (b) the quality of the assembled clusters and (c) whether an assembly captures the overall TE landscape of an organisms (i.e. the abundance, the number of SNPs and internal deletions of all TE families). The requirements for computing these metrics vary, ranging from annotations of piRNA clusters to consensus sequences of TEs and genomic sequencing data. Using these novel metrics, we evaluate the effect of assembly algorithm, polishing, read length, coverage, residual polymorphisms and finally identify strategies that yield reliable assemblies of piRNA clusters. Based on an optimized approach, we provide assemblies for the two\n                    <jats:italic>Drosophila melanogaster</jats:italic>\n                    strains Canton‐S and Pi2. About 80% of known piRNA clusters were assembled in both strains. Finally, we demonstrate the generality of our approach by extending our metrics to humans and\n                    <jats:italic>Arabidopsis thaliana</jats:italic>\n                    .\n                  </jats:p>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":2,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"34181811","pmcid":null,"openalex_id":null,"authors":[],"funders":[{"funder_name":"Austrian Science Fund","grant_id":"P30036‐B25 and W1225","title":null},{"funder_name":"Austrian Science Fund FWF","grant_id":"P 30036","title":null},{"funder_name":"Austrian Science Fund FWF","grant_id":"W 1225","title":null},{"funder_name":"Austrian Science Fund FWF","grant_id":"P30036-B25 and W1225","title":null}],"total_grants":4,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[],"oa_status":"hybrid","license":"cc-by","oa_locations":[{"url":"https://doi.org/10.1111/1755-0998.13455","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/pdf/10.1111/1755-0998.13455","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/full-xml/10.1111/1755-0998.13455","host_type":"publisher"}],"fields_of_study":[],"mesh_terms":["Animals","Humans","Drosophila melanogaster","Arabidopsis","RNA, Small Interfering","Genomics"],"keywords":["Drosophila melanogaster","Transposable elements","Genome Assembly","Pirna Clusters","Oxford Nanopore Sequencing"],"sdg_mappings":[],"linked_datasets":[{"doi":"10.6084/m9.figshare.24406636.v1","title":"Additional file 1 of The composition of piRNA clusters in Drosophila melanogaster deviates from expectations under the trap model","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.24406636","title":"Additional file 1 of The composition of piRNA clusters in Drosophila melanogaster deviates from expectations under the trap model","publisher":"figshare","resource_type":"JournalArticle"}],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-28T13:53:03.567395Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}