{"doi":"10.1093/bioinformatics/btad308","title":"TRASH: Tandem Repeat Annotation and Structural Hierarchy","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>The advent of long-read DNA sequencing is allowing complete assembly of highly repetitive genomic regions for the first time, including the megabase-scale satellite repeat arrays found in many eukaryotic centromeres. The assembly of such repetitive regions creates a need for their de novo annotation, including patterns of higher order repetition. To annotate tandem repeats, methods are required that can be widely applied to diverse genome sequences, without prior knowledge of monomer sequences.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>Tandem Repeat Annotation and Structural Hierarchy (TRASH) is a tool that identifies and maps tandem repeats in nucleotide sequence, without prior knowledge of repeat composition. TRASH analyses a fasta assembly file, identifies regions occupied by repeats and then precisely maps them and their higher order structures. To demonstrate the applicability and scalability of TRASH for centromere research, we apply our method to the recently published Col-CEN genome of Arabidopsis thaliana and the complete human CHM13 genome.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>TRASH is freely available at:https://github.com/vlothec/TRASH and supported on Linux.</jats:p>\n               </jats:sec>","journal":"Bioinformatics","year":2023,"id":626305,"datarank":0.6862066467755076,"base_score":4.574710978503383,"endowment":4.574710978503383,"self_citation_contribution":0.6862066467755076,"citation_network_contribution":0.0,"self_endowment_contribution":0.6862066467755076,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":96,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1240431,"name":"Michael Hong","orcid":"0000-0001-6638-6733","position":1,"is_corresponding":false},{"id":616075,"name":"Ian R. Henderson","orcid":"0000-0001-5066-1489","position":2,"is_corresponding":false},{"id":1620067,"name":"Piotr Wlodzimierz","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"TRASH: Tandem Repeat Annotation and Structural Hierarchy","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>The advent of long-read DNA sequencing is allowing complete assembly of highly repetitive genomic regions for the first time, including the megabase-scale satellite repeat arrays found in many eukaryotic centromeres. The assembly of such repetitive regions creates a need for their de novo annotation, including patterns of higher order repetition. To annotate tandem repeats, methods are required that can be widely applied to diverse genome sequences, without prior knowledge of monomer sequences.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>Tandem Repeat Annotation and Structural Hierarchy (TRASH) is a tool that identifies and maps tandem repeats in nucleotide sequence, without prior knowledge of repeat composition. TRASH analyses a fasta assembly file, identifies regions occupied by repeats and then precisely maps them and their higher order structures. To demonstrate the applicability and scalability of TRASH for centromere research, we apply our method to the recently published Col-CEN genome of Arabidopsis thaliana and the complete human CHM13 genome.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>TRASH is freely available at:https://github.com/vlothec/TRASH and supported on Linux.</jats:p>\n               </jats:sec>","is_dataset_classified":null,"base_score":4.532599493153256,"endowment":4.532599493153256,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"37162382","pmcid":"PMC10199239","openalex_id":"https://openalex.org/W4376121976","authors":[],"funders":[{"funder_name":"Biotechnology and Biological Sciences Research Council","grant_id":"BB/S006842/1","title":null},{"funder_name":"Biotechnology and Biological Sciences Research Council","grant_id":"BB/S020012/1","title":null},{"funder_name":"Biotechnology and Biological Sciences Research Council","grant_id":"BB/V003984/1","title":null},{"funder_name":"European Research Council","grant_id":"681987","title":null}],"total_grants":4,"fwci":26.9251,"citation_percentile":0.99720571,"influential_citations":0,"citation_trend":[{"year":2023,"count":7},{"year":2024,"count":25},{"year":2025,"count":42},{"year":2026,"count":18}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://academic.oup.com/bioinformatics/advance-article-pdf/doi/10.1093/bioinformatics/btad308/50267442/btad308.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/advance-article-pdf/doi/10.1093/bioinformatics/btad308/50267442/btad308.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/39/5/btad308/50408270/btad308.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/btad308","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/37162382","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/10199239","host_type":"repository"},{"url":"https://www.repository.cam.ac.uk/handle/1810/349646","host_type":"repository"},{"url":"https://doi.org/10.17863/cam.96590","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10199239/pdf/btad308.pdf","host_type":"repository"},{"url":"https://www.repository.cam.ac.uk/bitstreams/088106bb-69ec-43f7-92f1-836ae678c370/download","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC10199239","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC10199239?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Chromosomal and Genetic Variations","Genomics and Phylogenetic Studies","Genome Rearrangement Algorithms","Humans","Tandem Repeat Sequences","Repetitive Sequences, Nucleic Acid","Base Sequence","Genomics","Centromere","Sequence Analysis, DNA"],"mesh_terms":["Base Sequence","Centromere","Humans","Repetitive Sequences, Nucleic Acid","Sequence Analysis, DNA","Tandem Repeat Sequences","Genomics"],"keywords":["Annotation","Computer science","Hierarchy","Tandem repeat","Tandem","Natural language processing","Artificial intelligence","Computational biology","Biology","Genetics","Genome","Engineering"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-04T13:18:03.342751Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}