{"doi":"10.1073/pnas.2113075119","title":"AnchorWave: Sensitive alignment of genomes with high sequence diversity, extensive structural polymorphism, and whole-genome duplication","abstract":"<jats:title>Significance</jats:title>\n                  <jats:p>One fundamental analysis needed to interpret genome assemblies is genome alignment. Yet, accurately aligning regulatory and transposon regions outside of genes remains challenging. We introduce Anchored Wavefront alignment (AnchorWave), which implements a genome duplication informed longest path algorithm to identify collinear regions and performs base pair–resolved, end-to-end alignment for collinear blocks using an efficient two-piece affine gap cost strategy. AnchorWave improves the alignment under a number of scenarios: genomes with high similarity, large genomes with high transposable element activity, genomes with many inversions, and alignments between species with deeper evolutionary divergence and different whole-genome duplication histories. Potential use cases include genome comparison for evolutionary analysis of nongenic sequences and population genetics of taxa with large, repeat-rich genomes.</jats:p>","journal":"Proceedings of the National Academy of Sciences","year":2022,"id":593543,"datarank":0.7117398192544876,"base_score":4.74493212836325,"endowment":4.74493212836325,"self_citation_contribution":0.7117398192544876,"citation_network_contribution":0.0,"self_endowment_contribution":0.7117398192544876,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":114,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":2,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":30917,"name":"Santiago Marco‐Sola","orcid":"0000-0001-7951-3914","position":1,"is_corresponding":false},{"id":1025925,"name":"Miquel Moretó","orcid":"0000-0002-9848-8758","position":2,"is_corresponding":false},{"id":1175119,"name":"Lynn Johnson","orcid":"0000-0001-8103-2722","position":3,"is_corresponding":false},{"id":19346,"name":"Edward S. Buckler","orcid":"0000-0002-3100-371X","position":4,"is_corresponding":false},{"id":1325676,"name":"Michelle C. Stitzer","orcid":"0000-0003-4140-3765","position":5,"is_corresponding":false},{"id":1519116,"name":"Baoxing Song","orcid":"0000-0003-1478-9228","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"AnchorWave: Sensitive alignment of genomes with high sequence diversity, extensive structural polymorphism, and whole-genome duplication","abstract":"<jats:title>Significance</jats:title>\n                  <jats:p>One fundamental analysis needed to interpret genome assemblies is genome alignment. Yet, accurately aligning regulatory and transposon regions outside of genes remains challenging. We introduce Anchored Wavefront alignment (AnchorWave), which implements a genome duplication informed longest path algorithm to identify collinear regions and performs base pair–resolved, end-to-end alignment for collinear blocks using an efficient two-piece affine gap cost strategy. AnchorWave improves the alignment under a number of scenarios: genomes with high similarity, large genomes with high transposable element activity, genomes with many inversions, and alignments between species with deeper evolutionary divergence and different whole-genome duplication histories. Potential use cases include genome comparison for evolutionary analysis of nongenic sequences and population genetics of taxa with large, repeat-rich genomes.</jats:p>","is_dataset_classified":null,"base_score":4.74493212836325,"endowment":4.74493212836325,"datacite_reuse_total":2,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"34934012","pmcid":"PMC8740769","openalex_id":"https://openalex.org/W4200042858","authors":[],"funders":[],"total_grants":0,"fwci":13.5214,"citation_percentile":0.99005501,"influential_citations":0,"citation_trend":[{"year":2021,"count":1},{"year":2022,"count":10},{"year":2023,"count":17},{"year":2024,"count":38},{"year":2025,"count":38},{"year":2026,"count":10}],"oa_status":"hybrid","license":"cc-by-nc-nd","oa_locations":[{"url":"https://www.pnas.org/content/pnas/119/1/e2113075119.full.pdf","host_type":"journal"},{"url":"https://www.pnas.org/content/pnas/119/1/e2113075119.full.pdf","host_type":"publisher"},{"url":"https://pnas.org/doi/pdf/10.1073/pnas.2113075119","host_type":"publisher"},{"url":"https://doi.org/10.1073/pnas.2113075119","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/34934012","host_type":"repository"},{"url":"https://ddd.uab.cat/record/311389","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/8740769","host_type":"repository"},{"url":"http://hdl.handle.net/2117/363131","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC8740769","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC8740769?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Chromosomal and Genetic Variations","Genomics and Phylogenetic Studies","Plant Disease Resistance and Genetics","Genome, Plant","Polymorphism, Genetic","Sequence Alignment","Software","Zea mays"],"mesh_terms":["Zea mays","Polymorphism, Genetic","Software","Sequence Alignment","Genome, Plant"],"keywords":["Genome","Indel","Biology","Structural variation","Genetics","Gene duplication","Computational biology","Genome evolution","Bacterial genome size","Transposable element","Comparative genomics","INDEL Mutation","Genomics","Evolutionary biology","Gene","Single-nucleotide polymorphism","Genome comparison","Whole-genome Duplication","Regulatory Element Alignment","Sensitive Genome Alignment","Transposable Element Variation"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Life in Land"}],"linked_datasets":[{"doi":"10.6084/m9.figshare.25903882.v1","title":"Additional file 1 of ACMGA: a reference-free multiple-genome alignment pipeline for plant species","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.25903882","title":"Additional file 1 of ACMGA: a reference-free multiple-genome alignment pipeline for plant species","publisher":"figshare","resource_type":"JournalArticle"}],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"doi"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-27T11:24:46.420018Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}