{"doi":"10.1186/1471-2105-9-306","title":"Grammar-based distance in progressive multiple sequence alignment","abstract":null,"journal":"BMC Bioinformatics","year":2008,"id":649991,"datarank":0.5641800173540344,"base_score":3.7612001156935624,"endowment":3.7612001156935624,"self_citation_contribution":0.5641800173540344,"citation_network_contribution":0.0,"self_endowment_contribution":0.5641800173540344,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":42,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1694717,"name":"Hasan H Otu","orcid":null,"position":1,"is_corresponding":false},{"id":751051,"name":"Khalid Sayood","orcid":"0000-0001-8010-5272","position":2,"is_corresponding":false},{"id":1694716,"name":"David J Russell","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Grammar-based distance in progressive multiple sequence alignment","abstract":"BACKGROUND: We propose a multiple sequence alignment (MSA) algorithm and compare the alignment-quality and execution-time of the proposed algorithm with that of existing algorithms. The proposed progressive alignment algorithm uses a grammar-based distance metric to determine the order in which biological sequences are to be pairwise aligned. The progressive alignment occurs via pairwise aligning new sequences with an ensemble of the sequences previously aligned. RESULTS: The performance of the proposed algorithm is validated via comparison to popular progressive multiple alignment approaches, ClustalW and T-Coffee, and to the more recently developed algorithms MAFFT, MUSCLE, Kalign, and PSAlign using the BAliBASE 3.0 database of amino acid alignment files and a set of longer sequences generated by Rose software. The proposed algorithm has successfully built multiple alignments comparable to other programs with significant improvements in running time. The results are especially striking for large datasets. CONCLUSION: We introduce a computationally efficient progressive alignment algorithm using a grammar based sequence distance particularly useful in aligning large datasets.","is_dataset_classified":null,"base_score":3.7612001156935624,"endowment":3.7612001156935624,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"18616828","pmcid":"PMC2478692","openalex_id":"https://openalex.org/W2149492307","authors":[],"funders":[{"funder_name":"NIAID NIH HHS","grant_id":"K25 AI068151","title":null},{"funder_name":"National Institutes of Health","grant_id":"1K25AI068151-01","title":"Identification of Biological Materials of Unknown Origin"}],"total_grants":2,"fwci":1.8213,"citation_percentile":0.84952095,"influential_citations":0,"citation_trend":[{"year":2012,"count":5},{"year":2013,"count":5},{"year":2014,"count":1},{"year":2015,"count":2},{"year":2016,"count":2},{"year":2017,"count":3},{"year":2018,"count":4},{"year":2019,"count":1},{"year":2020,"count":1},{"year":2021,"count":1},{"year":2022,"count":2},{"year":2024,"count":1}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://bmcbioinformatics.biomedcentral.com/counter/pdf/10.1186/1471-2105-9-306","host_type":"journal"},{"url":"https://bmcbioinformatics.biomedcentral.com/counter/pdf/10.1186/1471-2105-9-306","host_type":"publisher"},{"url":"http://link.springer.com/content/pdf/10.1186/1471-2105-9-306.pdf","host_type":"publisher"},{"url":"http://link.springer.com/article/10.1186/1471-2105-9-306/fulltext.html","host_type":"publisher"},{"url":"https://doi.org/10.1186/1471-2105-9-306","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/18616828","host_type":"repository"},{"url":"http://nrs.harvard.edu/urn-3:HUL.InstRepos:4774190","host_type":"repository"},{"url":"https://digitalcommons.unl.edu/electricalengineeringfacpub/100","host_type":"repository"},{"url":"https://doaj.org/article/f85243be8a98482b80d06d2b9ea1d195","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/2478692","host_type":"repository"},{"url":"http://www.biomedcentral.com/1471-2105/9/306/abstract","host_type":"BioMedCentral"},{"url":"http://www.biomedcentral.com/1471-2105/9/306","host_type":"BioMedCentral"},{"url":"http://www.biomedcentral.com/content/pdf/1471-2105-9-306.pdf","host_type":"BioMedCentral"},{"url":"https://europepmc.org/articles/PMC2478692","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC2478692?pdf=render","host_type":"Europe_PMC"},{"url":"https://bmcbioinformatics.biomedcentral.com/track/pdf/10.1186/1471-2105-9-306","host_type":""},{"url":"http://dx.doi.org/10.1186/1471-2105-9-306","host_type":""},{"url":"https://dx.doi.org/10.1186/1471-2105-9-306","host_type":""},{"url":"https://doi.org/https://doi.org/10.1186/1471-2105-9-306","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Protein Structure and Dynamics","Microbial Community Ecology and Physiology","0301 basic medicine","0206 medical engineering","02 engineering and technology","03 medical and health sciences"],"mesh_terms":["Animals","Computing Methodologies","Humans","Information Theory","Natural Language Processing","Terminology as Topic","Semantics","Sequence Alignment","Sequence Analysis, DNA","Sequence Analysis, Protein","Genomics","Databases, Genetic"],"keywords":["Multiple sequence alignment","Computer science","Pairwise comparison","Sequence alignment","Alignment-free sequence analysis","Smith–Waterman algorithm","Set (abstract data type)","Structural alignment","Metric (unit)","Sequence (biology)","Algorithm","Software","Data mining","Cluster analysis","Artificial intelligence","Pattern recognition (psychology)","Biology","Peptide sequence","QH301-705.5","Computer applications to medicine. Medical informatics","R858-859.7","Information Theory","Biochemistry","Computing Methodologies","Sequence Analysis, Protein","Terminology as Topic","Databases, Genetic","Animals","Humans","Biology (General)","Molecular Biology","Natural Language Processing","Genomics","Sequence Analysis, DNA","Electrical and Computer Engineering","004","Computer Science Applications","Semantics","Research Article"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-10T04:39:18.538783Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}