{"doi":"10.1093/bioinformatics/btm496","title":"Genome-wide selection of tag SNPs using multiple-marker correlation","abstract":"<jats:title>Abstract</jats:title><jats:p>Motivations: The tag SNP approach is a valuable tool in whole genome association studies, and a variety of algorithms have been proposed to identify the optimal tag SNP set. Currently, most tag SNP selection is based on two-marker (pairwise) linkage disequilibrium (LD). Recent literature has shown that multiple-marker LD also contains useful information that can further increase the genetic coverage of the tag SNP set. Thus, tag SNP selection methods that incorporate multiple-marker LD are expected to have advantages in terms of genetic coverage and statistical power.</jats:p><jats:p>Results: We propose a novel algorithm to select tag SNPs in an iterative procedure. In each iteration loop, the SNP that captures the most neighboring SNPs (through pair-wise and multiple-marker LD) is selected as a tag SNP. We optimize the algorithm and computer program to make our approach feasible on today's typical workstations. Benchmarked using HapMap release 21, our algorithm outperforms standard pair-wise LD approach in several aspects. (i) It improves genetic coverage (e.g. by 7.2% for 200 K tag SNPs in HapMap CEU) compared to its conventional pair-wise counterpart, when conditioning on a fixed tag SNP number. (ii) It saves genotyping costs substantially when conditioning on fixed genetic coverage (e.g. 34.1% saving in HapMap CEU at 90% coverage). (iii) Tag SNPs identified using multiple-marker LD have good portability across closely related ethnic groups and (iv) show higher statistical power in association tests than those selected using conventional methods.</jats:p><jats:p>Availability: A computer software suite, multiTag, has been developed based on this novel algorithm. The program is freely available by written request to the author at ke_hao@merck.com</jats:p><jats:p>Contact:  ke_hao@163.com</jats:p><jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.</jats:p>","journal":"Bioinformatics","year":2007,"id":14984,"datarank":1.3376571066404717,"base_score":3.1780538303479458,"endowment":3.1780538303479458,"self_citation_contribution":0.47670807455219194,"citation_network_contribution":0.8609490320882797,"self_endowment_contribution":0.47670807455219194,"citer_contribution":0.8609490320882797,"corpus_percentile":null,"corpus_rank":null,"citation_count":23,"citer_count":23,"citers_with_citation_signal":19,"citers_with_endowment":19,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":116034,"name":"K. Hao","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"base_score":3.1780538303479458,"endowment":3.1780538303479458,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"18006555","pmcid":null,"openalex_id":"https://openalex.org/W2152596080","authors":[],"funders":[],"total_grants":0,"fwci":1.0909,"citation_percentile":0.80520865,"influential_citations":1,"citation_trend":[{"year":2012,"count":3},{"year":2013,"count":2},{"year":2014,"count":2},{"year":2015,"count":1},{"year":2016,"count":3},{"year":2021,"count":1},{"year":2022,"count":2}],"oa_status":"closed","license":null,"oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/23/23/3178/16860850/btm496.pdf","host_type":"BRONZE"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/23/23/3178/49824350/bioinformatics_23_23_3178.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/btm496","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/18006555","host_type":"repository"}],"fields_of_study":["Genetic Associations and Epidemiology","Genetic and phenotypic traits in livestock","Forensic and Genetic Research","Biology","Computer Science","Medicine","Algorithms","Base Sequence","Chromosome Mapping","Expressed Sequence Tags","Genetic Markers","Linkage Disequilibrium","Molecular Sequence Data","Polymorphism, Single Nucleotide","Sequence Analysis, DNA","Statistics as Topic"],"mesh_terms":["Algorithms","Base Sequence","Chromosome Mapping","Genetic Markers","Molecular Sequence Data","Statistics as Topic","Linkage Disequilibrium","Sequence Analysis, DNA","Expressed Sequence Tags","Polymorphism, Single Nucleotide"],"keywords":["International HapMap Project","Tag SNP","Single-nucleotide polymorphism","SNP","SNP genotyping","Selection (genetic algorithm)","Linkage disequilibrium","Computer science","Genetic association","Computational biology","Genetics","Data mining","Biology","Machine learning","Gene","Genotype"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Partnerships for the goals"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-06-01T15:53:28.422132Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}