{"doi":"10.1002/gepi.20028","title":"Tag SNP selection for association studies","abstract":"<jats:title>Abstract</jats:title><jats:p>This report describes current methods for selection of informative single nucleotide polymorphisms (SNPs) using data from a dense network of SNPs that have been genotyped in a relatively small panel of subjects. We discuss the following issues: (1) Optimal selection of SNPs based upon maximizing either the predictability of unmeasured SNPs or the predictability of SNP haplotypes as selection criteria. (2) The dependence of the performance of tag SNP selection methods upon the density of SNP markers genotyped for the purpose of haplotype discovery and tag SNP selection. (3) The likely power of case‐control studies to detect the influence upon disease risk of common disease‐causing variants in candidate genes in a haplotype‐based analysis. We propose a quasi‐empirical approach towards evaluating the power of large studies with this calculation based upon the SNP genotype and haplotype frequencies estimated in a haplotype discovery panel. In this calculation, each common SNP in turn is treated as a potential unmeasured causal variant and subjected to a correlation analysis using the remaining SNPs. We use a small portion of the HapMap ENCODE data (488 common SNPs genotyped over approximately a 500 kb region of chromosome 2) as an illustrative example of this approach towards power evaluation. © 2004 Wiley‐Liss, Inc.</jats:p>","journal":"Genetic Epidemiology","year":2004,"id":593312,"datarank":9.963389480052172,"base_score":5.293304824724492,"endowment":5.293304824724492,"self_citation_contribution":0.793995723708674,"citation_network_contribution":9.169393756343498,"self_endowment_contribution":0.793995723708674,"citer_contribution":9.169393756343498,"corpus_percentile":null,"corpus_rank":null,"citation_count":198,"citer_count":185,"citers_with_citation_signal":154,"citers_with_endowment":154,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":729,"name":"Daniel O. Stram","orcid":"0000-0001-7099-9022","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Tag SNP selection for association studies","abstract":"<jats:title>Abstract</jats:title><jats:p>This report describes current methods for selection of informative single nucleotide polymorphisms (SNPs) using data from a dense network of SNPs that have been genotyped in a relatively small panel of subjects. We discuss the following issues: (1) Optimal selection of SNPs based upon maximizing either the predictability of unmeasured SNPs or the predictability of SNP haplotypes as selection criteria. (2) The dependence of the performance of tag SNP selection methods upon the density of SNP markers genotyped for the purpose of haplotype discovery and tag SNP selection. (3) The likely power of case‐control studies to detect the influence upon disease risk of common disease‐causing variants in candidate genes in a haplotype‐based analysis. We propose a quasi‐empirical approach towards evaluating the power of large studies with this calculation based upon the SNP genotype and haplotype frequencies estimated in a haplotype discovery panel. In this calculation, each common SNP in turn is treated as a potential unmeasured causal variant and subjected to a correlation analysis using the remaining SNPs. We use a small portion of the HapMap ENCODE data (488 common SNPs genotyped over approximately a 500 kb region of chromosome 2) as an illustrative example of this approach towards power evaluation. © 2004 Wiley‐Liss, Inc.</jats:p>","is_dataset_classified":null,"base_score":5.293304824724492,"endowment":5.293304824724492,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"15372618","pmcid":null,"openalex_id":"https://openalex.org/W2146809629","authors":[],"funders":[{"funder_name":"NIGMS NIH HHS","grant_id":"GM58897","title":null},{"funder_name":"NCI NIH HHS","grant_id":"CA63464","title":null},{"funder_name":"NHGRI NIH HHS","grant_id":"HG002790","title":null}],"total_grants":3,"fwci":10.4248,"citation_percentile":0.9876686,"influential_citations":0,"citation_trend":[{"year":2012,"count":6},{"year":2013,"count":8},{"year":2014,"count":8},{"year":2015,"count":6},{"year":2016,"count":8},{"year":2017,"count":5},{"year":2018,"count":7},{"year":2019,"count":5},{"year":2020,"count":2},{"year":2021,"count":9},{"year":2022,"count":8},{"year":2023,"count":3},{"year":2024,"count":1},{"year":2025,"count":3},{"year":2026,"count":1}],"oa_status":"closed","license":"http://onlinelibrary.wiley.com/termsAndConditions#vor","oa_locations":[{"url":"https://api.wiley.com/onlinelibrary/tdm/v1/articles/10.1002%2Fgepi.20028","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/pdf/10.1002/gepi.20028","host_type":"publisher"},{"url":"https://doi.org/10.1002/gepi.20028","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/15372618","host_type":"repository"}],"fields_of_study":["Genetic Associations and Epidemiology","Genomics and Rare Diseases","RNA modifications and cancer","Chromosome Mapping","Genetic Markers","Genetic Predisposition to Disease","Genome, Human","Genotype","Haplotypes","Humans","Linkage Disequilibrium","Models, Genetic","Polymorphism, Single Nucleotide"],"mesh_terms":["Chromosome Mapping","Genetic Markers","Genotype","Haplotypes","Humans","Models, Genetic","Linkage Disequilibrium","Genome, Human","Genetic Predisposition to Disease","Polymorphism, Single Nucleotide"],"keywords":["Single-nucleotide polymorphism","Tag SNP","International HapMap Project","SNP","Haplotype","Genetics","Selection (genetic algorithm)","Biology","Haplotype estimation","Genetic association","SNP genotyping","Genotype","Computational biology","Gene","Computer science","Artificial intelligence"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-26T20:26:59.512544Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}