{"doi":"10.1002/gepi.20472","title":"Genome‐wide association studies: quality control and population‐based measures","abstract":"<jats:title>Abstract</jats:title><jats:p>Genome‐wide association studies, using hundreds of thousands of single‐nucleotide polymorphism (SNP) markers, have become a standard approach for identifying disease susceptibility genes. The change in the technology poses substantial computational and statistical challenges that have been addressed in the quality control, imputation, and population‐based measure groups of the Genetic Analysis Workshop 16. The computational challenges pertain to efficient memory management and computational speed of the statistical procedures, and we discuss an approach for efficient SNP storage. Accuracy and computational speed is relevant for genotype calling, and the results from a comparison of three calling algorithms are discussed. The first statistical challenge is related to statistical quality control, and we discuss two novel quality control procedures. These low‐level analyses have an effect on subsequent preparatory steps for high‐level analyses, e.g., the quality of genotype imputation approaches. After the conduct of a genome‐wide association study with successful replication and/or validation, measures of diagnostic accuracy, including the area under the curve, are investigated. The area under the curve can be constructed from summary data in some situations. Finally, we discuss how the population‐attributable risk of a genetic variant that is only measured in a reference data set can be determined. <jats:italic>Genet. Epidemiol</jats:italic>. 33 (Suppl. 1):S45–S50, 2009. © 2009 Wiley‐Liss, Inc.</jats:p>","journal":"Genetic Epidemiology","year":2009,"id":647264,"datarank":0.5641800173540344,"base_score":3.7612001156935624,"endowment":3.7612001156935624,"self_citation_contribution":0.5641800173540344,"citation_network_contribution":0.0,"self_endowment_contribution":0.5641800173540344,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":42,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":597401,"name":"Andreas Ziegler","orcid":"0000-0002-8386-5397","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Genome‐wide association studies: quality control and population‐based measures","abstract":"<jats:title>Abstract</jats:title><jats:p>Genome‐wide association studies, using hundreds of thousands of single‐nucleotide polymorphism (SNP) markers, have become a standard approach for identifying disease susceptibility genes. The change in the technology poses substantial computational and statistical challenges that have been addressed in the quality control, imputation, and population‐based measure groups of the Genetic Analysis Workshop 16. The computational challenges pertain to efficient memory management and computational speed of the statistical procedures, and we discuss an approach for efficient SNP storage. Accuracy and computational speed is relevant for genotype calling, and the results from a comparison of three calling algorithms are discussed. The first statistical challenge is related to statistical quality control, and we discuss two novel quality control procedures. These low‐level analyses have an effect on subsequent preparatory steps for high‐level analyses, e.g., the quality of genotype imputation approaches. After the conduct of a genome‐wide association study with successful replication and/or validation, measures of diagnostic accuracy, including the area under the curve, are investigated. The area under the curve can be constructed from summary data in some situations. Finally, we discuss how the population‐attributable risk of a genetic variant that is only measured in a reference data set can be determined. <jats:italic>Genet. Epidemiol</jats:italic>. 33 (Suppl. 1):S45–S50, 2009. © 2009 Wiley‐Liss, Inc.</jats:p>","is_dataset_classified":null,"base_score":3.7612001156935624,"endowment":3.7612001156935624,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19924716","pmcid":"PMC2996103","openalex_id":"https://openalex.org/W2025016633","authors":[],"funders":[{"funder_name":"NIGMS NIH HHS","grant_id":"R01 GM031575","title":null}],"total_grants":1,"fwci":2.5031,"citation_percentile":0.89641239,"influential_citations":0,"citation_trend":[{"year":2012,"count":5},{"year":2013,"count":2},{"year":2014,"count":6},{"year":2015,"count":2},{"year":2016,"count":4},{"year":2017,"count":2},{"year":2018,"count":1},{"year":2019,"count":6},{"year":2020,"count":2},{"year":2022,"count":1},{"year":2023,"count":1},{"year":2026,"count":1}],"oa_status":"green","license":"http://onlinelibrary.wiley.com/termsAndConditions#vor","oa_locations":[{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/2996103","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/2996103","host_type":"repository"},{"url":"https://api.wiley.com/onlinelibrary/tdm/v1/articles/10.1002%2Fgepi.20472","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/pdf/10.1002/gepi.20472","host_type":"publisher"},{"url":"https://doi.org/10.1002/gepi.20472","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/19924716","host_type":"repository"}],"fields_of_study":["Genetic Associations and Epidemiology","Genetic and phenotypic traits in livestock","Algorithms","Alleles","Female","Genome-Wide Association Study","Genotype","Humans","Male","Molecular Epidemiology","Polymorphism, Single Nucleotide","Public Health","Quality Control","ROC Curve"],"mesh_terms":["Algorithms","Alleles","Female","Genotype","Humans","Male","Public Health","Quality Control","ROC Curve","Molecular Epidemiology","Polymorphism, Single Nucleotide","Genome-Wide Association Study"],"keywords":["Imputation (statistics)","Genetic association","SNP","Genome-wide association study","Population","Population stratification","Single-nucleotide polymorphism","Computer science","Quality Score","Data mining","Computational biology","Genetics","Biology","Genotype","Machine learning","Missing data","Gene","Medicine"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-09T17:40:27.358862Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}