{"doi":"10.1093/bioinformatics/btr424","title":"Identifying disease-associated SNP clusters via contiguous outlier detection","abstract":"<jats:title>Abstract</jats:title><jats:p>Motivation: Although genome-wide association studies (GWAS) have identified many disease-susceptibility single-nucleotide polymorphisms (SNPs), these findings can only explain a small portion of genetic contributions to complex diseases, which is known as the missing heritability. A possible explanation is that genetic variants with small effects have not been detected. The chance is &amp;lt; 8 that a causal SNP will be directly genotyped. The effects of its neighboring SNPs may be too weak to be detected due to the effect decay caused by imperfect linkage disequilibrium. Moreover, it is still challenging to detect a causal SNP with a small effect even if it has been directly genotyped.</jats:p><jats:p>Results: In order to increase the statistical power when detecting disease-associated SNPs with relatively small effects, we propose a method using neighborhood information. Since the disease-associated SNPs account for only a small fraction of the entire SNP set, we formulate this problem as Contiguous Outlier DEtection (CODE), which is a discrete optimization problem. In our formulation, we cast the disease-associated SNPs as outliers and further impose a spatial continuity constraint for outlier detection. We show that this optimization can be solved exactly using graph cuts. We also employ the stability selection strategy to control the false positive results caused by imperfect parameter tuning. We demonstrate its advantage in simulations and real experiments. In particular, the newly identified SNP clusters are replicable in two independent datasets.</jats:p><jats:p>Availability: The software is available at: http://bioinformatics.ust.hk/CODE.zip.</jats:p><jats:p>Contact:  eeyu@ust.hk</jats:p><jats:p>Supplementary information:  Supplementary data are available at Bioinformatics online.</jats:p>","journal":"Bioinformatics","year":2011,"id":630038,"datarank":0.29188652235829704,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"self_citation_contribution":0.29188652235829704,"citation_network_contribution":0.0,"self_endowment_contribution":0.29188652235829704,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":6,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1632027,"name":"Xiaowei Zhou","orcid":null,"position":1,"is_corresponding":false},{"id":79027,"name":"Xiang Wan","orcid":"0000-0002-2324-6527","position":2,"is_corresponding":false},{"id":476827,"name":"Qiang Yang","orcid":"0000-0001-8367-1755","position":3,"is_corresponding":false},{"id":578407,"name":"Hong Xue","orcid":"0000-0002-8133-9828","position":4,"is_corresponding":false},{"id":1295708,"name":"Weichuan Yu","orcid":"0000-0002-5510-6916","position":5,"is_corresponding":false},{"id":261823,"name":"Can Yang","orcid":"0000-0002-4407-3055","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Identifying disease-associated SNP clusters via contiguous outlier detection","abstract":"<jats:title>Abstract</jats:title><jats:p>Motivation: Although genome-wide association studies (GWAS) have identified many disease-susceptibility single-nucleotide polymorphisms (SNPs), these findings can only explain a small portion of genetic contributions to complex diseases, which is known as the missing heritability. A possible explanation is that genetic variants with small effects have not been detected. The chance is &amp;lt; 8 that a causal SNP will be directly genotyped. The effects of its neighboring SNPs may be too weak to be detected due to the effect decay caused by imperfect linkage disequilibrium. Moreover, it is still challenging to detect a causal SNP with a small effect even if it has been directly genotyped.</jats:p><jats:p>Results: In order to increase the statistical power when detecting disease-associated SNPs with relatively small effects, we propose a method using neighborhood information. Since the disease-associated SNPs account for only a small fraction of the entire SNP set, we formulate this problem as Contiguous Outlier DEtection (CODE), which is a discrete optimization problem. In our formulation, we cast the disease-associated SNPs as outliers and further impose a spatial continuity constraint for outlier detection. We show that this optimization can be solved exactly using graph cuts. We also employ the stability selection strategy to control the false positive results caused by imperfect parameter tuning. We demonstrate its advantage in simulations and real experiments. In particular, the newly identified SNP clusters are replicable in two independent datasets.</jats:p><jats:p>Availability: The software is available at: http://bioinformatics.ust.hk/CODE.zip.</jats:p><jats:p>Contact:  eeyu@ust.hk</jats:p><jats:p>Supplementary information:  Supplementary data are available at Bioinformatics online.</jats:p>","is_dataset_classified":null,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"21784794","pmcid":null,"openalex_id":"https://openalex.org/W2123126806","authors":[],"funders":[],"total_grants":0,"fwci":0.5825,"citation_percentile":0.7263113,"influential_citations":0,"citation_trend":[{"year":2012,"count":1},{"year":2013,"count":1},{"year":2014,"count":1},{"year":2015,"count":1},{"year":2018,"count":1},{"year":2022,"count":1}],"oa_status":"bronze","license":null,"oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/27/18/2578/16899652/btr424.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/27/18/2578/16899652/btr424.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/27/18/2578/48866406/bioinformatics_27_18_2578.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/btr424","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/21784794","host_type":"repository"},{"url":"http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.999.2198","host_type":""},{"url":"http://bioinformatics.oxfordjournals.org/cgi/content/short/27/18/2578","host_type":"repository"},{"url":"http://repository.hkust.edu.hk/ir/Record/1783.1-33864","host_type":"repository"}],"fields_of_study":["Genetic Associations and Epidemiology","Bioinformatics and Genomic Networks","Genomics and Rare Diseases","Algorithms","Cluster Analysis","Disease","Genetic Predisposition to Disease","Genome-Wide Association Study","Genotype","Humans","Linkage Disequilibrium","Polymorphism, Single Nucleotide","Software"],"mesh_terms":["Algorithms","Disease","Genotype","Humans","Software","Linkage Disequilibrium","Cluster Analysis","Genetic Predisposition to Disease","Polymorphism, Single Nucleotide","Genome-Wide Association Study"],"keywords":["Linkage disequilibrium","Single-nucleotide polymorphism","Tag SNP","SNP","Genome-wide association study","Outlier","Genetic association","SNP genotyping","Computer science","Biology","Computational biology","Genetics","Data mining","Artificial intelligence","Genotype","Gene"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Partnerships for the goals"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-05T20:19:27.725280Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}