{"doi":"10.1093/bioinformatics/bti124","title":"Validation of alternative methods of data normalization in gene co-expression studies","abstract":"<jats:title>Abstract</jats:title><jats:p>Motivation: Clusters of genes encoding proteins with related functions, or in the same regulatory network, often exhibit expression patterns that are correlated over a large number of conditions. Protein associations and gene regulatory networks can be modelled from expression data. We address the question of which of several normalization methods is optimal prior to computing the correlation of the expression profiles between every pair of genes.</jats:p><jats:p>Results: We use gene expression data from five experiments with a total of 78 hybridizations and 23 diverse conditions. Nine methods of data normalization are explored based on all possible combinations of normalization techniques according to between and within gene and experiment variation. We compare the resulting empirical distribution of gene × gene correlations with the expectations and apply cross-validation to test the performance of each method in predicting accurate functional annotation. We conclude that normalization methods based on mixed-model equations are optimal.</jats:p><jats:p>Contact:  tony.reverter-gomez@csiro.au</jats:p>","journal":"Bioinformatics","year":2005,"id":589668,"datarank":1.696069655641009,"base_score":4.2626798770413155,"endowment":4.2626798770413155,"self_citation_contribution":0.6394019815561974,"citation_network_contribution":1.0566676740848118,"self_endowment_contribution":0.6394019815561974,"citer_contribution":1.0566676740848118,"corpus_percentile":null,"corpus_rank":null,"citation_count":70,"citer_count":31,"citers_with_citation_signal":21,"citers_with_endowment":21,"datacite_reuse_total":4,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1508653,"name":"Wes Barris","orcid":null,"position":1,"is_corresponding":false},{"id":1508654,"name":"Sean McWilliam","orcid":null,"position":2,"is_corresponding":false},{"id":1508655,"name":"Keren A. Byrne","orcid":null,"position":3,"is_corresponding":false},{"id":1508656,"name":"Yong H. Wang","orcid":null,"position":4,"is_corresponding":false},{"id":1508657,"name":"Siok H. Tan","orcid":null,"position":5,"is_corresponding":false},{"id":1508658,"name":"Nick Hudson","orcid":null,"position":6,"is_corresponding":false},{"id":1508659,"name":"Brian P. Dalrymple","orcid":null,"position":7,"is_corresponding":false},{"id":1508652,"name":"Antonio Reverter","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Validation of alternative methods of data normalization in gene co-expression studies","abstract":"<jats:title>Abstract</jats:title><jats:p>Motivation: Clusters of genes encoding proteins with related functions, or in the same regulatory network, often exhibit expression patterns that are correlated over a large number of conditions. Protein associations and gene regulatory networks can be modelled from expression data. We address the question of which of several normalization methods is optimal prior to computing the correlation of the expression profiles between every pair of genes.</jats:p><jats:p>Results: We use gene expression data from five experiments with a total of 78 hybridizations and 23 diverse conditions. Nine methods of data normalization are explored based on all possible combinations of normalization techniques according to between and within gene and experiment variation. We compare the resulting empirical distribution of gene × gene correlations with the expectations and apply cross-validation to test the performance of each method in predicting accurate functional annotation. We conclude that normalization methods based on mixed-model equations are optimal.</jats:p><jats:p>Contact:  tony.reverter-gomez@csiro.au</jats:p>","is_dataset_classified":null,"base_score":4.2626798770413155,"endowment":4.2626798770413155,"datacite_reuse_total":4,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"15564293","pmcid":null,"openalex_id":"https://openalex.org/W2113286112","authors":[],"funders":[],"total_grants":0,"fwci":1.9159,"citation_percentile":0.85077019,"influential_citations":0,"citation_trend":[{"year":2012,"count":3},{"year":2013,"count":2},{"year":2014,"count":2},{"year":2015,"count":2},{"year":2016,"count":2},{"year":2017,"count":1},{"year":2018,"count":4},{"year":2019,"count":5},{"year":2020,"count":6},{"year":2021,"count":2},{"year":2022,"count":3},{"year":2023,"count":4},{"year":2024,"count":2},{"year":2025,"count":4}],"oa_status":"bronze","license":null,"oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/21/7/1112/48966778/bioinformatics_21_7_1112.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/21/7/1112/48966778/bioinformatics_21_7_1112.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/bti124","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/15564293","host_type":"repository"}],"fields_of_study":["Gene expression and cancer classification","Bioinformatics and Genomic Networks","Gene Regulatory Network Analysis","Algorithms","Benchmarking","Computer Simulation","Data Interpretation, Statistical","Gene Expression Profiling","Gene Expression Regulation","Models, Genetic","Models, Statistical","Numerical Analysis, Computer-Assisted","Oligonucleotide Array Sequence Analysis","Software"],"mesh_terms":["Algorithms","Computer Simulation","Data Interpretation, Statistical","Gene Expression Regulation","Models, Genetic","Numerical Analysis, Computer-Assisted","Software","Models, Statistical","Benchmarking","Oligonucleotide Array Sequence Analysis","Gene Expression Profiling"],"keywords":["Normalization (sociology)","Database normalization","Computational biology","Gene","Gene expression","Correlation","Gene expression profiling","Biology","Expression (computer science)","Computer science","Data mining","Genetics","Artificial intelligence","Mathematics","Pattern recognition (psychology)"],"sdg_mappings":[],"linked_datasets":[{"doi":"10.6084/m9.figshare.13352208.v1","title":"Additional file 12 of A systems biology framework integrating GWAS and RNA-seq to shed light on the molecular basis of sperm quality in swine","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.13352208","title":"Additional file 12 of A systems biology framework integrating GWAS and RNA-seq to shed light on the molecular basis of sperm quality in swine","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.13352217.v1","title":"Additional file 2 of A systems biology framework integrating GWAS and RNA-seq to shed light on the molecular basis of sperm quality in swine","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.13352217","title":"Additional file 2 of A systems biology framework integrating GWAS and RNA-seq to shed light on the molecular basis of sperm quality in swine","publisher":"figshare","resource_type":"JournalArticle"}],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-24T02:44:04.856908Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}