{"doi":"10.1101/2021.04.05.438429","title":"Warped Bayesian Linear Regression for Normative Modelling of Big Data","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>Normative modelling is becoming more popular in neuroimaging due to its ability to make predictions of deviation from a normal trajectory at the level of individual participants. It allows the user to model the distribution of several neuroimaging modalities, giving an estimation for the mean and centiles of variation. With the increase in the availability of big data in neuroimaging, there is a need to scale normative modelling to big data sets. However, the scaling of normative models has come with several challenges.</jats:p>\n                <jats:p>So far, most normative modelling approaches used Gaussian process regression, and although suitable for smaller datasets (up to a few thousand participants) it does not scale well to the large cohorts currently available and being acquired. Furthermore, most neuroimaging modelling methods that are available assume the predictive distribution to be Gaussian in shape. However, deviations from Gaussianity can be frequently found, which may lead to incorrect inferences, particularly in the outer centiles of the distribution. In normative modelling, we use the centiles to give an estimation of the deviation of a particular participant from the ‘normal’ trend. Therefore, especially in normative modelling, the correct estimation of the outer centiles is of utmost importance, which is also where data are sparsest.</jats:p>\n                <jats:p>Here, we present a novel framework based on Bayesian Linear Regression with likelihood warping that allows us to address these problems, that is, to scale normative modelling elegantly to big data cohorts and to correctly model non-Gaussian predictive distributions. In addition, this method provides also likelihood-based statistics, which are useful for model selection.</jats:p>\n                <jats:p>To evaluate this framework, we use a range of neuroimaging-derived measures from the UK Biobank study, including image-derived phenotypes (IDPs) and whole-brain voxel-wise measures derived from diffusion tensor imaging. We show good computational scaling and improved accuracy of the warped BLR for certain IDPs and voxels if there was a deviation from normality of these parameters in their residuals.</jats:p>\n                <jats:p>The present results indicate the advantage of a warped BLR in terms of; computational scalability and the flexibility to incorporate non-linearity and non-Gaussianity of the data, giving a wider range of neuroimaging datasets that can be correctly modelled.</jats:p>","journal":null,"year":null,"id":601587,"datarank":0.4636563680037475,"base_score":3.091042453358316,"endowment":3.091042453358316,"self_citation_contribution":0.4636563680037475,"citation_network_contribution":0.0,"self_endowment_contribution":0.4636563680037475,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":21,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":275658,"name":"Richard Dinga","orcid":"0000-0003-3040-1297","position":1,"is_corresponding":false},{"id":341865,"name":"Christian F. Beckmann","orcid":"0000-0002-3373-3193","position":2,"is_corresponding":false},{"id":107839,"name":"André F. Marquand","orcid":"0000-0001-5903-203X","position":3,"is_corresponding":false},{"id":806379,"name":"Charlotte Fraza","orcid":"0000-0002-7088-9250","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Warped Bayesian Linear Regression for Normative Modelling of Big Data","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>Normative modelling is becoming more popular in neuroimaging due to its ability to make predictions of deviation from a normal trajectory at the level of individual participants. It allows the user to model the distribution of several neuroimaging modalities, giving an estimation for the mean and centiles of variation. With the increase in the availability of big data in neuroimaging, there is a need to scale normative modelling to big data sets. However, the scaling of normative models has come with several challenges.</jats:p>\n                <jats:p>So far, most normative modelling approaches used Gaussian process regression, and although suitable for smaller datasets (up to a few thousand participants) it does not scale well to the large cohorts currently available and being acquired. Furthermore, most neuroimaging modelling methods that are available assume the predictive distribution to be Gaussian in shape. However, deviations from Gaussianity can be frequently found, which may lead to incorrect inferences, particularly in the outer centiles of the distribution. In normative modelling, we use the centiles to give an estimation of the deviation of a particular participant from the ‘normal’ trend. Therefore, especially in normative modelling, the correct estimation of the outer centiles is of utmost importance, which is also where data are sparsest.</jats:p>\n                <jats:p>Here, we present a novel framework based on Bayesian Linear Regression with likelihood warping that allows us to address these problems, that is, to scale normative modelling elegantly to big data cohorts and to correctly model non-Gaussian predictive distributions. In addition, this method provides also likelihood-based statistics, which are useful for model selection.</jats:p>\n                <jats:p>To evaluate this framework, we use a range of neuroimaging-derived measures from the UK Biobank study, including image-derived phenotypes (IDPs) and whole-brain voxel-wise measures derived from diffusion tensor imaging. We show good computational scaling and improved accuracy of the warped BLR for certain IDPs and voxels if there was a deviation from normality of these parameters in their residuals.</jats:p>\n                <jats:p>The present results indicate the advantage of a warped BLR in terms of; computational scalability and the flexibility to incorporate non-linearity and non-Gaussianity of the data, giving a wider range of neuroimaging datasets that can be correctly modelled.</jats:p>","is_dataset_classified":null,"base_score":3.091042453358316,"endowment":3.091042453358316,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"34798518","pmcid":"PMC7613680","openalex_id":"https://openalex.org/W3149809682","authors":[],"funders":[{"funder_name":"European Research Council","grant_id":"101001118","title":"Integrating biology and behaviour for precision stratification of mental disorders"},{"funder_name":"Wellcome Trust","grant_id":"215698","title":"BRAINCHART: Normative brain charting for predicting and stratifying psychosis"},{"funder_name":"ZonMw","grant_id":"91716415","title":null},{"funder_name":"Wellcome Trust","grant_id":"215698/Z/19/Z","title":null},{"funder_name":"Medical Research Council","grant_id":"MC_PC_17228","title":null},{"funder_name":"Medical Research Council","grant_id":"MC_QA137853","title":null},{"funder_name":"Wellcome Trust","grant_id":"98369","title":"Integrated multimodal brain imaging for neuroscience research and clinical practice."},{"funder_name":"Wellcome Trust","grant_id":"098369","title":"Integrated multimodal brain imaging for neuroscience research and clinical practice."},{"funder_name":"European Research Council","grant_id":"","title":null},{"funder_name":"Wellcome Trust","grant_id":"","title":null},{"funder_name":"Dutch Research Council (NWO)","grant_id":"","title":null}],"total_grants":11,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2021,"count":5},{"year":2022,"count":5},{"year":2023,"count":2},{"year":2024,"count":4},{"year":2025,"count":3},{"year":2026,"count":2}],"oa_status":"green","license":"cc-by","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2021/04/06/2021.04.05.438429.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2021/04/06/2021.04.05.438429.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/2021.04.05.438429","host_type":"publisher"},{"url":"https://doi.org/10.1101/2021.04.05.438429","host_type":"repository"},{"url":"https://pubmed.ncbi.nlm.nih.gov/34798518","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/7613680","host_type":"repository"},{"url":"https://doi.org/10.1016/j.neuroimage.2021.118715","host_type":"Unpaywall"},{"url":"https://europepmc.org/articles/PMC7613680","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC7613680?pdf=render","host_type":"Europe_PMC"},{"url":"http://dx.doi.org/10.1101/2021.04.05.438429","host_type":""},{"url":"https://doaj.org/article/a7db1f76a25a4ddca2aa6236565dc09e","host_type":""},{"url":"https://hdl.handle.net/https://repository.ubn.ru.nl/handle/2066/244123","host_type":""},{"url":"https://dx.doi.org/10.1016/j.neuroimage.2021.118715","host_type":""},{"url":"https://dx.doi.org/10.1101/2021.04.05.438429","host_type":""},{"url":"https://hdl.handle.net/2066/244123","host_type":""},{"url":"https://repository.ubn.ru.nl//bitstream/handle/2066/244123/244123.pdf","host_type":""},{"url":"http://dx.doi.org/10.1016/j.neuroimage.2021.118715","host_type":""},{"url":"https://doi.org/https://doi.org/10.1016/j.neuroimage.2021.118715","host_type":""}],"fields_of_study":["Advanced Neuroimaging Techniques and Applications","Functional Brain Connectivity Studies","Machine Learning in Healthcare","0301 basic medicine","03 medical and health sciences","0302 clinical medicine"],"mesh_terms":["Big Data","Bayes Theorem","United Kingdom","Humans","Normal Distribution","Diffusion Tensor Imaging","Neuroimaging"],"keywords":["Normative","Bayesian probability","Bayesian linear regression","Computer science","Neuroimaging","Big data","Econometrics","Statistics","Artificial intelligence","Psychology","Bayesian inference","Data mining","Mathematics","Machine Learning","Uk Biobank","Normative Modelling","Radboudumc 7: Neurodevelopmental disorders DCMN: Donders Center for Medical Neuroscience","Normal Distribution","610","220 Statistical Imaging Neuroscience","Neurosciences. Biological psychiatry. Neuropsychiatry","Bayes Theorem","Article","United Kingdom","510","Diffusion Tensor Imaging","Medical Neuroscience - Radboud University Medical Center","Humans","RC321-571"],"sdg_mappings":[{"sdg_number":16,"sdg_label":"16. Peace & justice"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"doi"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-29T16:48:29.887371Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}