{"doi":"10.1101/2020.10.30.361253","title":"The unbiased estimation of the fraction of variance explained by a model","abstract":"Abstract The correlation coefficient squared, r 2 , is often used to validate quantitative models on neural data. Yet it is biased by trial-to-trial variability: as trial-to-trial variability increases, measured correlation to a model’s predictions decreases; therefore, models that perfectly explain neural tuning can appear to perform poorly. Many solutions to this problem have been proposed, but some prior methods overestimate model fit, the utility of even the best performing methods is limited by the lack of confidence intervals and asymptotic analysis, and no consensus has been reached on which is the least biased estimator. We provide a new estimator, , that outperforms all prior estimators in our testing, and we provide confidence intervals and asymptotic guarantees. We apply our estimator to a variety of neural data to validate its utility. We find that neural noise is often so great that confidence intervals of the estimator cover the entire possible range of values ([0, 1]), preventing meaningful evaluation of the quality of a model’s predictions. We demonstrate the use of the signal-to-noise ratio (SNR) as a quality metric for making quantitative comparisons across neural recordings. Analyzing a variety of neural data sets, we find ~ 40% or less of some neural recordings do not pass even a liberal SNR criterion. Author Summary Quantifying the similarity between a model and noisy data is fundamental to the verification of advances in scientific understanding of biological phenomena, and it is particularly relevant to modeling neuronal responses. A ubiquitous metric of similarity is the correlation coefficient. Here we point out how the correlation coefficient depends on a variety of factors that are irrelevant to the similarity between a model and data. While neuroscientists have recognized this problem and proposed corrected methods, no consensus has been reached as to which are effective. Prior methods have wide variation in their precision, and even the most successful methods lack confidence intervals, leaving uncertainty about the reliability of any particular estimate. We address these issues by developing a new estimator along with an associated confidence interval that outperforms all prior methods.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2020,"id":123084,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9476,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":565686,"name":"Wyeth Bair","orcid":"0009-0008-4642-8447","position":1,"is_corresponding":false},{"id":565685,"name":"Dean A. Pospisil","orcid":"0000-0002-5793-2517","position":0,"is_corresponding":true}],"reference_count":19,"raw_metadata":null,"created_at":"2026-07-18T23:14:59.547352Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}