{"doi":"10.1093/jrsssc/qlae028","title":"Two-phase biomarker studies for disease progression with multiple registries","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>We consider the design and analysis of two-phase studies of the association between an expensive biomarker and disease progression when phase I data are obtained by pooling registries having different outcome-dependent recruitment schemes. We utilize two analysis methods, namely maximum-likelihood and inverse probability weighting (IPW), to handle missing covariates arising from a two-phase design. In the likelihood framework, we derive a class of residual-dependent designs for phase II sub-sampling from an observed data likelihood accounting for the phase I sampling plans used by the different registries. In the IPW approach, we derive and evaluate optimal stratified designs that approximate Neyman allocation. Simulation studies and an application to a motivating example demonstrate the finite sample improvements from the proposed designs over simple random sampling and standard stratified sampling schemes.</jats:p>","journal":"Journal of the Royal Statistical Society Series C: Applied Statistics","year":2024,"id":19445,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":1,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":131775,"name":"Richard J Cook","orcid":"0000-0002-1414-4908","position":1,"is_corresponding":false},{"id":131774,"name":"Fangya Mao","orcid":"0000-0003-1162-4348","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"21071399","pmcid":null,"openalex_id":"https://openalex.org/W4399600420","authors":[],"funders":[{"funder_name":"Natural Sciences and Engineering Research Council of Canada","grant_id":"RGPIN-2017-04207 and DGDND-2017-00072","title":null},{"funder_name":"Natural Sciences and Engineering Research Council of Canada","grant_id":"unidentified","title":"unidentified"}],"total_grants":2,"fwci":1.2487,"citation_percentile":0.79463583,"influential_citations":0,"citation_trend":[{"year":2025,"count":1},{"year":2026,"count":1}],"oa_status":"bronze","license":"CC BY NC","oa_locations":[{"url":"https://academic.oup.com/jrsssc/advance-article-pdf/doi/10.1093/jrsssc/qlae028/58225896/qlae028.pdf","host_type":"journal"},{"url":"https://academic.oup.com/jrsssc/advance-article-pdf/doi/10.1093/jrsssc/qlae028/58225896/qlae028.pdf","host_type":"HYBRID"},{"url":"https://academic.oup.com/jrsssc/advance-article-pdf/doi/10.1093/jrsssc/qlae028/58225896/qlae028.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/jrsssc/article-pdf/73/5/1111/60661296/qlae028.pdf","host_type":"publisher"},{"url":"http://dx.doi.org/10.1093/jrsssc/qlae028","host_type":"journal"},{"url":"https://doi.org/10.1093/jrsssc/qlae028","host_type":""},{"url":"https://zbmath.org/8015840","host_type":""}],"fields_of_study":["Statistical Methods in Clinical Trials","Statistical Methods and Inference","Optimal Experimental Design Methods","Medicine","Mathematics","03 medical and health sciences","0302 clinical medicine","0101 mathematics","01 natural sciences"],"mesh_terms":[],"keywords":["Sampling design","Sampling (signal processing)","Inverse probability","Inverse probability weighting","Covariate","Pooling","Missing data","Stratified sampling","Statistics","Weighting","Simple random sample","Computer science","Mathematics","Bayesian probability","Artificial intelligence","Medicine","Posterior probability","Population","missing covariate","two-phase sampling","multistate analysis","psoriatic disease","Applications of statistics","truncated data"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-06-04T04:06:46.596159Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}