{"doi":"10.1093/clinchem/hvab272","title":"How Can We Ensure Reproducibility and Clinical Translation of Machine Learning Applications in Laboratory Medicine?","abstract":"Recent studies demonstrating problems with COVID (1) and sepsis (2) prediction models have helped raise awareness (3) about the need for better practices in developing and reporting machine learning (ML; or artificial intelligence) methods in healthcare. This so-called reproducibility (or replication) crisis has been recognized and extends beyond clinical applications (4). In fact, this issue is not specific to ML. Many scientific journals, including Clinical Chemistry, have adopted principles (specified in the article submission guidelines) to facilitate reproducibility, rigor, and transparency in published findings. Although there are common elements that support transparency and rigor in science, each technology has its own set of pitfalls that must be addressed. This is true for the rapidly developing field of ML. The growth and development of open-source software and digitized and publicly available data sources have made the application of ML methods highly accessible. While this has facilitated a surge in interest and publications, it has reduced the requirement for developers to have the necessary foundational or subject matter knowledge needed for quality publications and innovations. When ML methods are published in clinical journals, peer reviewers and editors may lack the expertise to appropriately evaluate technical aspects of submissions. Thus, we need best practices that help educate clinical experts and govern how ML for laboratory medicine should be developed and communicated. This is critical for ensuring practicality and reproducibility of ML applications and, ultimately, their successful translation to clinical practice. Regulatory guidance for ML-based diagnostics is evolving. For example, the US Food and Drug Administration recently outlined its plan for regulating artificial intelligence/ML-based software as a medical device (5). The Food and Drug Administration maintains a list of approved artificial intelligence/ML-enabled devices (6). The current reality in laboratory medicine (and many other clinical specialties outside of radiology) is that there are relatively few examples of ML-based products that have gone through the Food and Drug Administration’s proposed review pathway or have been implemented as lab developed tests or decision support tools following some type of validation. Most ML applications are limited to publication in scientific journals, and several groups have proposed checklists that describe the minimum information required when reporting ML methods and points to be considered in their review (4, 7–11). Here, we draw from this prior work to summarize practices (Table 1) that should be applied in the development, reporting, and review of ML applications. Summary of recommended practices for ML model development activities. State the objective of the ML problem, specifying inputs and outputs, with its prediction methodology. Provide adequate detail about the clinical scenario and population where the model applies. Describe details of data collection/aggregation, preparation, and case labeling steps. Use large and diverse data sets that are representative of the intended population and scenario. Provide descriptive statistics with demographics before and after processing steps for training and validation data sets. Use data preparation methods that prevent outcome and validation data leakage. Perform 5- or 10-fold cross validation versus simple train/test splits. Investigate a range of model types and tune hyperparameters. Provide rationale for final selection. Report performance metrics from validation set predictions. Evaluate model performance and clinical utility or impact. Report metrics appropriate for the clinical decision (e.g., area under the precision recall curve, positive predictive value, negative predictive value, number needed to alert). Include external validation for final method performance evaluation and note any recalibration strategies used or indicated. Interpre","journal":"Clinical Chemistry","year":2021,"id":166486,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":34,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9528,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":693022,"name":"Stephen R Master","orcid":null,"position":1,"is_corresponding":false},{"id":693021,"name":"Shannon Haymond","orcid":null,"position":0,"is_corresponding":true}],"reference_count":13,"raw_metadata":null,"created_at":"2026-07-18T23:45:49.644768Z","pmid":"35019992","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}