{"doi":"10.1093/jamia/ocae088","title":"Development and external validation of deep learning clinical prediction models using variable-length time series data","abstract":"OBJECTIVES: To compare and externally validate popular deep learning model architectures and data transformation methods for variable-length time series data in 3 clinical tasks (clinical deterioration, severe acute kidney injury [AKI], and suspected infection). MATERIALS AND METHODS: This multicenter retrospective study included admissions at 2 medical centers that spanned 2007-2022. Distinct datasets were created for each clinical task, with 1 site used for training and the other for testing. Three feature engineering methods (normalization, standardization, and piece-wise linear encoding with decision trees [PLE-DTs]) and 3 architectures (long short-term memory/gated recurrent unit [LSTM/GRU], temporal convolutional network, and time-distributed wrapper with convolutional neural network [TDW-CNN]) were compared in each clinical task. Model discrimination was evaluated using the area under the precision-recall curve (AUPRC) and the area under the receiver operating characteristic curve (AUROC). RESULTS: The study comprised 373 825 admissions for training and 256 128 admissions for testing. LSTM/GRU models tied with TDW-CNN models with both obtaining the highest mean AUPRC in 2 tasks, and LSTM/GRU had the highest mean AUROC across all tasks (deterioration: 0.81, AKI: 0.92, infection: 0.87). PLE-DT with LSTM/GRU achieved the highest AUPRC in all tasks. DISCUSSION: When externally validated in 3 clinical tasks, the LSTM/GRU model architecture with PLE-DT transformed data demonstrated the highest AUPRC in all tasks. Multiple models achieved similar performance when evaluated using AUROC. CONCLUSION: The LSTM architecture performs as well or better than some newer architectures, and PLE-DT may enhance the AUPRC in variable-length time series data for predicting clinical outcomes during external validation.","journal":"Journal of the American Medical Informatics Association","year":2024,"id":428159,"datarank":0.4944403070912921,"base_score":2.639057329615259,"endowment":2.639057329615259,"self_citation_contribution":0.3958585994422889,"citation_network_contribution":0.0985817076490032,"self_endowment_contribution":0.3958585994422889,"citer_contribution":0.0985817076490032,"corpus_percentile":61.19749361800882,"corpus_rank":5017,"citation_count":13,"citer_count":12,"citers_with_citation_signal":6,"citers_with_endowment":6,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.6029,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":252284,"name":"Kyle A. Carey","orcid":"0000-0002-5038-6074","position":1,"is_corresponding":false},{"id":1229342,"name":"Jennie Martin","orcid":"0009-0004-8488-2887","position":2,"is_corresponding":false},{"id":108946,"name":"Jay L. Koyner","orcid":"0000-0001-6873-8712","position":3,"is_corresponding":false},{"id":253346,"name":"Dana P. Edelson","orcid":null,"position":4,"is_corresponding":false},{"id":253342,"name":"Emily Gilbert","orcid":null,"position":5,"is_corresponding":false},{"id":692622,"name":"Anoop Mayampurath","orcid":"0000-0002-3010-6960","position":6,"is_corresponding":false},{"id":252285,"name":"Majid Afshar","orcid":"0000-0002-6368-4652","position":7,"is_corresponding":false},{"id":252287,"name":"Matthew M. Churpek","orcid":"0000-0002-4030-5250","position":8,"is_corresponding":false},{"id":947341,"name":"Fereshteh S. Bashiri","orcid":"0000-0002-0153-4958","position":0,"is_corresponding":true}],"reference_count":32,"raw_metadata":null,"created_at":"2026-07-19T01:58:57.592578Z","pmid":"38679906","pmcid":"PMC11105134","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}