{"doi":"10.3389/fendo.2022.853863","title":"Analysis of Half a Billion Datapoints Across Ten Machine-Learning Algorithms Identifies Key Elements Associated With Insulin Transcription in Human Pancreatic Islet Cells","abstract":"Machine learning (ML)-workflows enable unprejudiced/robust evaluation of complex datasets. Here, we analyzed over 490,000,000 data points to compare 10 different ML-workflows in a large (N=11,652) training dataset of human pancreatic single-cell (sc-)transcriptomes to identify genes associated with the presence or absence of insulin transcript(s). Prediction accuracy/sensitivity of each ML-workflow was tested in a separate validation dataset (N=2,913). Ensemble ML-workflows, in particular Random Forest ML-algorithm delivered high predictive power (AUC=0.83) and sensitivity (0.98), compared to other algorithms. The transcripts identified through these analyses also demonstrated significant correlation with insulin in bulk RNA-seq data from human islets. The top-10 features, (including IAPP, ADCYAP1, LDHA and SST ) common to the three Ensemble ML-workflows were significantly dysregulated in scRNA-seq datasets from Ire-1α β-/- mice that demonstrate dedifferentiation of pancreatic β-cells in a model of type 1 diabetes (T1D) and in pancreatic single cells from individuals with type 2 Diabetes (T2D). Our findings provide direct comparison of ML-workflows in big data analyses, identify key elements associated with insulin transcription and provide workflows for future analyses.","journal":"Frontiers in Endocrinology","year":2022,"id":284672,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":5,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.8071,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":820304,"name":"Vinod Thorat","orcid":null,"position":1,"is_corresponding":false},{"id":820012,"name":"Mugdha V. Joglekar","orcid":"0000-0001-5346-2266","position":2,"is_corresponding":false},{"id":820013,"name":"Charlotte X. Dong","orcid":"0000-0002-2726-4968","position":3,"is_corresponding":false},{"id":257323,"name":"Hugo Lee","orcid":null,"position":4,"is_corresponding":false},{"id":820014,"name":"Yi Vee Chew","orcid":"0000-0003-1484-8654","position":5,"is_corresponding":false},{"id":820305,"name":"Adwait Bhave","orcid":null,"position":6,"is_corresponding":false},{"id":820015,"name":"Wayne J. Hawthorne","orcid":"0000-0001-8638-3468","position":7,"is_corresponding":false},{"id":7366,"name":"Feyza Engin","orcid":"0000-0002-7987-5381","position":8,"is_corresponding":false},{"id":820306,"name":"Aniruddha Pant","orcid":null,"position":9,"is_corresponding":false},{"id":669023,"name":"Louise T. Dalgaard","orcid":"0000-0002-3598-2775","position":10,"is_corresponding":false},{"id":962441,"name":"Sharda Bapat","orcid":"0000-0001-5116-0977","position":11,"is_corresponding":false},{"id":820016,"name":"Anandwardhan A. Hardikar","orcid":"0000-0001-5587-2090","position":12,"is_corresponding":false},{"id":820011,"name":"Wilson K.M. Wong","orcid":"0000-0001-7736-6834","position":0,"is_corresponding":true}],"reference_count":44,"raw_metadata":null,"created_at":"2026-07-19T00:29:36.828109Z","pmid":"35399953","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}