{"doi":"10.1002/ctm2.70108","title":"From data to decision: Scaling artificial intelligence with informatics for epilepsy management","abstract":"The integration of artificial intelligence (AI) into epilepsy research presents a critical opportunity to revolutionize the management of this complex neurological disorder.1 Despite significant advancements in developing AI algorithms to diagnose and manage epilepsy, their translation into clinical practice remains limited. This gap underscores the urgent need for scalable AI and neuroinformatics approaches that can bridge the divide between research and real-world application.2 The ability to generalize AI models from controlled research environments to diverse clinical settings is crucial. Current efforts have made substantial progress, but they also reveal common pitfalls, such as overestimation of model performance due to data leakage and the challenges of small sample sizes, which hinder the generalization of these models. To address these challenges and fully realize the potential of AI in epilepsy care, a robust framework for data sharing and collaboration across research centres is essential. Cloud-based informatics platforms offer a promising solution by enabling the aggregation and harmonization of large, multisite datasets. These platforms can facilitate the development of AI models that are not only powerful but also scalable and generalizable across different patient populations and clinical scenarios. In this commentary, we will explore the common methodological errors that lead to overly optimistic AI models in epilepsy research and propose strategies to overcome these issues. We will also discuss the importance of collaborative data sharing in building robust, clinically relevant AI tools and highlight the role of advanced neuroinformatics infrastructures in supporting the translational pathway from research to clinical practice (Figure 1). The promise of AI in epilepsy research is often hampered by methodological errors that lead to overly optimistic performance metrics. One of the most significant issues is data leakage, which occurs when information from outside the training dataset influences the model, resulting in an overestimation of its predictive power. This can happen when features are derived from the entire dataset rather than just the training subset.3 To mitigate this, strict separation between training and test datasets is essential and feature selection must be performed within each fold of the cross-validation process independently. Nested cross-validation, where model selection and performance estimation are conducted separately, further reduces the risk of data leakage. Another common error is the improper application of cross-validation techniques. Often, researchers perform feature selection or hyperparameter tuning on the entire dataset before cross-validation, leading to inflated performance metrics. The correct approach is to embed these steps within each fold of the cross-validation process to ensure that the test data remain completely unseen until the final evaluation. This practice helps prevent overfitting and provides a more accurate estimate of how the model will perform on new data. Small sample size presents a third challenge, particularly in epilepsy research, where datasets are often of modest size and heterogeneous. Small datasets can lead to overfitting, where the model learns patterns specific to the training data but fails to generalize to new data. Addressing this requires both methodological rigour and collaborative efforts to pool data across multiple sites, thereby creating larger, more diverse datasets. Data augmentation techniques, such as generating synthetic data, can also help increase the effective size of the training set. The development of robust AI models in epilepsy is further strengthened by collaborative data sharing, which allows researchers to pool datasets from multiple sources, increasing both the size and diversity of the data available for training. Epilepsy is a highly heterogeneous disorder, and individual research centres often have access to onl","journal":"Clinical and Translational Medicine","year":2024,"id":507496,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9582,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":324340,"name":"Alfredo Lucas","orcid":"0000-0001-9439-735X","position":1,"is_corresponding":false},{"id":265854,"name":"Kathryn A. Davis","orcid":"0000-0002-7020-6480","position":2,"is_corresponding":false},{"id":644407,"name":"Nishant Sinha","orcid":"0000-0002-2090-4889","position":0,"is_corresponding":true}],"reference_count":10,"raw_metadata":null,"created_at":"2026-07-19T02:11:02.460057Z","pmid":"39673123","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}