{"doi":"10.1093/gigascience/giad030","title":"Strategies and techniques for quality control and semantic enrichment with multimodal data: a case study in colorectal cancer with eHDPrep","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:sec>\n                    <jats:title>Background</jats:title>\n                    <jats:p>Integration of data from multiple domains can greatly enhance the quality and applicability of knowledge generated in analysis workflows. However, working with health data is challenging, requiring careful preparation in order to support meaningful interpretation and robust results. Ontologies encapsulate relationships between variables that can enrich the semantic content of health datasets to enhance interpretability and inform downstream analyses.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Findings</jats:title>\n                    <jats:p>We developed an R package for electronic health data preparation, “eHDPrep,” demonstrated upon a multimodal colorectal cancer dataset (661 patients, 155 variables; Colo-661); a further demonstrator is taken from The Cancer Genome Atlas (459 patients, 94 variables; TCGA-COAD). eHDPrep offers user-friendly methods for quality control, including internal consistency checking and redundancy removal with information-theoretic variable merging. Semantic enrichment functionality is provided, enabling generation of new informative “meta-variables” according to ontological common ancestry between variables, demonstrated with SNOMED CT and the Gene Ontology in the current study. eHDPrep also facilitates numerical encoding, variable extraction from free text, completeness analysis, and user review of modifications to the dataset.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions</jats:title>\n                    <jats:p>eHDPrep provides effective tools to assess and enhance data quality, laying the foundation for robust performance and interpretability in downstream analyses. Application to multimodal colorectal cancer datasets resulted in improved data quality, structuring, and robust encoding, as well as enhanced semantic information. We make eHDPrep available as an R package from CRAN (https://cran.r-project.org/package=eHDPrep) and GitHub (https://github.com/overton-group/eHDPrep).</jats:p>\n                  </jats:sec>","journal":"GigaScience","year":2022,"id":607246,"datarank":0.24141568686511508,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"self_citation_contribution":0.24141568686511508,"citation_network_contribution":0.0,"self_endowment_contribution":0.24141568686511508,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1559187,"name":"Rashi Pancholi","orcid":null,"position":1,"is_corresponding":false},{"id":1111576,"name":"Paul Miller","orcid":"0000-0002-5800-1220","position":2,"is_corresponding":false},{"id":1559188,"name":"Thorsten Forster","orcid":null,"position":3,"is_corresponding":false},{"id":336739,"name":"Helen G. Coleman","orcid":"0000-0003-4872-7877","position":4,"is_corresponding":false},{"id":54360,"name":"Ian M. Overton","orcid":"0000-0003-1158-8527","position":5,"is_corresponding":false},{"id":1559186,"name":"Tom M Toner","orcid":"0000-0001-8059-5822","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Strategies and techniques for quality control and semantic enrichment with multimodal data: a case study in colorectal cancer with eHDPrep","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:sec>\n                    <jats:title>Background</jats:title>\n                    <jats:p>Integration of data from multiple domains can greatly enhance the quality and applicability of knowledge generated in analysis workflows. However, working with health data is challenging, requiring careful preparation in order to support meaningful interpretation and robust results. Ontologies encapsulate relationships between variables that can enrich the semantic content of health datasets to enhance interpretability and inform downstream analyses.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Findings</jats:title>\n                    <jats:p>We developed an R package for electronic health data preparation, “eHDPrep,” demonstrated upon a multimodal colorectal cancer dataset (661 patients, 155 variables; Colo-661); a further demonstrator is taken from The Cancer Genome Atlas (459 patients, 94 variables; TCGA-COAD). eHDPrep offers user-friendly methods for quality control, including internal consistency checking and redundancy removal with information-theoretic variable merging. Semantic enrichment functionality is provided, enabling generation of new informative “meta-variables” according to ontological common ancestry between variables, demonstrated with SNOMED CT and the Gene Ontology in the current study. eHDPrep also facilitates numerical encoding, variable extraction from free text, completeness analysis, and user review of modifications to the dataset.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions</jats:title>\n                    <jats:p>eHDPrep provides effective tools to assess and enhance data quality, laying the foundation for robust performance and interpretability in downstream analyses. Application to multimodal colorectal cancer datasets resulted in improved data quality, structuring, and robust encoding, as well as enhanced semantic information. We make eHDPrep available as an R package from CRAN (https://cran.r-project.org/package=eHDPrep) and GitHub (https://github.com/overton-group/eHDPrep).</jats:p>\n                  </jats:sec>","is_dataset_classified":null,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"37171130","pmcid":"PMC10176503","openalex_id":"https://openalex.org/W4376226077","authors":[],"funders":[{"funder_name":"Engineering and Physical Sciences Research Council","grant_id":"2280988","title":null},{"funder_name":"Cancer Research UK","grant_id":"15333","title":null},{"funder_name":"Public Health Agency","grant_id":"SPI/5151/15","title":null},{"funder_name":"Wellcome Trust","grant_id":"","title":null},{"funder_name":"Medical Research Council","grant_id":"","title":null},{"funder_name":"Chief Scientist Office","grant_id":"","title":null},{"funder_name":"British Heart Foundation","grant_id":"","title":null},{"funder_name":"Medical Research Council UK","grant_id":"","title":null},{"funder_name":"Wellcome Trust","grant_id":"","title":null},{"funder_name":"British Heart Foundation","grant_id":"","title":null},{"funder_name":"Medical Research Council","grant_id":"","title":null},{"funder_name":"Chief Scientist Office","grant_id":"","title":null}],"total_grants":12,"fwci":0.3135,"citation_percentile":0.52417858,"influential_citations":0,"citation_trend":[{"year":2022,"count":1},{"year":2025,"count":2}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://academic.oup.com/gigascience/article-pdf/doi/10.1093/gigascience/giad030/50383140/giad030.pdf","host_type":"journal"},{"url":"https://academic.oup.com/gigascience/article-pdf/doi/10.1093/gigascience/giad030/50383140/giad030.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/gigascience/article-pdf/doi/10.1093/gigascience/giad030/60705693/giad030.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/gigascience/giad030","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/37171130","host_type":"repository"},{"url":"https://pure.qub.ac.uk/en/publications/f0bc4292-838c-44b5-8a11-ca2f2bf6867c","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/10176503","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10176503/pdf/giad030.pdf","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC10176503","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC10176503?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Biomedical Text Mining and Ontologies","Genomics and Rare Diseases","Bioinformatics and Genomic Networks","Humans","Semantics","Gene Ontology","Data Accuracy","Quality Control","Colorectal Neoplasms"],"mesh_terms":["Data Accuracy","Humans","Quality Control","Semantics","Colorectal Neoplasms","Gene Ontology"],"keywords":["Interpretability","Computer science","Information retrieval","Data mining","Workflow","Data integration","Data quality","Ontology","Consistency (knowledge bases)","Machine learning","Artificial intelligence","Database","Quality control","Bioinformatics","Quality assessment","Colorectal Cancer","Medical Informatics","Health Data","Semantic Enrichment"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-30T06:02:02.265852Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}