{"doi":"10.1038/s42003-022-03579-3","title":"Variational autoencoders learn transferrable representations of metabolomics data","abstract":"Dimensionality reduction approaches are commonly used for the deconvolution of high-dimensional metabolomics datasets into underlying core metabolic processes. However, current state-of-the-art methods are widely incapable of detecting nonlinearities in metabolomics data. Variational Autoencoders (VAEs) are a deep learning method designed to learn nonlinear latent representations which generalize to unseen data. Here, we trained a VAE on a large-scale metabolomics population cohort of human blood samples consisting of over 4500 individuals. We analyzed the pathway composition of the latent space using a global feature importance score, which demonstrated that latent dimensions represent distinct cellular processes. To demonstrate model generalizability, we generated latent representations of unseen metabolomics datasets on type 2 diabetes, acute myeloid leukemia, and schizophrenia and found significant correlations with clinical patient groups. Notably, the VAE representations showed stronger effects than latent dimensions derived by linear and non-linear principal component analysis. Taken together, we demonstrate that the VAE is a powerful method that learns biologically meaningful, nonlinear, and transferrable latent representations of metabolomics data.","journal":"Communications Biology","year":2022,"id":238518,"datarank":1.3997699524809986,"base_score":4.007333185232471,"endowment":4.007333185232471,"self_citation_contribution":0.6010999777848708,"citation_network_contribution":0.7986699746961277,"self_endowment_contribution":0.6010999777848708,"citer_contribution":0.7986699746961277,"corpus_percentile":null,"corpus_rank":null,"citation_count":54,"citer_count":53,"citers_with_citation_signal":37,"citers_with_endowment":37,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9501,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":816121,"name":"Annalise Schweickart","orcid":"0000-0001-9691-3741","position":1,"is_corresponding":false},{"id":237480,"name":"Leandro Cerchietti","orcid":"0000-0003-0608-1350","position":2,"is_corresponding":false},{"id":241452,"name":"Elisabeth Paietta","orcid":"0000-0002-0561-5544","position":3,"is_corresponding":false},{"id":489811,"name":"Hugo F. Fernández","orcid":"0000-0002-7322-0392","position":4,"is_corresponding":false},{"id":816122,"name":"Hassen Al‐Amin","orcid":"0000-0001-6358-1541","position":5,"is_corresponding":false},{"id":2350,"name":"Karsten Suhre","orcid":"0000-0001-9638-3912","position":6,"is_corresponding":false},{"id":2337,"name":"Jan Krumsiek","orcid":"0000-0003-4734-3791","position":7,"is_corresponding":false},{"id":816562,"name":"Daniel P. Gomari","orcid":null,"position":0,"is_corresponding":true}],"reference_count":65,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T00:22:31.741535Z","pmid":"35773471","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}