{"doi":"10.1101/2021.05.10.443525","title":"Systematic tissue annotations of –omics samples by modeling unstructured metadata","abstract":"Abstract There are currently &gt;1.3 million human –omics samples that are publicly available. This valuable resource remains acutely underused because discovering particular samples from this ever-growing data collection remains a significant challenge. The major impediment is that sample attributes are routinely described using varied terminologies written in unstructured natural language. We propose a natural-language-processing-based machine learning approach (NLP-ML) to infer tissue and cell-type annotations for –omics samples based only on their free-text metadata. NLP-ML works by creating numerical representations of sample descriptions and using these representations as features in a supervised learning classifier that predicts tissue/cell-type terms. Our approach significantly outperforms an advanced graph-based reasoning annotation method (MetaSRA) and a baseline exact string matching method (TAGGER). Model similarities between related tissues demonstrate that NLP-ML models capture biologically-meaningful signals in text. Additionally, these models correctly classify tissue-associated biological processes and diseases based on their text descriptions alone. NLP-ML models are nearly as accurate as models based on gene-expression profiles in predicting sample tissue annotations but have the distinct capability to classify samples irrespective of the –omics experiment type based on their text metadata. Python NLP-ML prediction code and trained tissue models are available at https://github.com/krishnanlab/txt2onto .","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2021,"id":215888,"datarank":0.4057366678020522,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.1369727474178439,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.1369727474178439,"corpus_percentile":null,"corpus_rank":null,"citation_count":5,"citer_count":5,"citers_with_citation_signal":4,"citers_with_endowment":4,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.6273,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":811208,"name":"Marc Maldaver","orcid":"0000-0001-9689-2768","position":1,"is_corresponding":false},{"id":465588,"name":"Anna Yannakopoulos","orcid":"0000-0001-7806-984X","position":2,"is_corresponding":false},{"id":811209,"name":"Lindsay Guare","orcid":"0000-0001-6988-5319","position":3,"is_corresponding":false},{"id":34964,"name":"Krishnan, Arjun","orcid":"0000-0002-7980-4110","position":4,"is_corresponding":false},{"id":811207,"name":"Nathaniel T. Hawkins","orcid":"0000-0002-7193-4113","position":0,"is_corresponding":true}],"reference_count":44,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-18T23:53:00.192353Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}