{"doi":"10.1016/j.mcpro.2023.100591","title":"UniProt and Mass Spectrometry-Based Proteomics—A 2-Way Working Relationship","abstract":"•High-quality, comprehensive, free resource of protein sequence and function.•Reference sets of whole-proteome sequence, including human, made available.•Includes isoforms, postprocessed chains, PTMs, variants, and functional domains.•Imports and displays peptide and PTM identifications from public domain repositories.•Access via website, API or ftp, download data in multiple community data formats. The human proteome comprises of all of the proteins produced by the sequences translated from the human genome with additional modifications in both sequence and function caused by nonsynonymous variants and posttranslational modifications including cleavage of the initial transcript into smaller peptides and polypeptides. The UniProtKB database (www.uniprot.org) is the world’s leading high-quality, comprehensive and freely accessible resource of protein sequence and functional information and presents a summary of experimentally verified, or computationally predicted, functional information added by our expert biocuration team for each protein in the proteome. Researchers in the field of mass spectrometry–based proteomics both consume and add to the body of data available in UniProtKB, and this review highlights the information we provide to this community and the knowledge we in turn obtain from groups via deposition of large-scale datasets in public domain databases. The human proteome comprises of all of the proteins produced by the sequences translated from the human genome with additional modifications in both sequence and function caused by nonsynonymous variants and posttranslational modifications including cleavage of the initial transcript into smaller peptides and polypeptides. The UniProtKB database (www.uniprot.org) is the world’s leading high-quality, comprehensive and freely accessible resource of protein sequence and functional information and presents a summary of experimentally verified, or computationally predicted, functional information added by our expert biocuration team for each protein in the proteome. Researchers in the field of mass spectrometry–based proteomics both consume and add to the body of data available in UniProtKB, and this review highlights the information we provide to this community and the knowledge we in turn obtain from groups via deposition of large-scale datasets in public domain databases. Mass spectrometry (MS)-based proteomics researchers aim to isolate, identify, and functionally characterize the protein profile of a cell, tissue, or organism of interest. Scientists may be interested in understanding how this profile changes as cells develop, age, respond to changes in external environmental conditions, or how the application of a xenobiotic, such as a drug, changes cellular behavior at the molecular level. In order for researchers to be able to determine the protein complement of a biological sample, they need to be able to map peptide or protein sequences to a reference resource, which not only enables the required identifications but also links these to the relevant organism and provides details of protein nomenclature, function, and additional sequence features such as posttranslational modifications (PTMs) and functional domains. The UniProt data resource (1UniProt ConsortiumUniProt: the universal protein knowledgebase in 2023.Nucl. Acids Res. 2022; https://doi.org/10.1093/nar/gkac1052Crossref Scopus (713) Google Scholar) provides a complete compendium of all known protein sequence data linked to a summary of the experimentally verified, or computationally predicted, functional information about that protein and is the most widely used protein sequence database employed by the proteomics community. This review is aimed at enabling proteomics scientists to not only download the optimal set of sequences for use in their search engine(s) but also to subsequently access as much information on their identified proteins as possible. It also describes how MS data are used to inform an","journal":"Molecular & Cellular Proteomics","year":2023,"id":324671,"datarank":0.8963241713031699,"base_score":3.5263605246161616,"endowment":3.5263605246161616,"self_citation_contribution":0.5289540786924243,"citation_network_contribution":0.3673700926107456,"self_endowment_contribution":0.5289540786924243,"citer_contribution":0.3673700926107456,"corpus_percentile":77.48897656068694,"corpus_rank":2911,"citation_count":33,"citer_count":30,"citers_with_citation_signal":22,"citers_with_endowment":22,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9507,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":103487,"name":"Jun Fan","orcid":"0000-0003-3899-2191","position":1,"is_corresponding":false},{"id":103499,"name":"Jie Luo","orcid":"0000-0001-9404-1302","position":2,"is_corresponding":false},{"id":92022,"name":"Michele Magrane","orcid":"0000-0003-3544-996X","position":3,"is_corresponding":false},{"id":57226,"name":"María Martin","orcid":"0000-0001-5454-2815","position":4,"is_corresponding":false},{"id":5925,"name":"Sandra Orchard","orcid":"0000-0002-8878-3972","position":5,"is_corresponding":false},{"id":103478,"name":"Emily Bowler-Barnett","orcid":"0000-0003-4785-7231","position":0,"is_corresponding":true}],"reference_count":43,"raw_metadata":null,"created_at":"2026-07-19T01:08:16.809511Z","pmid":"37301379","pmcid":"PMC10404557","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}