{"doi":"10.1101/624494","title":"Proteomics Standards Initiative Extended FASTA Format (PEFF)","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  Mass spectrometry-based proteomics enables the high-throughput identification and quantification of proteins, including sequence variants and post-translational modifications (PTMs), in biological samples. However, most workflows require that such variations be included in the search space used to analyze the data, and doing so remains challenging with most analysis tools. In order to facilitate the search for known sequence variants and PTMs, the Proteomics Standards Initiative (PSI) has designed and implemented the PSI Extended FASTA Format (PEFF). PEFF is based on the very popular FASTA format but adds a uniform mechanism for encoding substantially more metadata about the sequence collection as well as individual entries, including support for encoding known sequence variants, PTMs, and proteoforms. The format is very nearly backwards compatible, and as such, existing FASTA parsers will require little or no changes to be able to read PEFF files as FASTA files, although without supporting any of the extra capabilities of PEFF. PEFF is defined by a full specification document, controlled vocabulary terms, a set of example files, software libraries, and a file validator. Popular software and resources are starting to support PEFF, including the sequence search engine Comet and the knowledge bases neXtProt and UniProtKB. Widespread implementation of PEFF is expected to further enable proteogenomics and top-down proteomics applications by providing a standardized mechanism for encoding protein sequences and their known variations. All the related documentation, including the detailed file format specification and example files, are available at\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"http://www.psidev.info/peff\">http://www.psidev.info/peff</jats:ext-link>\n                  .\n                </jats:p>","journal":null,"year":null,"id":621034,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":558129,"name":"Jim Shofstahl","orcid":"0000-0001-5968-1742","position":1,"is_corresponding":false},{"id":17887,"name":"Juan Antonio Vizcaino","orcid":"0000-0002-3905-4335","position":2,"is_corresponding":false},{"id":530587,"name":"Harald Barsnes","orcid":"0000-0001-9696-2148","position":3,"is_corresponding":false},{"id":628461,"name":"Robert J. Chalkley","orcid":"0000-0002-9757-7302","position":4,"is_corresponding":false},{"id":17853,"name":"Gerben Menschaert","orcid":"0000-0002-7575-2085","position":5,"is_corresponding":false},{"id":103477,"name":"Emanuele Alpi","orcid":"0000-0003-4822-9472","position":6,"is_corresponding":false},{"id":578045,"name":"Karl Clauser","orcid":null,"position":7,"is_corresponding":false},{"id":55406,"name":"Jimmy K. Eng","orcid":"0000-0001-6352-6737","position":8,"is_corresponding":false},{"id":121997,"name":"Lydie Lane","orcid":"0000-0002-9818-3030","position":9,"is_corresponding":false},{"id":1603447,"name":"Sean L. Seymour","orcid":"0000-0002-9571-8220","position":10,"is_corresponding":false},{"id":1603448,"name":"Luis Francisco Hernández Sánchez","orcid":"0000-0003-2809-1517","position":11,"is_corresponding":false},{"id":79930,"name":"Gerhard Mayer","orcid":"0000-0002-1767-2343","position":12,"is_corresponding":false},{"id":79931,"name":"Martin Eisenacher","orcid":"0000-0003-2687-7444","position":13,"is_corresponding":false},{"id":19125,"name":"Yasset Perez-Riverol","orcid":"0000-0001-6579-6941","position":14,"is_corresponding":false},{"id":1603449,"name":"Eugene A. Kapp","orcid":"0000-0002-3283-4259","position":15,"is_corresponding":false},{"id":557229,"name":"Luis Mendoza","orcid":"0000-0003-0128-8643","position":16,"is_corresponding":false},{"id":819020,"name":"Peter R. Baker","orcid":"0000-0001-7826-9251","position":17,"is_corresponding":false},{"id":593456,"name":"Andrew Collins","orcid":"0000-0001-5916-6871","position":18,"is_corresponding":false},{"id":557230,"name":"Tim Van Den Bossche","orcid":"0000-0002-5916-2587","position":19,"is_corresponding":false},{"id":6362,"name":"Eric W. Deutsch","orcid":"0000-0001-8732-0928","position":20,"is_corresponding":false},{"id":6396,"name":"Pierre‐Alain Binz","orcid":"0000-0002-0045-7698","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Proteomics Standards Initiative Extended FASTA Format (PEFF)","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  Mass spectrometry-based proteomics enables the high-throughput identification and quantification of proteins, including sequence variants and post-translational modifications (PTMs), in biological samples. However, most workflows require that such variations be included in the search space used to analyze the data, and doing so remains challenging with most analysis tools. In order to facilitate the search for known sequence variants and PTMs, the Proteomics Standards Initiative (PSI) has designed and implemented the PSI Extended FASTA Format (PEFF). PEFF is based on the very popular FASTA format but adds a uniform mechanism for encoding substantially more metadata about the sequence collection as well as individual entries, including support for encoding known sequence variants, PTMs, and proteoforms. The format is very nearly backwards compatible, and as such, existing FASTA parsers will require little or no changes to be able to read PEFF files as FASTA files, although without supporting any of the extra capabilities of PEFF. PEFF is defined by a full specification document, controlled vocabulary terms, a set of example files, software libraries, and a file validator. Popular software and resources are starting to support PEFF, including the sequence search engine Comet and the knowledge bases neXtProt and UniProtKB. Widespread implementation of PEFF is expected to further enable proteogenomics and top-down proteomics applications by providing a standardized mechanism for encoding protein sequences and their known variations. All the related documentation, including the detailed file format specification and example files, are available at\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"http://www.psidev.info/peff\">http://www.psidev.info/peff</jats:ext-link>\n                  .\n                </jats:p>","is_dataset_classified":null,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19910364","pmcid":null,"openalex_id":"https://openalex.org/W2942833306","authors":[],"funders":[{"funder_name":"Wellcome Trust","grant_id":"101477","title":"PRIDE Atlas."},{"funder_name":"National Institutes of Health","grant_id":"1R24GM127667-01","title":"Advancing data and metadata standards for proteomics mass spectra"},{"funder_name":"National Institutes of Health","grant_id":"5U54EB020406-04","title":"Big Data for Discovery Science"},{"funder_name":"National Institutes of Health","grant_id":"5U19AG023122-04","title":"Consortium to Study the Genetics of Longevity"},{"funder_name":"National Institutes of Health","grant_id":"3U24HG007822-05S2","title":"Identification of synaptic gene sets for CNS synapse taxonomy and antibodies for their engagement"},{"funder_name":"Wellcome Trust","grant_id":"208391","title":"The PRIDE database: A proteomics data hub in the life sciences"},{"funder_name":"National Institutes of Health","grant_id":"5U24HG007822-12","title":"UniProt: A Protein Sequence and Function Resource for Biomedical Science"},{"funder_name":"National Institutes of Health","grant_id":"3U54EB020406-03S1","title":"Big Data for Discovery Science"},{"funder_name":"National Institutes of Health","grant_id":"5R01GM087221-10","title":"Development of Trans Proteomic Pipeline, an Analysis Suite for Mass Spectrometry"}],"total_grants":9,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2020,"count":2}],"oa_status":"green","license":"cc-by","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2019/05/08/624494.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2019/05/08/624494.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/624494","host_type":"publisher"},{"url":"https://doi.org/10.1101/624494","host_type":"repository"},{"url":"http://hdl.handle.net/1854/LU-01JAT7KNH6TJ4384B64ACEESZA","host_type":"repository"},{"url":"https://dx.doi.org/10.1101/624494","host_type":""},{"url":"http://dx.doi.org/10.1101/624494","host_type":""}],"fields_of_study":["Advanced Proteomics Techniques and Applications","Mass Spectrometry Techniques and Applications","Genomics and Phylogenetic Studies","0301 basic medicine","03 medical and health sciences","0206 medical engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Computer science","Metadata","Documentation","UniProt","Software","Proteomics","File format","Information retrieval","World Wide Web","Database","Biology","Programming language","Genetics"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Partnerships for the goals"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-03T12:59:05.044325Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}