{"doi":"10.1101/2023.09.27.559809","title":"The Integration of Proteogenomics and Ribosome Profiling Circumvents Key Limitations to Increase the Coverage and Confidence of Novel Microproteins","abstract":"There has been a dramatic increase in the identification of non-conical translation and a significant expansion of the protein-coding genome and proteome. Among the strategies used to identify novel small ORFs (smORFs), Ribosome profiling (Ribo-Seq) is the gold standard for the annotation of novel coding sequences by reporting on smORF translation. In Ribo-Seq, ribosome-protected footprints (RPFs) that map to multiple sites in the genome are computationally removed since they cannot unambiguously be assigned to a specific genomic location, or to a specific transcript in the case of multiple isoforms. Furthermore, RPFs necessarily result in short (25-34 nucleotides) reads, increasing the chance of ambiguous and multi-mapping alignments, such that smORFs that reside in these regions cannot be identified by Ribo-Seq. Here, we show that the inclusion of proteogenomics to create a Ribosome Profiling and Proteogenomics Pipeline (RP3) bypasses this limitation to identify a group of microprotein-encoding smORFs that are missed by current Ribo-Seq pipelines. Moreover, we show that the microproteins identified by RP3 have different sequence compositions from the ones identified by Ribo-Seq-only pipelines, which can affect proteomics identification. In aggregate, the development of RP3 maximizes the detection and confidence of protein-encoding smORFs and microproteins.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2023,"id":395153,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.7853,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":417048,"name":"Angie L. Bookout","orcid":null,"position":1,"is_corresponding":false},{"id":237004,"name":"Christopher A. Barnes","orcid":"0000-0003-2312-7082","position":2,"is_corresponding":false},{"id":291038,"name":"Brendan Miller","orcid":"0000-0002-7461-7973","position":3,"is_corresponding":false},{"id":1169981,"name":"Pablo Machado","orcid":"0000-0001-5616-9583","position":4,"is_corresponding":false},{"id":1169982,"name":"Luiz Augusto Basso","orcid":"0000-0003-0903-2407","position":5,"is_corresponding":false},{"id":1022820,"name":"Cristiano Valim Bizarro","orcid":"0000-0002-2609-8996","position":6,"is_corresponding":false},{"id":17868,"name":"Alan Saghatelian","orcid":"0000-0002-0427-563X","position":7,"is_corresponding":false},{"id":1022819,"name":"Eduardo Vieira de Souza","orcid":"0000-0003-2773-8550","position":0,"is_corresponding":true}],"reference_count":39,"raw_metadata":null,"created_at":"2026-07-19T01:19:18.686170Z","pmid":"37808637","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}