{"doi":"10.1101/2020.05.19.102285","title":"Deep learning the collisional cross sections of the peptide universe from a million training samples","abstract":"<h4>ABSTRACT</h4> The size and shape of peptide ions in the gas phase are an under-explored dimension for mass spectrometry-based proteomics. To explore the nature and utility of the entire peptide collisional cross section (CCS) space, we measure more than a million data points from whole-proteome digests of five organisms with trapped ion mobility spectrometry (TIMS) and parallel accumulation – serial fragmentation (PASEF). The scale and precision (CV <1%) of our data is sufficient to train a deep recurrent neural network that accurately predicts CCS values solely based on the peptide sequence. Cross section predictions for the synthetic ProteomeTools library validate the model within a 1.3% median relative error (R > 0.99). Hydrophobicity, position of prolines and histidines are main determinants of the cross sections in addition to sequence-specific interactions. CCS values can now be predicted for any peptide and organism, forming a basis for advanced proteomics workflows that make full use of the additional information.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2020,"id":9024,"datarank":0.32544306105105864,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.05667914066685035,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.05667914066685035,"corpus_percentile":47.42786416028468,"corpus_rank":6797,"citation_count":5,"citer_count":3,"citers_with_citation_signal":2,"citers_with_endowment":2,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.5896,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-05-21","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":20758,"name":"Niklas D. Köhler","orcid":"0000-0003-2726-0518","position":1,"is_corresponding":false},{"id":20986,"name":"Andreas-David Brunner","orcid":"0000-0002-2733-7899","position":2,"is_corresponding":false},{"id":77576,"name":"Jean-Marc H. Wanka","orcid":null,"position":3,"is_corresponding":false},{"id":77577,"name":"Eugenia Voytik","orcid":"0000-0003-4776-0771","position":4,"is_corresponding":false},{"id":77578,"name":"Maximilian T. Strauss","orcid":"0000-0003-3320-6833","position":5,"is_corresponding":false},{"id":42,"name":"Fabian Joachim Theis","orcid":"0000-0002-2419-1943","position":6,"is_corresponding":false},{"id":2937,"name":"Matthias Mann","orcid":"0000-0003-1292-4799","position":7,"is_corresponding":false},{"id":49057,"name":"Florian Meier","orcid":"0000-0003-4451-5025","position":0,"is_corresponding":true}],"reference_count":72,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-03-01T18:20:47.508186Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}