{"doi":"10.1101/2025.04.17.649309","title":"Predicting Protein Electrostatics with Protein Language Models","abstract":"Ionization states play crucial roles in protein function, yet predicting protein pKa values remains a formidable challenge despite decades of research. Here we present KaMLESM2 and KaML-ESMC, neural network task heads built on ESM protein language models (pLMs) and trained on the PKAD-3r experimental dataset augmented via GAINES, a latent-space sampling strategy for addressing data scarcity. KaML-ESM2/ESMC significantly outperform current structure- and sequence-based approaches across four benchmarks, achieving root-mean-square errors of about 0.5 units across six titratable residue types in native proteins. Performance degradation on 89 buried engineered OBTRUDEs (iOnizable suBsTitutions foR bUrieD rEsidues) that lack evolutionary support (KaML-ESMC RMSE = 1.89) reveals a key limitation of the current framework, which may be overcome through supervised training. Based on these and other data presented in the work, we hypothesize that protein sequence, through its evolutionary context, encodes not only structure and function but indirectly also electrostatic characteristics. We applied KaML-ESM2 to the human proteome, demonstrating that predicted pKa values can potentially identify functional sites and infer catalytic mechanisms. We offer KaML, a sequence-based, end-to-end platform to support applications spanning biological exploration, drug design, protein engineering, and biomolecular simulation. Although additional research is needed, GAINES could offer a general framework for addressing data scarcity in machine learning approaches for protein-related problems.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2025,"id":555498,"datarank":0.20794415416798362,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"self_citation_contribution":0.20794415416798362,"citation_network_contribution":0.0,"self_endowment_contribution":0.20794415416798362,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9358,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":952051,"name":"Guy W. Dayhoff","orcid":"0000-0003-4271-7125","position":1,"is_corresponding":false},{"id":1038113,"name":"Daniel Kortzak","orcid":"0000-0001-9098-3466","position":2,"is_corresponding":false},{"id":236994,"name":"Zhong‐Yin Zhang","orcid":"0000-0002-3234-0769","position":3,"is_corresponding":false},{"id":439755,"name":"Mingzhe Shen","orcid":"0000-0001-7461-3764","position":0,"is_corresponding":true}],"reference_count":33,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T02:54:59.329539Z","pmid":"41928971","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}