{"doi":"10.1016/j.jmb.2025.169538","title":"Structure-Based Classification of CRISPR/Cas9 Proteins: A Machine Learning Approach to Elucidating Cas9 Allostery","abstract":"The CRISPR/Cas9 system is a powerful gene-editing tool. Its specificity and stability rely on complex allosteric regulation. Understanding these allosteric regulations is essential for developing high-fidelity Cas9 variants with reduced off-target effects. Here, we used a novel structure-based machine learning (ML) approach to systematically identify long-range allosteric networks in Cas9. Our ML model was trained using all available Cas9 structures, ensuring a comprehensive representation of Cas9's structural landscape. We then applied this model to Streptococcus pyogenes Cas9 (SpCas9) to demonstrate the feature selection process. Using Cα-Cα inter-residue distances, we mapped key allosteric networks and refined them through a two-stage SHAP feature selection (FS) strategy, reducing a vast feature space to 28 critical Lysine-Arginine (Lys-Arg) residue pairs that mediate SpCas9 interdomain communication, stability, and specificity. These Lys-Arg pairs initially shared a 46.5 Å inter-residue distance, but molecular dynamics simulations revealed distinct stabilization behaviors, indicating a hierarchical allosteric network. Further mutational analysis of R78A-K855A (M1) and R765A-K1246A (M2) identified an \"electrostatic valley,\" a stabilizing network where positively charged residues interact with negatively charged DNA to maintain SpCas9's structural integrity. Disrupting this valley through direct (M2) or allosteric (M1) mutations destabilized SpCas9's DNA-bound conformation, leading to distinct pathways for improving SpCas9 specificity. This study provides a new framework for understanding allostery in Cas9, integrating ML-driven structural analysis with MD simulations. By identifying key allosteric residues and introducing the electrostatic valley as a central concept, we offer a rational strategy for engineering high-fidelity Cas9 variants. Beyond Cas9, our approach can be applied to uncover allosteric hotspots in other enzyme regulations and rational protein design.","journal":"Journal of Molecular Biology","year":2025,"id":580734,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9556,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1162555,"name":"Vindi M. Jayasinghe‐Arachchige","orcid":"0000-0002-5493-6328","position":1,"is_corresponding":false},{"id":1492195,"name":"Charlene R Norgan Radler","orcid":null,"position":2,"is_corresponding":false},{"id":420647,"name":"Shouyi Wang","orcid":"0000-0002-9046-6474","position":3,"is_corresponding":false},{"id":497425,"name":"Jin Liu","orcid":"0000-0002-1067-4063","position":4,"is_corresponding":false},{"id":1011439,"name":"Sita Sirisha Madugula","orcid":"0000-0001-9944-117X","position":0,"is_corresponding":true}],"reference_count":70,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T02:58:43.046112Z","pmid":"41218722","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}