{"doi":"10.1101/2025.06.08.658499","title":"Simpatico: accurate and ultra-fast virtual drug screening with atomic embeddings","abstract":"Abstract Virtual screening, the in-silico assessment of large libraries of small molecules for binding to a therapeutic protein target, is a critical early step in drug discovery. The dominant approach, molecular docking, requires a separate calculation for each protein-molecule pair, and is too slow to apply alone at the billion-compound scale of modern compound libraries. A recent embedding-retrieval paradigm, exemplified by DrugCLIP, addresses this bottleneck by training deep models to map proteins and small molecules into a shared embedding space, such that proteins are co-located with their likely binding partners; candidate ligands can then be retrieved directly by nearest-neighbors search, with no per-pair calculation. However, current embedding-retrieval methods collapse each protein and each ligand into a single embedding, creating an information bottleneck that limits the representation of partial or alternative binding compatibility. We present simpatico , an embedding-retrieval virtual screening tool that instead produces a unique embedding for each atom in a protein pocket or ligand. A CLIP-style contrastive objective trains these atomic embeddings so that protein-ligand atom pairs known to interact are nearby in embedding space. To screen a protein target, each protein-atom embedding is used as a query against a vector database of precomputed small-molecule atomic embeddings, returning the closest atoms in the library; a simple aggregation step assigns a binding score to each candidate molecule containing retrieved atoms. Where prior retrieval-based methods index one vector per ligand, simpatico indexes one per heavy atom; query time grows sublinearly in library size. On challenging decoy benchmarks, simpatico achieves state-of-the-art predictive accuracy, outperforming recent dense-retrieval methods despite training on only ∼15,000 protein-ligand complexes from PDBBind, with no pretraining and no 3D ligand pose estimation. Simpatico also exceeds the accuracy of physics-based docking and deep-learning-augmented docking methods, is competitive with diffusion-based docking, and runs orders of magnitude faster than all three. Simpatico is open source software; all code, weights, and data may be accessed at https://github.com/TravisWheelerLab/Simpatico .","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2025,"id":567841,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.95,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":264756,"name":"Travis J. Wheeler","orcid":"0000-0003-2004-1785","position":1,"is_corresponding":false},{"id":912457,"name":"Jeremiah Gaiser","orcid":"0000-0003-3046-8686","position":0,"is_corresponding":true}],"reference_count":27,"raw_metadata":null,"created_at":"2026-07-19T02:56:48.183032Z","pmid":"40661404","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}