{"doi":"10.1101/2024.11.04.620735","title":"Localization of Macromolecules in Crowded Cellular Cryo-electron Tomograms from Extremely Sparse Labels","abstract":"Abstract Motivation Localizing macromolecules in crowded cellular cryo-electron tomograms (cryo-ET) is crucial for determining their in situ structures. Traditional template matching-based approaches for this task suffer from template-specific biases and have low throughput. Given these problems, learning-based solutions are necessary. However, the paucity of annotated data for training poses substantial challenges for such learning-based methods. Moreover, preparing extensively annotated cellular cryo-ET tomograms for training macromolecule localization methods is extremely time-consuming and burdensome due to the large volume and low signal-to-noise ratio of the tomograms. Results In this work, we developed TomoPicker, an annotation-efficient macromolecule localization method for cryo-ET tomograms. To achieve such annotation-efficiency, TomoPicker regards macromolecule localization as a voxel classification problem and solves it with two different positive-unlabeled learning approaches. We evaluated TomoPicker on two experimental cryo-electron tomography (cryo-ET) datasets of crowded eukaryotic cells and one experimental dataset of relatively less crowded prokaryotic cell. We observed that, with only 10 annotated macromolecule locations, TomoPicker with positive unlabeled learning achieved a performance comparable to that of state-of-the-art supervised methods trained with several hundred annotations. In other words, TomoPicker achieved plausible segmentation with up to 98% less data compared to supervised learning-based methods. Furthermore, it demonstrated substantial improvements over existing learning-based macromolecule localization methods under sparse annotation scenarios. Code The code to train and use TomoPicker is available on https://github.com/DuranRafid/TomoPicker .","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2024,"id":486432,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9429,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1330216,"name":"Ajmain Yasar Ahmed","orcid":"0009-0007-2557-9564","position":1,"is_corresponding":false},{"id":1330217,"name":"H.M. Shadman Tabib","orcid":"0009-0008-2890-7268","position":2,"is_corresponding":false},{"id":1330218,"name":"Md Toki Tahmid","orcid":"0000-0002-1152-3726","position":3,"is_corresponding":false},{"id":1330652,"name":"Md. Zarif Ul Alam","orcid":null,"position":4,"is_corresponding":false},{"id":307539,"name":"Zachary Freyberg","orcid":"0000-0001-6460-0118","position":5,"is_corresponding":false},{"id":383213,"name":"Min Xu","orcid":"0000-0002-0881-5891","position":6,"is_corresponding":false},{"id":741562,"name":"Mostofa Rafid Uddin","orcid":null,"position":0,"is_corresponding":true}],"reference_count":23,"raw_metadata":null,"created_at":"2026-07-19T02:08:01.404471Z","pmid":"39574774","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}