{"doi":"10.1101/2023.10.13.562258","title":"Negative Binomial Mixture Model for Identification of Noise in Antigen-Specificity Predictions by LIBRA-seq","abstract":"Structured Abstract Motivation LIBRA-seq (linking B cell receptor to antigen specificity by sequencing) provides a powerful tool for interrogating the antigen-specific B cell compartment and identifying antibodies against antigen targets of interest. Identification of noise in LIBRA-seq antigen count data is critical for improving antigen binding predictions for downstream applications including antibody discovery and machine learning technologies. Results In this study, we present a method for denoising LIBRA-seq data by clustering antigen counts into signal and noise components with a negative binomial mixture model. This approach leverages the VRC01 negative control cells included in a recent LIBRA-seq study(Abu-Shmais et al .) to provide a data-driven means for identification of technical noise. We apply this method to a dataset of nine donors representing separate LIBRA-seq experiments and show that our approach provides improved predictions for in vitro antibody-antigen binding when compared to the standard scoring method used in LIBRA-seq, despite variance in data size and noise structure across samples. This development will improve the ability of LIBRA-seq to identify antigen-specific B cells and contribute to providing more reliable datasets for future machine learning based approaches to predicting antibody-antigen binding as the corpus of LIBRA-seq data continues to grow. Availability and Implementation Jupyter notebooks detailing model fitting and figure generation in Python are available at https://github.com/perrywasdin/mixture_model_denoising . Contact Email: Ivelin.Georgiev@Vanderbilt.edu Supplementary Information Supplementary figures are provided in the attached PDF.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2023,"id":412728,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9512,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1064865,"name":"Alexandra A. Abu-Shmais","orcid":"0000-0001-5514-3277","position":1,"is_corresponding":false},{"id":813982,"name":"Michael Irvin","orcid":null,"position":2,"is_corresponding":false},{"id":1142125,"name":"Matthew J. Vukovich","orcid":"0000-0002-7110-9623","position":3,"is_corresponding":false},{"id":290894,"name":"Ivelin S. Georgiev","orcid":"0000-0002-6312-7696","position":4,"is_corresponding":false},{"id":790105,"name":"Perry T. Wasdin","orcid":"0000-0001-7351-2048","position":0,"is_corresponding":true}],"reference_count":19,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T01:21:50.261851Z","pmid":"37904915","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}