{"doi":"10.1101/2021.08.28.457846","title":"The Need for Transfer Learning in CRISPR-Cas Off-Target Scoring","abstract":"Abstract Motivation The scalable design of safe guide RNA sequences for CRISPR gene editing depends on the computational “scoring” of DNA locations that may be edited. As there is no widely accepted benchmark dataset to compare scoring models, we present a curated “TrueOT” dataset that contains thoroughly validated datapoints to best reflect the properties of in vivo editing. Many existing models are trained on data from high throughput assays. We hypothesize that such models may suboptimally transfer to the low throughput data in TrueOT due to fundamental biological differences between proxy assays and in vivo behavior. We developed new Siamese convolutional neural networks, trained them on a proxy dataset, and compared their performance against existing models on TrueOT. Results Our simplest model with a single convolutional and pooling layer surprisingly exhibits state-of-the-art performance on TrueOT. Adding subsequent layers improved performance on a proxy dataset while compromising performance on TrueOT. We demonstrate improved generalization on TrueOT with a Siamese model of higher complexity when we apply transfer learning techniques. These results suggest an urgent need for the CRISPR community to agree upon a benchmark dataset such as TrueOT and highlight that various sources of CRISPR data cannot be assumed to be equivalent. Availability and Implementation Our code base and datasets are available on GitHub at github.com/baolab-rice/CRISPR_OT_scoring .","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2021,"id":216096,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":5,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.5158,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":280589,"name":"Yidan Pan","orcid":"0000-0002-9702-8530","position":1,"is_corresponding":false},{"id":337683,"name":"Hoang-Anh Vu","orcid":"0000-0003-3467-9106","position":2,"is_corresponding":false},{"id":811469,"name":"Mingming Cao","orcid":"0000-0002-4234-1175","position":3,"is_corresponding":false},{"id":488301,"name":"Richard G. Baraniuk","orcid":"0000-0002-0721-8999","position":4,"is_corresponding":false},{"id":241982,"name":"Gang Bao","orcid":"0000-0001-5501-554X","position":5,"is_corresponding":false},{"id":488297,"name":"Pavan K. Kota","orcid":"0000-0001-7008-8405","position":0,"is_corresponding":true}],"reference_count":60,"raw_metadata":null,"created_at":"2026-07-18T23:53:00.192353Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}