{"doi":"10.1101/2023.02.16.23286056","title":"Semantic Retrieval of Similar Radiological Images using Vision Transformers","abstract":"Abstract Background Identifying visually and semantically similar radiological images in a database can facilitate the creation of decision support tools, teaching files, and research cohorts. Existing content-based image retrieval tools are often limited to searching by pixel-wise difference or vector distance of model predictions. Vision transformers (ViT) use attention to simultaneously take into account radiological diagnosis and visual appearance. Purpose We aim to develop a ViT-based image retrieval framework and evaluate the algorithm on NIH Chest Radiographs (CXR) and NLST Chest CTs. Materials and Methods The model was trained on 112,120 CXR and 111,955 CT images. For CXR, a ViT binary classifier was trained on 4 ground truth labels (Cardiomegaly, Opacity, Emphysema, No Finding) and ensembled to produce multilabel classifications for each CXR. For CT, a regression model was trained to minimize L1 loss on the continuous ground truth labels of patient weight. The ViT image embedding layer was treated as a global image descriptor, using the L2 distance between descriptors as a similarity measure. To qualitatively evaluate the model, five radiologists performed a reader performance study with random query images (25 CT, 25 CXR). For each image, they chose the 5 most similar images from a set of 10 images (the 5 closest and 5 furthest images from the query in model space). Inter-radiologist and radiologist-model agreement statistics were calculated. Results The CXR model achieved nDCG@5 of 0.73 (p&lt;0.001) and Cardiomegaly mAP@5 of 0.76 (p&lt;0.001) among other results. The CT model achieved nDCG of 16.85 (p&lt;0.001). The model prediction agreed with radiologist consensus on 86% of CXR samples and 79.2% of CT samples. Inter-radiologist Fleiss Kappa of 0.51 and radiologist-consensus-to-model Cohen’s Kappa of 0.65 were observed. A t-SNE of the CT model latent space was generated to validate similar image clustering. Conclusion Our ViT architecture retrieved visually and semantically similar radiological images. Summary Statement This study evaluates the efficacy of using ViT based image embeddings for CBIR tasks for CXR and CT images, finding that it performs well on visual and semantic recognition tasks. Key Results The CXR model achieved nDCG@5 of 0.73 (p&lt;0.001) and Cardiomegaly mAP@5 of 0.76 (p&lt;0.001) among other results for CXR. The CT model achieved nDCG of 16.85 (p&lt;0.001). The model prediction agreed with radiologist consensus on 86% of CXR samples and 79.2% of CT samples. Inter-radiologist Fleiss Kappa of 0.51 and radiologist consensus to model Cohen’s Kappa of 0.65 were observed.","journal":"medRxiv","year":2023,"id":394178,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9718,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":70396,"name":"Michael Jayasuriya","orcid":"0000-0003-2366-841X","position":1,"is_corresponding":false},{"id":1168868,"name":"Adrian Serapio","orcid":null,"position":2,"is_corresponding":false},{"id":1168557,"name":"Xiao Wu","orcid":"0000-0002-2679-7447","position":3,"is_corresponding":false},{"id":1168558,"name":"Eric Davis","orcid":"0000-0001-8025-3742","position":4,"is_corresponding":false},{"id":1168559,"name":"Jamie Schroeder","orcid":"0000-0002-3710-911X","position":5,"is_corresponding":false},{"id":409192,"name":"Maya Vella","orcid":"0000-0002-4816-5563","position":6,"is_corresponding":false},{"id":325837,"name":"Jae Ho Sohn","orcid":"0000-0002-6733-7551","position":7,"is_corresponding":false},{"id":1168556,"name":"Anjali Thakrar","orcid":"0000-0002-3081-1168","position":0,"is_corresponding":true}],"reference_count":17,"raw_metadata":null,"created_at":"2026-07-19T01:19:10.334330Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}