{"doi":"10.1109/cvpr52688.2022.01567","title":"Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning","abstract":"Vision Transformers (ViTs) and their multi-scale and hierarchical variations have been successful at capturing image representations but their use has been generally studied for low-resolution images (e.g. 256 × 256, 384 × 384). For gigapixel whole-slide imaging (WSI) in computational pathology, WSIs can be as large as 150000 × 150000 pixels at 20 × magnification and exhibit a hierarchical structure of visual tokens across varying resolutions: from 16 × 16 images capturing individual cells, to 4096 × 4096 images characterizing interactions within the tissue microenvironment. We introduce a new ViT architecture called the Hierarchical Image Pyramid Transformer (HIPT), which leverages the natural hierarchical structure inherent in WSIs using two levels of self-supervised learning to learn high-resolution image representations. HIPT is pretrained across 33 cancer types using 10,678 gigapixel WSIs, 408,218 4096 × 4096 images, and 104M 256 × 256 images. We benchmark HIPT representations on 9 slide-level tasks, and demonstrate that: 1) HIPT with hierarchical pretraining outperforms current state-of-the-art methods for cancer subtyping and survival prediction, 2) self-supervised ViTs are able to model important inductive biases about the hierarchical structure of phenotypes in the tumor microenvironment.","journal":"2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","year":2022,"id":296299,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":527,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9615,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":840503,"name":"Chengkuan Chen","orcid":null,"position":1,"is_corresponding":false},{"id":819294,"name":"Yicong Li","orcid":"0000-0002-5659-793X","position":2,"is_corresponding":false},{"id":615212,"name":"Tiffany Chen","orcid":"0000-0003-2546-2941","position":3,"is_corresponding":false},{"id":834557,"name":"Andrew D. Trister","orcid":"0000-0002-7070-0139","position":4,"is_corresponding":false},{"id":982997,"name":"Rahul G. Krishnan","orcid":"0000-0002-7955-3956","position":5,"is_corresponding":false},{"id":95256,"name":"Faisal Mahmood","orcid":"0000-0001-7587-1562","position":6,"is_corresponding":false},{"id":615211,"name":"Richard J. Chen","orcid":"0000-0003-0389-1331","position":0,"is_corresponding":true}],"reference_count":108,"raw_metadata":null,"created_at":"2026-07-19T00:31:16.555318Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}