{"doi":"10.6084/m9.figshare.3487685.v5","title":"TCGA Pan-Cancer sample, expression, and mutation data for Project Cognoma","abstract":"The following datasets were created for Project Cognoma:<br>`expression-matrix.tsv.bz2` is a sample × gene matrix indicating a gene's expression level for a given sample. This dataset will be the feature/x/predictor information for Project Cognoma.<br>`mutation-matrix.tsv.bz2` is a sample × gene matrix indicating whether a gene is mutated for a given sample. Select columns (or unions of several columns) in this dataset will be the status/y/outcome for Project Cognoma.<br>`samples.tsv` is a sample × attribute matrix providing sample information and clinical measures for each sample.<br>`covariates.tsv` is a sample × attribute matrix for modeling that encodes categorical variables in samples.tsv using dummies.<br>All datasets contain the same samples as rows (in the same order). No two samples correspond to the same patient.<br>The data was retrieved from the UCSC Xena Browser. All original work in the data is released under CC0 1.0. TCGA data is public domain and the license of Xena data is currently unclear.<br>These datasets were created by the GitHub repository commit below, although the large files were not tracked due to file size. See the download directory of the cancer-data repository for metadata files with the version info for the Xena downloads this release is based on.<br>See the data/subset directory of the cancer-data repository on GitHub to browse small subsets of the expression and mutation datasets.","journal":"Figshare","year":2016,"id":2050,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":0.0,"corpus_rank":10062,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.95,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2016-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":301,"name":"Gregory P. Way","orcid":"0000-0002-0503-9348","position":1,"is_corresponding":false},{"id":23977,"name":"C McLeod","orcid":null,"position":2,"is_corresponding":false},{"id":23978,"name":"Stephen D. Shank","orcid":"0000-0003-0734-9953","position":3,"is_corresponding":false},{"id":308,"name":"Casey S. Greene","orcid":"0000-0001-8713-9213","position":4,"is_corresponding":false},{"id":1208,"name":"Daniel S. Himmelstein","orcid":"0000-0002-3012-7446","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-03-01T18:20:47.508186Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":"green","license":"public-domain","views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}