{"doi":"10.5281/zenodo.3742218","title":"Curation and ISA representation of a SARS-Cov2/Covid-19 Proteomics Dataset - PXD107710 - ISA representation","abstract":"Curation and ISA representation of a SARS-Cov2/Covid-19 Proteomics Dataset deposited in PRIDE database with accession number: PXD107710 ISA-Tab annotation for the \"SARS-CoV-2 infected host cell proteomics reveal potential therapy targets\" publication. Github repository: https://github.com/ISA-tools/PXD017710 This is part of an effort to (re-)annotate: https://dx.doi.org/10.21203/rs.3.rs-17218/v1 Additional work done as part of: https://github.com/virtual-biohackathons/covid-19-bh20 https://github.com/virtual-biohackathons/covid-19-bh20/wiki/FairData <strong>Proteomics data</strong> Available from PRIDE at https://www.ebi.ac.uk/pride/archive/projects/PXD017710<br> and [MassIVE/CCMS Maestro+MSstats reanalysis of MSV000085096 / PXD017710] <strong>ISA-Tab representation:</strong> Rationale: Demonstrate suitability of the ISA format for representing MS based protein profiling experiment with more granularity and details, thus providing a better representation of the experiment design.<br> The formatting and re-annotation are based on information extracted from:<br> - the original publication<br> - the supplementary tables available from the publishers site<br> - the 'filtered-results.csv' helper file as supplied to @sneumann during the HUPO-PSI meeting March 2020 <br> Viewing the ISA-tab formatted and re-annotated PXD017710 with ISATab-Viewer Viewing the ISA-tab formatted and re-annotated PXD017710 locally, do the following: ```bash<br> python -m http.server 8000<br> ``` Then point your browser to `http://0.0.0.0:8000/isaviewer-demo.html` <strong>Curation tasks performed:</strong> * initial structure of the study design in ISA format: * linkage of Proteome and Translatome data (supplementary material) to ISA assay tables (via Derived Data File) * processing the Proteome and Translatome data (supplementary material) with python pandas library to generate the following csv files: - proteome_intensities_long_table_ggplot2.txt<br> - proteome_diffanal_ratio_pvalue_long_table_ggplot2.txt<br> - translatome_intensities_long_table_ggplot2.txt <br> - translatome_diffanal_ratio_pvalue_long_table_ggplot2<br> <br> The files are `long table` corresponding to a `melt` on the Excel file originally generated by the users and can be readily loaded in R ggplot2 library for graphical representation.<br> The statistical relevant elements have been annotated with the STATO ontology and the tables comply with a Frictionless.io Data Package.<br> The jupyter notebook for the transformation is available. * conversion of raw data to mzML format: detailed in https://github.com/ISA-tools/PXD017710 install docker: <br> ```bash<br> &gt;brew update<br> &gt;brew install docker<br> ``` sign in to docker<br> ```bash<br> &gt;docker start<br> &gt;docker login<br> ``` pull docker container for ProteoWizard:<br> ```bash<br> &gt;docker pull chambm/pwiz-i-agree-to-the-vendor-licenses<br> ``` :warning: be sure to sign-up and login to https://hub.docker.com/ in order to be able to reach https://hub.docker.com/r/chambm/pwiz-skyline-i-agree-to-the-vendor-licenses <br> run the pwiz tool from the container over the raw data:<br> ```bash<br> docker run -it --rm -e WINEDEBUG=-all -v /Users/Downloads/PXD017710/raw/:/data chambm/pwiz-skyline-i-agree-to-the-vendor-licenses wine msconvert /data/*.raw --mzML<br> ``` <br> * ontology markup for:<br> * declaration of independent variables as ISA Study Factors:{biological agent, dose, time point, replicate} -&gt;OBI<br> * Taxonomic information (host cells and virus) -&gt; NCBITaxonomy<br> * Cell line: CaCo-2 cells -&gt; Cell Line Ontology<br> * Disease: Colon Cancer -&gt; Human Phenotype Ontology<br> * MS specific aspect (TMT reagent, instrument ... ) -&gt; PSI-MS<br> * Statistical Tests -&gt; STATO <br> <strong>Unresolved curatorial issues:</strong> 1. ambiguities related to Tandem Mass Tag labelling protocol<br> - the publication mentions TMT11 (see Figure 2 in https://www.researchsquare.com/article/rs-17218/v1)<br> - the information available from PRIDE mentions TMT6 (https://www.ebi.ac.uk/pride/archive/projects/PXD017710)<br> This may require another round of annotation on the TMT agents and fractions in the ISA a_assay representation <br> 2. SARS-Cov2 isolate: no clear NCBI Taxonomic anchoring and unclear origin: -&gt; the markup is made to the parent class (as of 06.04.2020) <strong>Release and packaging as a BDBAG:</strong> The tgz file associated with this upload has been producing using https://github.com/fair-research/bdbag. It contains several manifest files detailing metadata and data files, providing md5 and sha256 checksums. <strong>Github repository:</strong> https://github.com/ISA-tools/PXD017710","journal":"Zenodo (CERN European Organization for Nuclear Research)","year":2020,"id":8096,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":0.0,"corpus_rank":10062,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9378,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":1,"downloads":53,"has_version_chain":false,"published_date":"2020-04-06","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":290,"name":"Susanna‐Assunta Sansone","orcid":"0000-0001-5306-5690","position":1,"is_corresponding":false},{"id":289,"name":"Rocca-Serra, Philippe","orcid":"0000-0001-9853-5668","position":0,"is_corresponding":false}],"reference_count":1,"raw_metadata":null,"created_at":"2026-03-01T18:20:47.508186Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":"green","license":"cc-by","views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}