{"doi":"10.1002/1873-3468.14067","title":"Sharing biological data: why, when, and how","abstract":"Data sharing is an essential element of the scientific method, imperative to ensure transparency and reproducibility. Researchers often reuse shared data for meta-analyses or to accompany new data. Different areas of research collect fundamentally different types of data, such as tabular data, sequence data, and image data. These types of data differ greatly in size and require different approaches for sharing. Here, we outline good practices to make your biological data publicly accessible and usable, generally and for several specific kinds of data. Sharing data proves more useful when others can easily find and access, interpret, and reuse the data. To maximize the benefit of sharing your data, follow the findable, accessible, interoperable, and reusable (FAIR) guiding principles of data sharing [[1]] (Box 1), which optimize reuse of generated data. The FAIR principles outline clear standards for ensuring that others can find and access your data and that once accessed, users can easily understand and reuse the data. The FAIR principles provide a clear collection of important details to include within your data and metadata (see ‘Data metadata and documentation’). The first step in (re)using data is to find them. Metadata and data should be easy to find for both humans and computers. Machine-readable metadata are essential for automatic discovery of datasets and services, so this is an essential component of the FAIRification process. Once the user finds the required data, she/he needs to know how can they be accessed, possibly including authentication and authorization. The data usually need to be integrated with other data. In addition, the data need to interoperate with applications or workflows for analysis, storage, and processing. The ultimate goal of FAIR is to optimize the reuse of data. To achieve this, metadata and data should be well-described so that they can be replicated and/or combined in different settings. By GO FAIR [[1]] (https://www.go-fair.org/fair-principles/), provided under the Creative Commons Attribution 4.0 International license. The repositories and practices we recommend below fulfill some of these principles and make it easier for you to follow others. This will not only help others using your data, but can also save you time in the future (see ‘The benefits of sharing data to individual researchers’). The National Institutes of Health (NIH), Canadian Institutes of Health Research (CIHR), Monarch Initiative [[2, 3]], and the Research Data Alliance (https://www.rd-alliance.org/) all recommend FAIR principles for data sharing. Amendments to these recommendations that add measures for traceability (such as evidence and provenance), licensing, and connectedness (such as identifiers and versioning) further improve data reusability [[4, 5]]. Sharing data allows for transparency in scientific studies and allows one to fully understand what occurred in an analysis and reproduce the results. Without complete data, metadata (see ‘Data and metadata’), and information about resources used to generate the data, reproducing a study proves impossible [[6, 7]]. Within the biological sciences, we have a problem of data waste—ostensibly shared data that no one ever uses. Many otherwise useful datasets go underused because researchers cannot effectively reuse the data. The inability to reuse arises from lack of discoverability, lack of important information provided, inconsistencies in data and metadata, and licensing issues. When shared effectively, we can multiply the benefits of large datasets that cost large amounts of funds and research time. Combining previously shared biological data accelerates development of analytical methods used to analyze biological data. Reusing rare samples increases the sample impact. Combining data together in meta-analyses increases study power. Data sharing also leads to fewer duplicate studies. Researchers can build on previous studies to corroborate or falsify their findings ","journal":"FEBS Letters","year":2021,"id":152735,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":73,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9499,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":301,"name":"Gregory P. Way","orcid":"0000-0002-0503-9348","position":1,"is_corresponding":false},{"id":255007,"name":"Wout Bittremieux","orcid":"0000-0002-3105-1359","position":2,"is_corresponding":false},{"id":98191,"name":"Jean‐Paul Armache","orcid":"0000-0001-9195-2282","position":3,"is_corresponding":false},{"id":4026,"name":"Melissa A Haendel","orcid":"0000-0001-9114-8737","position":4,"is_corresponding":false},{"id":15548,"name":"Michael M. Hoffman","orcid":"0000-0002-4517-1562","position":5,"is_corresponding":false},{"id":648227,"name":"Samantha L. Wilson","orcid":"0000-0003-4346-9696","position":0,"is_corresponding":true}],"reference_count":123,"raw_metadata":null,"created_at":"2026-07-18T23:43:34.972031Z","pmid":"33843054","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}