{"doi":"10.18174/546089","title":"Paving the way for FAIR data in plant phenotyping","abstract":"The increasing nutritional demands of the world as well as the need for crops that perform reliably, in spite of diverse environmental conditions (abiotic and biotic stresses and variable weather conditions), put the plant sciences at the forefront of domains where progress is urgently needed. To be able to do so, plant phenotyping and genotyping are extremely important. Especially in plant phenotyping, research is met with challenges related to poor data management, and thereby inefficient exploitation - let alone reuse of datasets. The challenges to phenotypic data reuse and integration arise due to the highly distributed nature of data in the domain (as there are no central plant phenotypic data repositories) and their multifaceted heterogeneity. The variety of experimental goals and the sheer number of species studied may necessitate different approaches (e.g. for crops, model organisms, forest trees). Experiments may be conducted in open fields, greenhouses or other locations, follow different designs and produce different types of data (e.g. visual observation of a score, images, manual and automatic measurements, molecular assays). Even when everything else matches, the data files produced may have different formats and structures, which is a challenge for data integration. Moreover, good data documentation practices are often lacking, which hinders interpretation and reuse. In the vast majority of cases, plant phenotyping datasets are used only once, solely to address the research question for which they were originally generated. It is the exception, rather than the rule, when different datasets, produced by different, uncoordinated parties, are analyzed to generate further knowledge. Even rarer, though much more useful, are cases where independently created datasets are integrated for the purpose of meta-analyses or improvement of statistical and predictive models. Such work is crucial, for example, for multi-environment studies investigating the adaptability of crops to different conditions. This relative rarity of meta-analyses and integrative studies indicates that researchers conduct experiments and collect data anew for every new study they wish to undertake, which is in many cases a suboptimal use of resources. This may not be a serious issue on a low level (i.e., single experiments) but on a higher level where multiple independent experiments may be reused and integrated in e.g. multi-environment trials, this has a greater impact.The challenges mentioned in the previous paragraph for plant phenotyping are not specific, but generic for the data life cycle in research. To address this challenge, the FAIR (Findable, Accessible, Interoperable, Reusable) data principles have been proposed as guidelines to alleviate generic reusability bottlenecks. However, FAIR data principles require domain-specific solutions. With them, datasets become more easily discoverable, interpretable, integratable and reusable. Furthermore, the principles emphasize that there should be an equal focus on human and machine readability, so that automated techniques can facilitate every step of the process. It is up to each community to devise ways to implement the FAIR principles. In this thesis, we investigated the application of the FAIR data principles in the domain of plant phenotyping.Our initial research question focused on a core requirement of FAIR, domain-relevant community standards. We identified and tackled shortcomings of the MIAPPE (Minimum Information About a Plant Phenotyping Experiment) metadata standard, which was initially presented as a flat checklist. We produced a new, refined version, MIAPPE 1.1, which can cover experiments involving a broader range of plant species (including forest trees), boasts improved documentation, and can now support FAIR data through its explicit data model and ontology (Chapter 2). We tested the new version of the standard by using it to describe a wide range of different plant phenotyping ex","journal":"Data Archiving and Networked Services (DANS)","year":2021,"id":219965,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.8901,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":95360,"name":"Evangelia A. Papoutsoglou","orcid":"0000-0001-8209-1900","position":0,"is_corresponding":true}],"reference_count":248,"raw_metadata":null,"created_at":"2026-07-18T23:53:42.928212Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}