{"doi":"10.1101/2021.07.19.452952","title":"Customization of a <i>dada2</i> -based pipeline for fungal Internal Transcribed Spacer 1 (ITS 1) amplicon datasets","abstract":"Abstract Identification and analysis of fungal communities commonly rely on internal transcribed spacer (ITS)-based amplicon sequencing. Currently, there is no gold standard to infer and classify fungal constituents, in part since methodologies have been adapted from analyses of bacterial communities. To achieve high resolution inference of fungi in clinical samples, we customized a DADA2-based pipeline using a mock community of eleven medically relevant fungi. While DADA2 allowed the discrimination of ITS1 sequences differing by a single nucleotide, quality filtering, sequencing bias, and database selection were identified as key variables determining the accuracy of sample inference. By fine-tuning quality filtering, we decreased the number of wrongly discarded sequences attributed to Aspergillus species, Saccharomyces cerevisiae , and Candida glabrata reads. We confirmed this effect in patient samples. By adapting a wobble nucleotide in the ITS1 forward primer region, we further increased the yield of S. saccharomyces and C. glabrata sequences. Finally, we showed that a BLAST-based algorithm based on the UNITE+INSD or the NCBI NT database achieved a higher reliability in species-level taxonomic annotation than the naïve Bayesian classifier implemented in DADA2. These steps optimized a robust fungal ITS1 sequencing pipeline that, in most instances, enables species level-assignment of community members.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2021,"id":219082,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9486,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":288396,"name":"Bing Zhai","orcid":"0000-0001-6571-7465","position":1,"is_corresponding":false},{"id":817615,"name":"John V. Frame","orcid":null,"position":2,"is_corresponding":false},{"id":228380,"name":"Tobias M. Hohl","orcid":"0000-0002-9097-5412","position":3,"is_corresponding":false},{"id":108835,"name":"Ying Taur","orcid":"0000-0002-6601-8284","position":4,"is_corresponding":false},{"id":461113,"name":"Thierry Rolling","orcid":"0000-0002-0277-5067","position":0,"is_corresponding":true}],"reference_count":30,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-18T23:53:33.915158Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}