{"doi":"10.1101/2022.11.19.517216","title":"Long-read genomes reveal pangenomic variation underlying yeast phenotypic diversity","abstract":"Abstract Understanding the genetic causes of trait variation is a primary goal of genetic research. One way that individuals can vary genetically is through the existence of variable pangenomic genes – genes that are only present in some individuals in a population. The presence or absence of entire genes could have large effects on trait variation. However, variable pangenomic genes can be missed in standard genotyping workflows, due to reliance on aligning short-read sequencing to reference genomes. A popular method for studying the genetic basis of trait variation is linkage mapping, which identifies quantitative trait loci (QTLs), regions of the genome that harbor causative genetic variants. Large-scale linkage mapping in the budding yeast Saccharomyces cerevisiae has found thousands of QTLs affecting myriad yeast phenotypes. To enable the resolution of QTLs caused by variable pangenomic genes, we used long-read sequencing to generate highly complete de novo assemblies of 16 diverse yeast isolates. With these assemblies we resolved growth QTLs to specific genes that are absent from the reference genome but present in the broader yeast population at appreciable frequency. Copies of genes also recombine onto chromosomes where they are absent in the reference genome, and we found that these copies generate additional QTLs whose resolution requires pangenome characterization. Our findings demonstrate the power of long-read sequencing to identify the genetic basis of trait variation.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2022,"id":304896,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9429,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":998321,"name":"Ilya Andreev","orcid":"0000-0001-6414-4570","position":1,"is_corresponding":false},{"id":340059,"name":"Michael Chambers","orcid":"0000-0002-0658-6984","position":2,"is_corresponding":false},{"id":20064,"name":"Morgan Park","orcid":null,"position":3,"is_corresponding":false},{"id":406002,"name":"NISC Comparative Sequencing Program","orcid":null,"position":4,"is_corresponding":false},{"id":291464,"name":"Joshua S. Bloom","orcid":"0000-0002-7241-1648","position":5,"is_corresponding":false},{"id":678046,"name":"Meru J. Sadhu","orcid":"0000-0002-8636-9102","position":6,"is_corresponding":false},{"id":303836,"name":"Cory A. Weller","orcid":"0000-0001-6965-5599","position":0,"is_corresponding":true}],"reference_count":55,"raw_metadata":null,"created_at":"2026-07-19T00:32:37.185846Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}