{"doi":"10.1111/1755-0998.70008","title":"Seeing the Forest Despite the Trees in Repeat‐Rich Genomic Regions","abstract":"Technological advances are producing genome assemblies of increasing quality at steadily decreasing costs. These assemblies enable the extraction of rich biological information from previously inaccessible genomic regions (e.g., repeat-rich regions) and from diverse organisms underrepresented in genomic research. Gaining functional insights from new assemblies often requires generating additional data sets, experimental approaches and complex analysis. Novel analytical methods that substantially shorten the path to biological insights are valuable, particularly if they draw conclusions from the direct analysis of assemblies. In this issue of Molecular Ecology Resources, Elphinstone et al. (2025) present RepeatOBserver—a tool to visualise repeat organisation through direct analysis of chromosome-scale assemblies. This tool facilitates the summary and visualisation of large- and fine-scale patterns of repetitive DNA sequence structure across assemblies. Their approach borrows metrics from information theory, which have found uses in ecology (i.e., the Shannon Diversity Index), to help infer functional regions within repetitive sequences including putative centromeres. Importantly, RepeatOBserver does not require annotations, repeat libraries or functional genomic data—just a high-quality assembly. This type of tool addresses ongoing challenges in mapping the structure and functions of repeat-rich chromosomal regions, which remain the least well-understood components of genomes. The availability of chromosome-scale genome assemblies is growing rapidly, as advances in long-read sequencing technology make assembly-based approaches accessible to more taxa. These genome assemblies can reveal important insights into genome biology, biomedicine and biodiversity. Our ability to extract these insights from assemblies is built on decades of hard-won work in early genomic model organisms. For example, early work on gene structure, regulation and evolution provided a knowledge base for ab initio gene prediction from nothing more than a DNA sequence. While annotation tools for non-coding sequences like the abundant repetitive DNAs found in most eukaryotic genomes are now accessible, the methods to extract insights from these regions are less mature. Repetitive DNAs evolve rapidly: their composition, organisation and abundance varies across species (Yunis and Yasmineh 1971), making predictions based on sequence conservation difficult. Many insights require functional genomic data (e.g., ChIP-seq, methylation and ATAC-seq), which may be challenging to access in non-model systems. Despite recent progress in resolving repeats in chromosome-scale assemblies, their assembly and annotation remain non-trivial problems (Lower et al. 2018). Some genome regions with critical functions are enriched in, or entirely composed of, repeated DNA sequences. Centromeres—the essential structures that guide chromosome segregation during cell division—are typically embedded in repetitive regions and remain perhaps the most challenging genome regions to predict. Centromere prediction is especially challenging because: (1) they are generally defined by the presence of a centromere-specific histone variant (CENP-A) rather than by specific DNA sequences; and (2) they tend to occur in repeat-dense chromosomal regions enriched in satellite DNA and/or transposable elements, which can be arranged as higher order repeats (reviewed in Allshire and Karpen 2008). Centromeres vary widely across species in their size and organisation (reviewed in Hartley and O'Neill 2019). Genomic studies have only recently begun to reveal the detailed organisation of centromeres of some fungi, plants and animals. Some interesting patterns are emerging from these studies: the functional centromere core can correlate with regions of low repeat diversity and a regional dip in DNA methylation (e.g., Altemose et al. 2022). However, inferring that a repetitive DNA is part of one of these functional ","journal":"Molecular Ecology Resources","year":2025,"id":569643,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.7239,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":343006,"name":"John S. Sproul","orcid":"0000-0002-6747-3537","position":1,"is_corresponding":false},{"id":289357,"name":"Amanda M. Larracuente","orcid":"0000-0001-5944-5686","position":0,"is_corresponding":true}],"reference_count":11,"raw_metadata":null,"created_at":"2026-07-19T02:57:03.510013Z","pmid":"40608064","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}