{"doi":"10.3389/fmolb.2023.1286716","title":"The emerging field of opportunities for single-cell DNA methylation studies in hematology and beyond","abstract":null,"journal":"Frontiers in Molecular Biosciences","year":2023,"id":632051,"datarank":0.10397207708399181,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"self_citation_contribution":0.10397207708399181,"citation_network_contribution":0.0,"self_endowment_contribution":0.10397207708399181,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1630722,"name":"Agostina Bianchi","orcid":null,"position":1,"is_corresponding":false},{"id":556953,"name":"Renée Beekman","orcid":"0000-0001-7081-7874","position":2,"is_corresponding":false},{"id":1638230,"name":"Leone Albinati","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"The emerging field of opportunities for single-cell DNA methylation studies in hematology and beyond","abstract":"DNA methylation (DNAm) commonly refers to the methylation of cytosines in the mammalian genome, which mainly occurs at CpG sites (cytosines followed by guanines). Its existence was initially predicted in 1948 [1] and the first analyses of eukaryotic DNAm states were performed in 1978 [2]. Since then, and speeded up by recent genome-wide analyses, a large body of evidence suggests that proper DNAm is vital for hematopoiesis. For example, aberrations affecting the function of the DNAm machinery, e.g., DNA methyltransferases (DNMTs) and demethylating enzymes (TETs among others) (Fig. 1a), have profound effects on the hematological system, from differentiation defects to malignant transformation, as reviewed by Gore et al. and Medeiros et al. [3,4] While DNAm is usually associated with gene expression changes, a clear link between these two layers is still lacking. Less commonly considered is that DNAm can be used as read out of other layers of information, such as cell identity [5][6][7] and cell proliferation [8][9][10][11][12]. This, together with advancements in single-cell technologies, opens up new avenues to study cellular processes from the single-cell DNAm (scDNAm) perspective. In this opinion piece, we address the main processes that shape the DNA methylome during hematopoiesis and beyond, and provide insights into how to utilise scDNAm studies to extract underlying layers of biological information. Altogether, allowing to expand the opportunities for new discoveries in the single-cell era.DNAm changes in regulatory elements are widely studied since early observations that promoters of tumor suppressor genes (TSGs) are hypermethylated in cancer with p16 as hallmark example [13,14]. However, in many cases silencing of TSGs precedes promoter methylation and demethylation does not necessarily reactivate gene expression as reviewed by Jones et al. [15] In addition, promoter DNAm is not directly linked to gene expression, but majorly depends on promoter CpG density [16]. Nevertheless, for a fraction of promoters, DNAm and gene expression levels show clear correlations, for instance mediated by methylation sensitive transcription factors (TF) [17][18][19]. Interestingly, with the advancement of genome-wide methods, the epigenomics field has opened up to explore DNAm atin enhancers. Even though active enhancers harbour low DNAm levels [6,20], however,it remains unclear if DNA demethylation is necessary for enhancer activation. For most enhancers chromatin accessibility is not linked to low DNAm levels at the single-molecule level [21]. This suggests that DNA demethylation is not instructive for activation of most enhancers. Nevertheless, enhancer DNAm is a powerful tool to study lineage specification and cell differentiation, both guided by TF programs driving unique enhancer activation patterns. In fact, it is becoming increasingly clear that differential DNAm among cell lineage and differentiation stages are mainly located in enhancers, as shown by detailed DNAm map of hematopoietic cells [7,22], as well as for a broad set of other cell types [5,6]. Enhancer DNAm can furthermore be used as an indirect readout of cellular TF activity, both in healthy and tumor cells. This can be achieved by applying TF motif enrichment analyses in differentially methylated regions, as shown for hematopoeitic cells types and their derived malignancies, as well as for non-hematopoietic cells [5-8, 22, 23]. Thirdly, unmethylated regions can harbour epigenetic imprints of past enhancer activation. An example is a set of more than 100 enhancer regions that is active during B-cell differentiation at the stage of germinal center B cells, but neither active in healthy pre-and post-germinal B cells, nor in their derived neoplasms. Interestingly, while these regions have high DNAm levels in tumors arising from pre-germinal center B cells, they reduce their DNAm levels in neoplasms originating from post-germinal center B cells (determined combining data of [23][24][25][26][27]). This shows that in post-germinal center derived B-cell tumors these inactive regions harbour the memory of previous enhancer activation, revealed by their low DNAm levels. Finally, genomic regions with low DNAm levels can represent inactive sites (lacking histone marks H3K27ac and H3K4me1) that were active enhancers during earlier cell stages. Clear examples are murine embryonic enhancers with low DNAm levels in adult tissues [24,25]. Hence, low DNAm levels at these regions serve as an imprint of past enhancer activity and can be used to study cellular history.The clear segregation of cell types and cell stages based on DNAm levels at enhancers implies that DNAm profiles can be very informative to determine the cellular origin of tumors. This is of particular interest in poorly differentiated tumors with unclear cell of origin, or in cases of metastasis where the primary tumor cannot be determined. In fact, ample examples exist where DNAm of tumor DNA or circulating cell-free DNA can be used as biomarker to diagnose tumors, as reviewed or analysed in large cohort studies [26][27][28][29]. Moreover, DNAm cannot only be exploited to study tumor cell of origin at the global cell type level, but also at the more detailed level of differentiation stage. For instance, different Bcell related tumors, such as chronic lymphocytic leukemia (CLL) and mantle cell lymphoma, can be further divided into subtypes originating from pre-or post-germinal center B cells [30,31], being relevant for clinical decision making and prognostic outcome.Intriguingly, DNAm outside enhancers can also be leveraged to understand cellular history. Cell division for example is accompanied by global DNAm changes, leaving an epigenetic mark of proliferation. It is one of the main contributors that shape the DNA methylome in tumors as well as healthy cells [8][9][10][11][12]. This proliferation-associated drift is characterized by stochastic loss of DNAm in heterochromatic regions (lacking the main regulatory histone marks or containing H3K9me3). In contrast, polycomb-repressed regions (marked by H3K27me3), stochasticly accumulate DNAm upon proliferation [8]. Importantly, while this is a relatively novel concept, the underlying observations are not new; global DNAm loss in tumors and DNAm gain at TSG promoters were already observed long ago [13,14,32,33], followed by many similar findings. All these likely represent proliferation-associated drift. Of further note, the two contrasting behaviours -loss and gain of DNAm -at different genomic loci usually co-exist within a tumor, but some neoplasms show a clear bias for one of the two (e.g., hypermethylation in acute lymphoblastic leukemia and hypomethylation in multiple myelome) [8]. Thirdly, while some CpGs targeted by proliferation-associated drift overlap with CpGs affected by ageing, others do not, suggesting that these two processes are related but not necessarily affect the same genomic regions [8,12]. Though the functional impact of proliferation-associated DNAm drift is still unknown, notion of this process is crucial to investigate cancer development and progression. It can for example be used to study the impact of genetic alterations on proliferation or as powerful independent prognostic marker [8]. In contrast, it also can hamper DNAm interpretations. Particularly, the presence of partially methylated domains due to the stochastic nature of proliferation-associated drift limits the ability to distinguish functional hemimethylation events, e.g., imprinted or mono-allelic changes, unless moving to the single-cell and/or singlemolecule level.In summary, in this section we have outlined the power of DNAm data to characterise underlying biological processes shaping important underlying biological features of normal and tumor cells. All stand or falls though, with selection of the right CpGs. Namely, to study cell identity, CpGs in enhancers marked by histone marks H3K27ac and H3K4me1, will be most powerful. In contrast, for proliferative history, CpGs in H3K9me3-and H3K27me3-marked regions need to be analysed, together with those in heterochromatic regions lacking main regulatory histone marks (Fig. 1b). Of note, these regions can differ dependent on the biological system. To define them, detailed, cell-type specific histone-mark based chromatin state maps are needed. In this respect, large-scale efforts to generate reference epigenomes are invaluable [34,35], with the BLUEPRINT consortium being at the forefront of providing detailed characterisations of normal and malignant hematopoiesis [36]. They provide excellent resources enhancing the opportunities to read out underlying biological information from DNAm data.As shown in the previous paragraphs, bulk DNAm studies have contributed to discoveries in normal hematological development and its derived malignancies. However, these analyses do not enable the exploration of cellular heterogeneity, nor the detection of rare cell types such as low abundant healthy or pre-malignant cell states, respectively arising during normal differentiation and tumor formation. Neither do they allow us to study DNAm patterns in isolated small cell populations, due to the high amount of input material needed. To overcome these limitations, the DNAm field has directed major efforts towards the development of single-cell technologies. In this section, we will give an overview of the main recent advancements of scDNAm methods and highlight key aspects to take into consideration to choose the appropriate method dependent on the biological question of interest. An overview of the highlighted methods and their main features can be found in Table 1.Single-cell DNAm methods face multiple challenges, both from the biological and technical perspective. A biological bottleneck is the large number of CpGs present in every human cell, namely two copies of roughly 29M CpGs. Therefore, analyzing the entire DNA methylome in large numbers of cells is very difficult, if not impossible, due to sequencing limitations. Covering all CpGs in a single cell would require 100-1,000 times more reads/cell in comparison with a scRNA-seq experiment. Even with extreme sequencing efforts (5-7M reads/cell), genome-wide single-cell methods can only cover a random set of 6-7.5% of the CpGs/cell [37,38]. Consequently, scDNAm methods suffer from high dropout rates, generating very sparse datasets. This sparsity is furthermore augmented by other technical reasons such as degradation of the DNA upon bisulfite treatment, and PCR amplification biases and failures. Importantly, the high level of randomness of CpG coverage in genome-wide methods implies that it is unlikely to have the same CpGs covered in different cells, resulting in a low so-called overlapping coverage. In other words, for each cell one obtains information about a different set of CpGs (Fig. 1c). This leads to the loss of single-CpG resolution when comparing DNAm readouts among cells, since binarization of the data into larger regions is necessary to compare scDNAm readouts among cells. A related biological limitation is that most CpGs have static DNAm levels; within each tissue only a minor fraction of CpGs show dynamic methylation [5,23,39]. This means that only a small fraction of the DNA methylome is informative. For the above-mentioned reasons, aiming to analyze the entire DNA methylome, as in genome-wide methods, hinders the ability to obtain information about biologically relevant CpGs.Nevertheless, high-throughput genome-wide bisulfite-based scDNAm methods, such as sci-MET [38], sci-METv2 [40], snmC-seq [37], snmC-seq2 [41] and Drop-BS [42] offer exciting opportunities to perform hypothesis-free DNAm analyses, with high-throughput methods enabling readouts of thousands of cells. However, one must keep in mind that the overlapping coverage and resolution obtained with these methods largely depend on the sequencing depth one can afford. At high sequencing depth they are effective in distinguishing cell types. Hence, taking advantage of their unbiasedness, genome-wide methods have been used to build cell-type specific DNAm atlases [37], investigate hematopoiesis during embryonic development [43], and study hematopoietic stem cell populations [44]. This illustrates their power to uncover new cell types and intermediate differentiation stages. Additionally, even with low sequencing depth we consider that they can be employed to study proliferation-associated drift, as this can be calculated using DNAm levels over large genomic regions. While sequencing depth is of major importance, increasing the ed sequencing efforts cannot control for the loss of information due to degradation of DNA during bisulfite conversion. For this reason, alternative genome-wide methods based on methylation-sensitive restriction enzymes (MSRE) have been developed, such as epi-gSCAR [45] and DARE [46]. MSRE-based methods improve overlapping coverage, but at the same time they are not completely unbiased, since CpGs should be located inside an MSRE restriction site to be studied. Furthermore, they are plate-based, which limits cellular throughput to a few cells per experiment, unless one can count on the use of a robot to automate the plate-based processing.To overcome constraints related to the analysis of the entire DNA methylome, we and others have developed targeted scDNAm methods. We saw the opportunity of combining MSRE with a microfluidics system, the Tapestri platform from Mission Bio, and developed scTAM-seq [47]. Our methodology builds upon an important development to reduce the needed sequencing depth (and thus the price) by targeting outputs to biological informative CpGs (Fig. 1d). Other targeted methods have so far focused on covering specific functional regions in the genome such as regulatory regions in sciMET-cap [48], promoters in scRRBS [49][50][51], or LINEs and SINEs in scTEM-seq [52]. While these methods do reduce sequencing requirements, they still suffer from low overlapping coverage.none of them aimed to reach single-CpG resolution. Furthermore, scRRBS and scTEM-seq are plate-based which limits cellular throughput as mentioned above. NeverthelessYet, such approaches have shed important light onto DNAm dynamics in relation to clonal hematopoiesis and chronic lymphocytic leukemiaCLL [50,53]. To reach single-base pair resolution of a highly informative set of CpGs, we developed a microfluidics-based method, scTAM-seq, using a rigorous selection of approximately 400 genomic sites to maximize overlapping coverage. To maximize overlapping coverage and allow cell comparisons of a highly informative set of methylation sites at the single-CpG level, we developed the microfluidics-based method scTAM-seq, using a rigorous selection of approximately 400 genomic sites. This method allowed us not only to read out distinct B-cell stages but also to capture the more continuous process of memory B-cell formation [47]. Of notice, this method requires prior knowledge of the system to design the CpG panel. Nevertheless, Hence, we believe that ourthis method will open new doors to characterise complex, continuous biological systems with transitional cell states such as hematological differentiation [54,55], cellular reprogramming [56], as well as tumor development and plasticity [57,58].Overall, in this section, we have discussed the major aspects to consider when opting to use scDNAm methods. The most suitable method to profile the DNA methylome in single cells will depend on 1) the knowledge of the model, 2) the biological question to be answered, and 3) the budget. That being said, to obtain a global snapshot of the DNAm landscape or build cellular DNAm atlases, unbiased genomewide methods are the most appropriate [40][41][42] , followed by less-expensive targeted approaches such as scRRBS [51,59] or sciMET-cap [48]. Based on the number of input cells, either plate-based (low input) or droplet-based/combinatorial-indexing-based (high input) methods would be preferred. Finally, if one aims to study a small set of highly informative CpGs (up to 9001,000) at single-base pairCpG resolutionlevel, we believe that scTAM-seq is optimal [47]. The overlapping coverage that can ce be achieved by this method allows the study of DNAm variability at individual genomic loci as well as their combinatorial DNAm behaviors in unprecedented detail.The single-cell field has evolved extensively in recent years, providing us highly informative atlases of many tissue types through large-scale initiatives such as the Human Cell Atlas [60]. The field has largely focused on transriptomic and chromatin accessibility readouts though, with single-cell explorations of the DNA methylome lagging far behind. This implies that our view of the DNA methylome at single-cell level is still in its infancy, while highly relevant information can be extracted from this epigenetic layer. In this opinion piece, we describe the underlying biological processesfeatures that can be read out from DNAm data and we lined out the main developments of the scDNAm field in recent years. In this way, we aim to underline the biological relevance of DNAm studies in the single-cell era, an emerging field with many opportunities. Importantly, the scDNAm field is rapidly progressing and while we write this piece, new technologies are being developed with the aim to increase the targeted amount of CpGs, overlapping coverage, and resolution, while reducing sequencing efforts and therefore the costs. Furthermore, though out of the scope of this piece, techniques are developed to combine DNAm studies with other (epi)genetic layers of information [45-47, 49, 50, 52, 61-65]. Such single-cell multi-omics technologies shall aid in better understanding of the correlations between DNAm and other molecular features such as somatic mutations, copy number variations, chromatin accessibility, gene expression, cell-surface proteins, and alternative splicing. In summary, with all recent biological insights and technical advancements, the scDNAm field has exciting times ahead, heading towards deep characterisations of complex biological processes such as cell differentiation, cellular reprogramming, and tumor formation, capturing the potential paths cells can follow within these continuous landscapes.The authors declare the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.R.B. conceptualized the outline of the opinion piece, L.A., A.B., and R.B. discussed all its content, designed the figure, and wrote the manuscript and performed the revisions.We acknowledge support from the Spanish Ministry of Science and Innovation to the EMBL partnership, from the Centro de Excelencia Severo Ochoa and the CERCA Programme/Generalitat de Catalunya. A.B. is supported by an FPI fellowship from the Spanish Ministry of Science and Innovation (PRE2019-087574). L.A. is supported by a fellowship from the EMERALD International PhD Programme for Medical Doctors (101034290). R.B. was supported by a Junior Leader Fellowship from the la Caixa foundation and grants from the Spanish Ministry of Science and Innovation (RTI2018-096359-A-I00) and the European Hematology      and is supported by grants from the European    and the Spanish Ministry of Science and Innovation   enzymes such as  and  were not  in this  as they  a    in this     of CpGs that can be analysed to  underlying biological information.   histone              overview of overlapping coverage,  by the number of cells in which information of a specific CpG is   scDNAm methods to a specific set of the CpGs in the genome  overlapping coverage  in comparison to  covering the same  of CpGs in each cell    methods reduce sequencing efforts and  while increasing overlapping coverage and resolution. They can be based on targeting genomic features such as CpG  regions in scRRBS and  elements in scTEM-seq  this method the number of CpGs is an  based on the number of  elements covered per  or designed to capture information at specific genomic loci marked by their   and  The  methods are most powerful to increase overlapping coverage and can  single-CpG resolution,   to  CpGs by   the  of  developments though, this number will    sequencing depth of scRNA-seq    will  the  we believe that single-CpG resolution can be  for  CpGs per   sequencing will    this number will increase even  The  methods are most powerful to increase overlapping coverage, which we   a  when a low number of CpGs are  This  is dependent on the  and sequencing  We believe that this  is   at   CpGs, using   the  of  developments and  of sequencing  this number will    of highlighted single-cell DNA methylation methods.  overview of the methods mentioned in the  with their main   considered as high-throughput methods those that  allow to reach an  of  cells per    methods based on being targeted or genome-wide and we highlighted the number of CpGs per cell when   in the   MSRE  methylation-sensitive restriction    copy number    single","is_dataset_classified":null,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"37954981","pmcid":"PMC10637949","openalex_id":"https://openalex.org/W4388202020","authors":[],"funders":[{"funder_name":"European Research Council","grant_id":"101039265","title":"Deciphering translocation-based genome topology effects and their role in lymphoma formation"},{"funder_name":"European Hematology Association","grant_id":"RG59","title":null},{"funder_name":"European Commission","grant_id":"101034290","title":"International PhD Programme for Medical Doctors"}],"total_grants":3,"fwci":0.0925,"citation_percentile":0.41630865,"influential_citations":0,"citation_trend":[{"year":2025,"count":1}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://www.frontiersin.org/articles/10.3389/fmolb.2023.1286716/pdf?isPublishedV2=False","host_type":"journal"},{"url":"https://www.frontiersin.org/articles/10.3389/fmolb.2023.1286716/pdf?isPublishedV2=False","host_type":"GOLD"},{"url":"https://www.frontiersin.org/articles/10.3389/fmolb.2023.1286716/pdf?isPublishedV2=False","host_type":"publisher"},{"url":"https://www.frontiersin.org/articles/10.3389/fmolb.2023.1286716/full","host_type":"publisher"},{"url":"http://dx.doi.org/10.3389/fmolb.2023.1286716","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/37954981","host_type":"repository"},{"url":"http://hdl.handle.net/10230/58996","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/10637949","host_type":"repository"},{"url":"https://doaj.org/article/fd8caa9e2d994de28e6a114cf98d8915","host_type":"repository"},{"url":"http://repositori.upf.edu/bitstream/10230/58996/1/Albinati_fmb_emer.pdf","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10637949/pdf/fmolb-10-1286716.pdf","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC10637949","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC10637949?pdf=render","host_type":"Europe_PMC"},{"url":"https://doi.org/10.3389/fmolb.2023.1286716","host_type":""}],"fields_of_study":["Epigenetics and DNA Methylation","Prenatal Screening and Diagnostics","Single-cell and spatial transcriptomics","Medicine","Biology","0301 basic medicine","03 medical and health sciences"],"mesh_terms":[],"keywords":["Hematology","DNA methylation","Internal medicine","DNA","Field (mathematics)","Biology","Computational biology","Medicine","Genetics","Gene","Mathematics","Hematopoiesis","Hematological malignancies","epigenomics","Single-cell Technologies","QH301-705.5","Molecular Biosciences","Biology (General)"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-06T01:31:30.280944Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}