{"doi":"10.1002/ctd2.194","title":"One toolkit to bring them all, and <i>in silico</i> analyze them","abstract":"Circulating tumor DNA (ctDNA) is a class of short tumor-derived DNA molecules that are typically detected in body fluids including plasma. Recent evidence suggests that certain molecular characteristics of ctDNA associated with various aspects of cancer transcriptome and ctDNA abundance can inform disease burden.1 These findings have shown the promise of utilizing liquid biopsy for tracing disease progression and guiding clinical decisions. Despite the increasing advances in ctDNA research and the diverse computational workflows developed to support such research, computational toolsets that tackle multiple critical questions at once are lacking yet are highly needed by the science community. To address this need, Li et al. recently developed a comprehensive toolkit for assessing cancer transcriptomic dynamics using the whole-genome sequencing (WGS) data of cell free DNA (cfDNA), which contains DNA of normal cells and ctDNA.2 This tool, named Integrated analysis toolkit for whole-genome-wide features of cfDNA (INAC), carries out multiple functions including generating whole-genome copy number profiles, estimating gene expression using fragment-based analyses, and identifying disease predictive features using machine learning (see Figure 1). Human cancers often undergo massive chromosomal rearrangements, resulting in aneuploidy that is detected through somatic copy number profiling. Depending on how much ctDNA species are captured (also known as the tumor or ctDNA fraction), the somatic copy number profile of a cfDNA sample can reflect tumor aneuploidy to a certain degree. In cases where multiple diseases co-occur in an individual, a mixture of the copy number alterations of both diseases is likely observed when analyzing cfDNA. Therefore, the copy number profile of a cfDNA sample may inform tumor burden over the time and disease classification, which is evaluated by the INAC_CNV module (see Table 1). In addition, multiple analyses of ctDNA fragmentation patterns that are provided by the toolkit are essential to transcriptomic modeling. The ctDNA species identified in plasma originate from fragmented DNA that remains bound by histones but is released from tumor cells in its nucleosome-bound form.9 Therefore, the distribution of ctDNA reads mapped to a genomic locus can be used to infer chromatin accessibility. In this context, open chromatin is associated with a higher percentage of short ctDNA fragments (or a high short/long fragment ratio) and a lower read abundance at promoters may indicate active gene expression, and vice versa10; see Figure 1). To enable the assessment of expression rate, INAC calculates the long-to-short fragment ratio and nucleosome-depleted region sequence depths at various regions flanking transcriptional start sites (TSSs) using the INAC_TSS module. Additionally, a recently published article described promoter fragmentation entropy (PFE) as another measure of gene expression state.5 The underlying basis is that highly expressed genes are associated with an open chromatin state and thus more diverse histone binding pattern, which is recognized as a greater variety of ctDNA fragment sizes with a high PFE (see Figure 1). The algorithm that assesses this fragmentation feature is also included in the INAC for evaluating expression states as INAC_PFE. Finally, one potential clinical application of ctDNA analysis is using its various molecular characteristics for predicting disease states and clinical outcome. Li et al. showed that ctDNA fragmentation patterns at TSSs could predict cancer and normal tissues better than other tested features using different machine learning (ML) methods. It is possible that additional fragmentation features that can be captured by the toolkit and have not yet been explored may be better associated with cancers, or a specific subtype of a cancer using the provided ML algorithm. Clearly, the accessibility to the toolset reveals new possibilities of identifying potential predict","journal":"Clinical and Translational Discovery","year":2023,"id":378556,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9506,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1142805,"name":"Anna Baj","orcid":"0009-0004-3636-8912","position":1,"is_corresponding":false},{"id":465481,"name":"Adam G. Sowalsky","orcid":"0000-0003-2760-1853","position":2,"is_corresponding":false},{"id":859272,"name":"Chennan Li","orcid":"0000-0002-4520-4694","position":0,"is_corresponding":true}],"reference_count":11,"raw_metadata":null,"created_at":"2026-07-19T01:16:48.594070Z","pmid":"37220531","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}