{"doi":"10.1101/2025.01.06.631595","title":"DNALongBench: A Benchmark Suite for Long-Range DNA Prediction Tasks","abstract":"Modeling long-range DNA dependencies is crucial for understanding genome structure and function across a wide range of biological contexts. However, effectively capturing these extensive dependencies, which may span millions of base pairs in tasks such as three-dimensional (3D) chromatin folding prediction, remains a significant challenge. Furthermore, a comprehensive benchmark suite for evaluating tasks that rely on long-range dependencies is notably absent. To address this gap, we introduce DNALongBench, a benchmark dataset encompassing five important genomics tasks that consider long-range dependencies up to 1 million base pairs: enhancer-target gene interaction, expression quantitative trait loci, 3D genome organization, regulatory sequence activity, and transcription initiation signals. To comprehensively assess DNALongBench, we evaluate the performance of five methods: a task-specific expert model, a convolutional neural network (CNN)-based model, and three fine-tuned DNA foundation models - HyenaDNA, Caduceus-Ph, and Caduceus-PS. We envision DNALongBench as a standardized resource with the potential to facilitate comprehensive comparisons and rigorous evaluations of emerging DNA sequence-based deep learning models that account for long-range dependencies.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2025,"id":554838,"datarank":0.29188652235829704,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"self_citation_contribution":0.29188652235829704,"citation_network_contribution":0.0,"self_endowment_contribution":0.29188652235829704,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":6,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.7461,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1408134,"name":"Zhenqiao Song","orcid":"0009-0006-6125-900X","position":1,"is_corresponding":false},{"id":1452722,"name":"Yang Zhang","orcid":"0000-0002-0432-5772","position":2,"is_corresponding":false},{"id":516608,"name":"Shike Wang","orcid":"0000-0003-3007-533X","position":3,"is_corresponding":false},{"id":1408135,"name":"Danqing Wang","orcid":"0009-0002-4611-628X","position":4,"is_corresponding":false},{"id":51369,"name":"Muyu Yang","orcid":"0000-0002-0142-0328","position":5,"is_corresponding":false},{"id":1452723,"name":"Lei Li","orcid":"0000-0001-7927-233X","position":6,"is_corresponding":false},{"id":51348,"name":"Jian Ma","orcid":"0000-0002-4202-5834","position":7,"is_corresponding":false},{"id":1452721,"name":"Wenduo Cheng","orcid":"0009-0001-8819-2451","position":0,"is_corresponding":true}],"reference_count":43,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T02:54:54.542303Z","pmid":"39829833","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}