{"doi":"10.1007/s11103-025-01604-7","title":"Genomic language models with k-mer tokenization strategies for plant genome annotation and regulatory element strength prediction","abstract":"<jats:title>Abstract</jats:title>\n          <jats:p>Recent advances in genomic language models have improved the accuracy of <jats:italic>in silico</jats:italic> analyses, yet many rely on resource-intensive architectures. In this study, we focus on the impact of <jats:italic>k</jats:italic>-mer tokenization strategies–specifically varying window sizes (three to eight) and overlap schemes–on the performance of transformer-based genomic language models. Through extensive evaluation across multiple plant genomic tasks, including splice site and alternative polyadenylation site prediction, we show that thoughtful design of the <jats:italic>k</jats:italic>-mer tokenizer plays a critical role in model performance, often outweighing model scale. In particular, overlap-based tokenization generally enhances performance by preserving local sequence context, while certain non-overlap configurations achieve competitive accuracy with improved computational efficiency in some tasks. Despite using a smaller model, our approach performs on par with the state-of-the-art AgroNT model in many cases. These results emphasize that <jats:italic>k</jats:italic>-mer tokenization, not merely model size, is a key determinant of success in genomic sequence modeling. Our findings provide practical guidance for designing efficient genomic language models tailored to plant biology.</jats:p>","journal":"Plant Molecular Biology","year":2025,"id":602608,"datarank":0.26876392038420827,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.0,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":5,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1482453,"name":"Kazumasa Horie","orcid":"0000-0002-4809-2552","position":1,"is_corresponding":false},{"id":1545537,"name":"Toshiyuki Amagasa","orcid":null,"position":2,"is_corresponding":false},{"id":1545539,"name":"Naoya Fukuda","orcid":null,"position":3,"is_corresponding":false},{"id":1545533,"name":"Shosuke Suzuki","orcid":"0009-0001-3354-195X","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Genomic language models with k-mer tokenization strategies for plant genome annotation and regulatory element strength prediction","abstract":"<jats:title>Abstract</jats:title>\n          <jats:p>Recent advances in genomic language models have improved the accuracy of <jats:italic>in silico</jats:italic> analyses, yet many rely on resource-intensive architectures. In this study, we focus on the impact of <jats:italic>k</jats:italic>-mer tokenization strategies–specifically varying window sizes (three to eight) and overlap schemes–on the performance of transformer-based genomic language models. Through extensive evaluation across multiple plant genomic tasks, including splice site and alternative polyadenylation site prediction, we show that thoughtful design of the <jats:italic>k</jats:italic>-mer tokenizer plays a critical role in model performance, often outweighing model scale. In particular, overlap-based tokenization generally enhances performance by preserving local sequence context, while certain non-overlap configurations achieve competitive accuracy with improved computational efficiency in some tasks. Despite using a smaller model, our approach performs on par with the state-of-the-art AgroNT model in many cases. These results emphasize that <jats:italic>k</jats:italic>-mer tokenization, not merely model size, is a key determinant of success in genomic sequence modeling. Our findings provide practical guidance for designing efficient genomic language models tailored to plant biology.</jats:p>","is_dataset_classified":null,"base_score":1.791759469228055,"endowment":1.791759469228055,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"40745106","pmcid":"PMC12313756","openalex_id":"https://openalex.org/W4412799752","authors":[],"funders":[],"total_grants":0,"fwci":2.0642,"citation_percentile":0.87120738,"influential_citations":0,"citation_trend":[{"year":2025,"count":3},{"year":2026,"count":2}],"oa_status":"hybrid","license":"cc-by","oa_locations":[{"url":"https://link.springer.com/content/pdf/10.1007/s11103-025-01604-7.pdf","host_type":"journal"},{"url":"https://link.springer.com/content/pdf/10.1007/s11103-025-01604-7.pdf","host_type":"publisher"},{"url":"https://link.springer.com/article/10.1007/s11103-025-01604-7/fulltext.html","host_type":"publisher"},{"url":"https://doi.org/10.1007/s11103-025-01604-7","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/40745106","host_type":"repository"},{"url":"https://tsukuba.repo.nii.ac.jp/records/2016574","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/12313756","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12313756/pdf/11103_2025_Article_1604.pdf","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC12313756","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC12313756?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Genomics and Phylogenetic Studies","Genomics and Chromatin Dynamics","RNA and protein synthesis mechanisms","Genome, Plant","Molecular Sequence Annotation","Genomics","Models, Genetic","Computational Biology","Regulatory Sequences, Nucleic Acid","Computer Simulation"],"mesh_terms":["Computer Simulation","Models, Genetic","Regulatory Sequences, Nucleic Acid","Genome, Plant","Computational Biology","Genomics","Molecular Sequence Annotation"],"keywords":["Lexical analysis","Biology","Computer science","Computational biology","Language model","In silico","Annotation","Genome","Artificial intelligence","Genetics","Gene","DNA","Genome annotation","k-mer","Transformer","Genomic Language Model"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-29T19:49:29.909346Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}