{"doi":"10.3390/rel17020192","title":"Identifying Phrase Boundaries in the Samaritan Pentateuch with Machine Learning","abstract":"<jats:p>In this paper, we present our approach to adding the first layers of syntactic information to an open dataset containing the text of the Samaritan Pentateuch with linguistic annotations. The Samaritan Pentateuch, written in Samaritan Hebrew, is the sacred text of the Samaritan community. The updated dataset now includes boundaries for phrase atoms and phrases. These were added using a parser previously employed to annotate Hebrew and Syriac texts morphologically. While the model predicts the locations of phrase atom and phrase boundaries, its output is not always accurate, so manual corrections were applied. This project highlights the potential of machine learning to enrich ancient texts. The syntactic features developed here can support linguistic research in the humanities and facilitate textual comparisons across ancient corpora—extending beyond morphological analysis to deepen our understanding of textual transmission. Additionally, the corrected data can be incorporated into the existing training set to improve future models.</jats:p>","journal":"Religions","year":2026,"id":12716,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":0.0,"corpus_rank":10062,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.8233,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2026-02-04","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":98963,"name":"Martijn Naaijer","orcid":"0009-0006-3325-0614","position":1,"is_corresponding":false},{"id":98964,"name":"Christian Canu Højgaard","orcid":"0000-0002-1855-1017","position":2,"is_corresponding":false},{"id":98965,"name":"Oliver Glanz","orcid":"0000-0001-7968-6719","position":3,"is_corresponding":false},{"id":98966,"name":"Saulo de Oliveira Cantanhede","orcid":null,"position":4,"is_corresponding":false},{"id":98962,"name":"Saulo de Oliveira Cantanhêde","orcid":"0000-0002-9695-7119","position":0,"is_corresponding":true}],"reference_count":17,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-03-01T18:20:47.508186Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":"gold","license":"cc-by","views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}