{"doi":"10.1121/10.0042820","title":"Automation of real-time vocal tract image segmentation with SAM 2.0 and morphological operation implementation","abstract":"<jats:p>Modeling articulatory representations is critical to the scientific study of speech production, including its relation to speech acoustics. However, discretizing articulatory dynamics in continuous speech has proven computationally taxing. For example, segmentation analyses of real-time vocal tract images deploying contour-tracking methods, while successful, require manual creation of templates and human supervised assessment [e.g., Bresch and Narayanan (2009). IEEE Trans. Med. Imaging. 28(3), 323–338]. In this paper, we utilize Segment Anything Model 2 (SAM 2.0) [Ravi et al. (2024). arXiv:2408.00714] to efficiently segment critical articulators in real-time magnetic resonance imaging speech production data without fine-tuning and with global nonlinear image filtering to examine such systems' ability to segment speech dynamics, which have both language- and subject-specific characteristics.</jats:p>","journal":"JASA Express Letters","year":2026,"id":647254,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1686209,"name":"Kyle Ng","orcid":null,"position":1,"is_corresponding":false},{"id":1686210,"name":"Ameen Qureshi","orcid":null,"position":2,"is_corresponding":false},{"id":589238,"name":"Dani Byrd","orcid":"0000-0003-3319-5871","position":3,"is_corresponding":false},{"id":1676317,"name":"Khalil Iskarous","orcid":"0000-0003-3269-7343","position":4,"is_corresponding":false},{"id":1686208,"name":"Haley Hsu","orcid":"0009-0001-4357-0092","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Automation of real-time vocal tract image segmentation with SAM 2.0 and morphological operation implementation","abstract":"<jats:p>Modeling articulatory representations is critical to the scientific study of speech production, including its relation to speech acoustics. However, discretizing articulatory dynamics in continuous speech has proven computationally taxing. For example, segmentation analyses of real-time vocal tract images deploying contour-tracking methods, while successful, require manual creation of templates and human supervised assessment [e.g., Bresch and Narayanan (2009). IEEE Trans. Med. Imaging. 28(3), 323–338]. In this paper, we utilize Segment Anything Model 2 (SAM 2.0) [Ravi et al. (2024). arXiv:2408.00714] to efficiently segment critical articulators in real-time magnetic resonance imaging speech production data without fine-tuning and with global nonlinear image filtering to examine such systems' ability to segment speech dynamics, which have both language- and subject-specific characteristics.</jats:p>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"41778895","pmcid":null,"openalex_id":"https://openalex.org/W7133567347","authors":[],"funders":[{"funder_name":"National Science Foundation","grant_id":"2240349","title":"CompCog: Deep causal inference grounds the perception of cognitive objects in speech"}],"total_grants":1,"fwci":0.0,"citation_percentile":0.20382079,"influential_citations":0,"citation_trend":[],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://doi.org/10.1121/10.0042820","host_type":"journal"},{"url":"https://doi.org/10.1121/10.0042820","host_type":"publisher"},{"url":"https://pubs.aip.org/asa/jel/article-pdf/doi/10.1121/10.0042820/20928290/035202_1_10.0042820.pdf","host_type":"publisher"},{"url":"https://pubmed.ncbi.nlm.nih.gov/41778895","host_type":"repository"},{"url":"https://doaj.org/article/df226852df4845318372c0df4074b197","host_type":"repository"}],"fields_of_study":["Voice and Speech Disorders","Phonetics and Phonology Research","Speech Recognition and Synthesis","0202 electrical engineering, electronic engineering, information engineering","02 engineering and technology","Humans","Magnetic Resonance Imaging","Image Processing, Computer-Assisted","Speech","Vocal Cords","Speech Acoustics","Speech Production Measurement"],"mesh_terms":["Humans","Image Processing, Computer-Assisted","Magnetic Resonance Imaging","Speech","Speech Acoustics","Speech Production Measurement","Vocal Cords"],"keywords":["Vocal tract","Segmentation","Speech production","Automation","Speech synthesis","Image (mathematics)","Image segmentation","Relation (database)","Speech Production Measurement","Image Processing, Computer-Assisted","Humans","Speech","Vocal Cords","Magnetic Resonance Imaging","Speech Acoustics"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-09T17:38:11.190563Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}