{"doi":"10.3390/electronics15091904","title":"A Multimodal Audiovisual Deep Learning Framework for Early Detection of Parkinson’s Disease","abstract":"<jats:p>Parkinson’s disease (PD) is a progressive neurodegenerative disorder primarily caused by the degeneration of dopamine-producing neurons in the substantia nigra, leading to characteristic motor symptoms such as tremors, rigidity, and bradykinesia, as well as non-motor manifestations including depression, sleep disturbances, and speech impairments. Among these symptoms, speech abnormalities affect approximately 90% of individuals with PD, making acoustic analysis a promising non-invasive cue for early detection. However, subtle speech variations are often imperceptible to the human ear, and speech-only analysis may overlook complementary visual manifestations, such as hypomimia—reduced facial expressivity commonly observed in PD patients. To address these limitations, we propose Parkinson’s Detection via Attentional Fusion Network (PDAF-Net), a novel multimodal deep learning framework for early PD detection that jointly models acoustic and facial dynamic features in a binary classification setting. The proposed architecture consists of a Dual-Stream Feature Encoder (DSFE), with an audio branch based on a one-dimensional convolutional neural network (1D-CNN) and bidirectional long short-term memory (BiLSTM), and a visual branch built upon a two-dimensional convolutional neural network (2D-CNN) and a Transformer encoder. Multimodal integration is achieved through a Cross-Attention-guided Attentional Feature Fusion (CA-AFF) module, which explicitly models bidirectional cross-modal interactions and performs adaptive feature recalibration via an iterative attentional fusion mechanism. We conducted experiments on a self-collected Chinese multimodal dataset comprising 100 PD patients and 100 healthy controls. Although the data are balanced at the subject level, sliding-window segmentation introduces sample-level imbalance; to address this issue, a class-balanced focal loss is employed. Model performance was evaluated using subject-wise five-fold cross-validation. The results demonstrate that PDAF-Net consistently outperforms unimodal baselines across multiple evaluation metrics, achieving an accuracy of 89.3%, an F1-score of 0.884, and an AUC of 0.916. These findings highlight the effectiveness of explicit cross-modal interaction modeling and adaptive feature fusion for improving automated early PD screening in real-world clinical settings.</jats:p>","journal":"Electronics","year":2026,"id":640774,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1665544,"name":"Hua Huo","orcid":"0000-0001-9545-5443","position":1,"is_corresponding":false},{"id":1665545,"name":"Yulong Pei","orcid":null,"position":2,"is_corresponding":false},{"id":1665546,"name":"Lan Ma","orcid":null,"position":3,"is_corresponding":false},{"id":1665547,"name":"Shilu Kang","orcid":"0000-0002-8291-963X","position":4,"is_corresponding":false},{"id":1303392,"name":"Jiaxin Xu","orcid":"0000-0001-9830-3189","position":5,"is_corresponding":false},{"id":1665548,"name":"Aokun Mei","orcid":"0009-0006-7919-5550","position":6,"is_corresponding":false},{"id":1665543,"name":"Yinpeng Guo","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"A Multimodal Audiovisual Deep Learning Framework for Early Detection of Parkinson’s Disease","abstract":"<jats:p>Parkinson’s disease (PD) is a progressive neurodegenerative disorder primarily caused by the degeneration of dopamine-producing neurons in the substantia nigra, leading to characteristic motor symptoms such as tremors, rigidity, and bradykinesia, as well as non-motor manifestations including depression, sleep disturbances, and speech impairments. Among these symptoms, speech abnormalities affect approximately 90% of individuals with PD, making acoustic analysis a promising non-invasive cue for early detection. However, subtle speech variations are often imperceptible to the human ear, and speech-only analysis may overlook complementary visual manifestations, such as hypomimia—reduced facial expressivity commonly observed in PD patients. To address these limitations, we propose Parkinson’s Detection via Attentional Fusion Network (PDAF-Net), a novel multimodal deep learning framework for early PD detection that jointly models acoustic and facial dynamic features in a binary classification setting. The proposed architecture consists of a Dual-Stream Feature Encoder (DSFE), with an audio branch based on a one-dimensional convolutional neural network (1D-CNN) and bidirectional long short-term memory (BiLSTM), and a visual branch built upon a two-dimensional convolutional neural network (2D-CNN) and a Transformer encoder. Multimodal integration is achieved through a Cross-Attention-guided Attentional Feature Fusion (CA-AFF) module, which explicitly models bidirectional cross-modal interactions and performs adaptive feature recalibration via an iterative attentional fusion mechanism. We conducted experiments on a self-collected Chinese multimodal dataset comprising 100 PD patients and 100 healthy controls. Although the data are balanced at the subject level, sliding-window segmentation introduces sample-level imbalance; to address this issue, a class-balanced focal loss is employed. Model performance was evaluated using subject-wise five-fold cross-validation. The results demonstrate that PDAF-Net consistently outperforms unimodal baselines across multiple evaluation metrics, achieving an accuracy of 89.3%, an F1-score of 0.884, and an AUC of 0.916. These findings highlight the effectiveness of explicit cross-modal interaction modeling and adaptive feature fusion for improving automated early PD screening in real-world clinical settings.</jats:p>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19910364","pmcid":null,"openalex_id":"https://openalex.org/W4414400843","authors":[],"funders":[{"funder_name":"National Natural Science Foundation of China","grant_id":"61672210","title":null},{"funder_name":"the Major Science and Technology Program of Henan Province","grant_id":"221100210500","title":null},{"funder_name":"the Central Government Guiding Local Science and Technology Development Fund Program of Henan Province","grant_id":"Z20221343032","title":null}],"total_grants":3,"fwci":0.0,"citation_percentile":0.00265779,"influential_citations":0,"citation_trend":[],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://doi.org/10.3390/electronics15091904","host_type":"journal"},{"url":"https://doi.org/10.3390/electronics15091904","host_type":"publisher"},{"url":"https://www.mdpi.com/2079-9292/15/9/1904/pdf","host_type":"publisher"},{"url":"https://doi.org/10.2139/ssrn.5518239","host_type":"repository"}],"fields_of_study":["Voice and Speech Disorders","Parkinson's Disease Mechanisms and Treatments","Subtitles and Audiovisual Media"],"mesh_terms":[],"keywords":["Feature (linguistics)","Deep learning","Feature learning","Face (sociological concept)","Multimodality","Mechanism (biology)","Feature extraction","Eye tracking"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-07T13:48:26.899555Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}