{"doi":"10.1101/2025.02.19.639032","title":"Detecting Primary Progressive Aphasia (PPA) from Text: A Benchmarking Study","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>Classifying subtypes of primary progressive aphasia (PPA) from connected speech presents significant diagnostic challenges due to overlapping linguistic markers. This study benchmarks the performance of traditional machine learning models with various feature extraction techniques, transformer-based models, and large language models (LLMs) for PPA classification. Our results indicate that while transformerbased models and LLMs exceed chance-level performance in terms of balanced accuracy, traditional classifiers combined with contextual embeddings remain highly competitive. Notably, SVM using RoBERTa’s embeddings achieves the highest classification accuracy. These findings underscore the potential of machine learning in enhancing the automatic classification of PPA subtypes.</jats:p>","journal":null,"year":null,"id":647212,"datarank":0.20794415416798362,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"self_citation_contribution":0.20794415416798362,"citation_network_contribution":0.0,"self_endowment_contribution":0.20794415416798362,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1686063,"name":"Fabian Lecron","orcid":null,"position":1,"is_corresponding":false},{"id":1686065,"name":"Philippe Fortemps","orcid":null,"position":2,"is_corresponding":false},{"id":19606,"name":"Bradford C. Dickerson","orcid":"0000-0002-5958-3445","position":3,"is_corresponding":false},{"id":1686068,"name":"Mascha Kurpicz-Briki","orcid":null,"position":4,"is_corresponding":false},{"id":987260,"name":"Neguine Rezaii","orcid":"0000-0002-8967-9237","position":5,"is_corresponding":false},{"id":1686062,"name":"Ghofrane Merhbene","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Detecting Primary Progressive Aphasia (PPA) from Text: A Benchmarking Study","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>Classifying subtypes of primary progressive aphasia (PPA) from connected speech presents significant diagnostic challenges due to overlapping linguistic markers. This study benchmarks the performance of traditional machine learning models with various feature extraction techniques, transformer-based models, and large language models (LLMs) for PPA classification. Our results indicate that while transformerbased models and LLMs exceed chance-level performance in terms of balanced accuracy, traditional classifiers combined with contextual embeddings remain highly competitive. Notably, SVM using RoBERTa’s embeddings achieves the highest classification accuracy. These findings underscore the potential of machine learning in enhancing the automatic classification of PPA subtypes.</jats:p>","is_dataset_classified":null,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19910364","pmcid":null,"openalex_id":"https://openalex.org/W4407893515","authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2025,"count":1},{"year":2026,"count":2}],"oa_status":"green","license":"cc-by","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2025/02/24/2025.02.19.639032.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2025/02/24/2025.02.19.639032.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/2025.02.19.639032","host_type":"publisher"},{"url":"https://doi.org/10.1101/2025.02.19.639032","host_type":"repository"},{"url":"https://orbi.umons.ac.be/handle/20.500.12907/55857","host_type":"repository"}],"fields_of_study":["Neurobiology of Language and Bilingualism","Text Readability and Simplification"],"mesh_terms":[],"keywords":["Primary progressive aphasia","Benchmarking","Aphasia","Natural language processing","Primary (astronomy)","Computer science","Psychology","Artificial intelligence","Medicine","Cognitive psychology","Business","Physics","Pathology","Marketing"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-09T17:23:25.584081Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}