{"doi":"10.1159/000541456","title":"The Imperative of Voice Data Collection in Clinical Trials","abstract":"Voice and speech are remarkably rich and versatile mediums that hold untapped potential for transforming healthcare. We propose that a systematic collection of voice and speech data in clinical trials can offer profound insights, enhance patient monitoring, and drive forward the decentralization of clinical research.Voice and speech, by their very nature, encapsulate a lot of health-related information, such as, but not restricted to our emotional state, cognitive function, respiratory health, or neuromuscular changes [1]. Voice is the sound wave produced by the larynx caused by the vibration of the vocal folds through air pressure from the lungs. That sound wave is altered by the shape of the throat, mouth, and nose (the resonators), similar to a musical instrument. Speech, on the other hand, involves the articulation of words and the production of speech content, linguistics, and prosody. The fluctuations and nuances in voice and speech can serve as sensitive indicators of underlying health conditions, either as a vocal biomarker (fulfilling the criteria of the BEST biomarker definition [2]) or a digital outcome measure in clinical trials [3]. Yet, despite its potential, voice and speech data remain a largely underutilized resource in clinical research, mostly due to the lack of standards and processes to integrate voice biomarkers into clinical trials.The longitudinal tracking of health through voice and speech data can profoundly modify our approach to patient monitoring. Unlike traditional biomarkers that often require invasive procedures, voice data can be collected non-invasively and regularly, providing a dynamic picture of an individual’s health over time. This frequent monitoring could be particularly valuable for remote patient monitoring allowing for timely and personalized interventions [4]. Moreover, voice data collection can significantly reduce the burden on patients. Traditional clinical trials often require frequent and inconvenient visits to clinical sites. In contrast, voice data can be collected remotely, contributing to decentralization of clinical trials and reducing the need for physical visits. This aspect is especially critical in the context of chronic diseases, where patients might find frequent clinical visits to be a substantial burden. By allowing patients to contribute data remotely, we can expand the reach of clinical trials beyond traditional clinical settings. This decentralization is particularly beneficial in expanding access to trials for individuals in remote or underserved areas, thereby enhancing the diversity and generalizability of clinical research findings.The low-cost implications of integrating voice data into clinical trials are noteworthy. Traditional biomarkers and diagnostic tools can be expensive and resource-intensive. Voice data collection, however, requires minimal equipment – often just a smartphone or a computer with a microphone. This accessibility not only reduces costs but also democratizes participation in clinical trials, making it easier for a broader population to contribute to and benefit from research.Despite the clear advantages, the integration of voice data into clinical trials is not without its challenges. One of the primary hurdles is the lack of standardized protocols for voice data collection and analysis. We have not yet fully cracked the “code” of how to systematically leverage voice data. There is ongoing work by a large multidisciplinary consortium [5] on the development and validation of standardized methods for capturing, processing, and analyzing voice data in clinical research contexts [6]. Furthermore, many voice features are not specific to a single disease or symptom. Rather, they may indicate a range of health conditions, reflecting the complex interplay of various biological, physiological, or even cognitive processes. This non-specificity, while challenging, also underscores the potential of voice as a holistic biomarker for health. By capt","journal":"Digital Biomarkers","year":2024,"id":445942,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":7,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9293,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":990917,"name":"Yaël Bensoussan","orcid":"0000-0002-1635-8627","position":1,"is_corresponding":false},{"id":932177,"name":"Guy Fagherazzi","orcid":"0000-0001-5033-5966","position":0,"is_corresponding":true}],"reference_count":4,"raw_metadata":null,"created_at":"2026-07-19T02:01:46.243344Z","pmid":"39539994","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}