Recognizing anxiety and depression in cancer patients based on speech and facial expressions

PurposeTo address the anxiety and depression experienced by cancer patients due to the stress of diagnosis and treatment, as well as the limitations of traditional assessment methods characterized by high subjectivity and low efficiency, this study aims to develop a multimodal fusion approach for the simultaneous and precise evaluation of these two psychological states.Patients and methodsA speech-video dataset of clinically diagnosed cancer patients was used. This study proposes a multimodal fusion approach: for depression recognition, We employ the HuBERT pre-training architecture based on Transformers, integrating specific acoustic features of depression with textual content to achieve accurate classification of depression through a voice-text modality. For anxiety recognition, a multi-task convolutional neural network is designed to infer anxiety status from the facial expressions.ResultsExperiments conducted on a speech-video dataset of clinically diagnosed cancer patients demonstrates that the multimodal fusion model achieves a depression recognition F1 value of 0.85 and an anxiety recognition F1 value of 0.74, significantly outperforming the unimodal model.ConclusionThe results of the two modalities are fused by decision-level weighted averaging to realize the simultaneous assessment of anxiety and depression in cancer patients. The study may provide technical support for rapid, noninvasive screening of psychological status in cancer patients.