Multimodal Stress Detection Using Sensor, Voice, and Facial Image Data for Real-Time Health Monitoring
Main Article Content
Abstract
Stress is a large-scale issue in today's world, with regard to its physical health, mental health, and cognitive effects. It also does not go easy trying to identify what we put out for stress, which can very well be person-based and what we see in different sets of data. Presently, most of what is done is a single data type approach, which is not very effective in reality. This paper reports on a multimodal stress detection platform that we have put together from physiological sensor data, speech inputs, and face images for real-time healthcare. We process heart rate variability in the physiological data, Mel-Frequency Cepstral Coefficients from speech, and face images through specialized-for each of the types of data. Also, we use independent machine learning models. In particular, a Random Forest model is used for physiological data, a Multi-Layer Perceptron for speech characteristics, and a ResNet-50 Convolutional Neural Network for facial images. On the other hand, based on our research, the speech base model performs well at a 97% accuracy level. Physiological and images perform at 88.28% and 72.07% accuracy levels, respectively. Similarly, in our research, we use decision level fusion, which is ensuring the stability of our performance, and if either of the sources of data is inadequate, we will still obtain accurate outcomes. Its aptness for use in real-time health care, in our view, constitutes a big advantage.