Multimodal Stress Detection Using Sensor, Voice, and Facial Image Data for Real-Time Health Monitoring

Main Article Content

Purva Soni
Dr. Dinesh Jain

Abstract

Stress is a large-scale issue in today's world, with regard to its physical health, mental health, and cognitive effects. It also does not go easy trying to identify what we put out for stress, which can very well be person-based and what we see in different sets of data. Presently, most of what is done is a single data type approach, which is not very effective in reality. This paper reports on a multimodal stress detection platform that we have put together from physiological sensor data, speech inputs, and face images for real-time healthcare. We process heart rate variability in the physiological data, Mel-Frequency Cepstral Coefficients from speech, and face images through specialized-for each of the types of data. Also, we use independent machine learning models. In particular, a Random Forest model is used for physiological data, a Multi-Layer Perceptron for speech characteristics, and a ResNet-50 Convolutional Neural Network for facial images. On the other hand, based on our research, the speech base model performs well at a 97% accuracy level. Physiological and images perform at 88.28% and 72.07% accuracy levels, respectively. Similarly, in our research, we use decision level fusion, which is ensuring the stability of our performance, and if either of the sources of data is inadequate, we will still obtain accurate outcomes. Its aptness for use in real-time health care, in our view, constitutes a big advantage.

Article Details

Section

Articles

How to Cite

Soni, P., & Jain, D. D. (2026). Multimodal Stress Detection Using Sensor, Voice, and Facial Image Data for Real-Time Health Monitoring . International Journal of Aquatic Research and Environmental Studies, 6(S5), 618-624. https://injoere.com/index.php/injoere/article/view/1342

Similar Articles

You may also start an advanced similarity search for this article.