A Multimodal Stress Detection System Using Text, Audio, and Video Analysis with CMII-Based Weighted Fusion
Main Article Content
Abstract
Mental stress has become one of the widely debated issues of relative importance to the overall productivity and quality of life, and traditional self-reported measurement systems are biased and subjective. The present paper is a proposal of privacy-saving multimodal stress detection system that takes into account three modalities, namely text-based sentiment analysis, audio-based vocal pattern analysis and facial micro-expression recognition. A zero-retention policy is used in the architecture, which means that raw biometric data are never stored and are only processed in volatile memory. The system combines Support Vector Machines and TF-IDF to analyze text, CNN-LSTM to analyze audio, and ResNet-based deep learning to analyze facial expressions. An innovative approach to weighted late fusion strategy and the use of a new Cross-Modal Inconsistency Index makes it possible to classify stress and identify masked stress. The non-regulatory adherence to the GDPR and the HIPAA is guaranteed by role-based access control, metadata cryptography protection, and audit logs. The findings prove that it is possible to identify multimodal stress without interfering with individual privacy and this is relevant to the research in affective computing.
