Multi-Horizon Stress Forecasting From Emotional Speech Sequences Using Temporal Convolutional Networks and Self-Attention
DOI:
https://doi.org/10.67440/ahj.vi.2500Keywords:
Stress Forecasting, Temporal Convolutional Networks, Multi-Horizon Prediction, Affective Computing, Speech Emotion Recognition, Self-Attention Mechanism.Abstract
Classic stress detection is based on fixed classification of the existing states which restricts proactive interventions. The paper presents a new multi-horizon forecasting model to estimate the future stress levels (3, 6 and 9 seconds into the future) using temporal speech-derived sequences. We introduce a Multi-Task Temporal Convolutional Network with Attention (MT-TCN-Att), that combines dilated causal convolution, multi-head self-attention, and learnable positional encoding to learn complex temporal relationships. A composite loss of time smoothness and mixup augmentation is used to optimize the model. The model performs well on CREMA-D dataset, with an F1-score of 0.9208 at a horizon of 9 seconds. Ablation experiments indicate that a 30-second lookback window is optimal in capturing deterministic emotional transitions whereas the learnable positional encoding can be naturally explained by emphasising the significance of historical baselines. Zero-shot cross-dataset tests on RAVDESS and EMOdB, however, reveal the presence of important domain shifts, demonstrating the difficulty of trained performance generalization to different acoustic contexts. Finally, this paper has provided the first step towards adaptable horizons of acoustic stress prediction allowing the real-time, proactive approach to mental health interventions.

