Spatio-Spectral Hypergraph Attention Network for Multi-Speaker Emotion Separation in High Noise Acoustic Environments
DOI:
https://doi.org/10.67440/ahj.vi.2038Keywords:
Multi-speaker speech emotion recognition, Hypergraph neural network, Speech separation, Acoustic noise, Spatio-spectral representation, Speaker disentanglement, Deep learning.Abstract
Multi-speaker speech emotion recognition becomes challenging when speech signals overlap and are corrupted by acoustic noise, as interference between speakers can obscure emotion-related acoustic characteristics. Existing approaches often model individual speech representations or pairwise relationships but have limited ability to capture higher-order interactions among concurrent speakers. This study proposes a Spatio-Spectral Hypergraph Attention Network (SS-HAN) for multi-speaker emotion recognition under noisy and overlapping acoustic conditions. The framework combines spatial-spectral acoustic representation with hypergraph-based relational learning to capture interactions among multiple simultaneous speech components. In contrast, we introduce an emotional resonance propagation mechanism to enhance the emotion-related representations and adopt a cross-speaker interference suppression mechanism to alleviate the emotional information leakage between speakers. We perform experiments on IEMOCAP, RAVDESS, and MSP-IMPROV with overlapping mixtures created under controlled acoustic conditions. The proposed framework is evaluated by accuracy of emotion recognition, F1-score, improvement of signal-to-interference ratio, and speaker disentanglement metrics. The experimental results show that SS-HAN provides improved performance when compared to the evaluated baseline approaches while maintaining competitive performance for different datasets and acoustic conditions. Ablation experiments further examine the contribution of the individual components of the proposed framework. The results indicate that modeling higher-order relationships between concurrent speakers can improve emotion representation and speaker-level emotion recognition in complex acoustic environments.

