Lightweight Emotional Speech Transformation using F0 Transplantation and Acoustic Feature Reprojection
Main Article Content
Abstract
In emotional environments, speaker recognition systems often experience significant performance degradation due to the variability introduced by emotional speech. To address this challenge, this paper proposes a lightweight emotional speech transformation framework that synthesizes emotionally expressive speech directly from neutral utterances, reducing the need for large emotional corpora or deep generative models. The proposed approach employs an F0 transplantation and prosodic modulation technique, where the fundamental frequency (F0) and energy contours from emotionally expressive donor utterances are imposed onto neutral speech while preserving the speaker's identity. Using the Berlin Emotional Speech Database (Emo-DB), the system synthesizes emotional variants of neutral samples and quantitatively evaluates them against real emotional references. The evaluation shows that 73.6% of the transformed samples achieved a positive emotional shift with an average similarity gain of 0.0154, while maintaining high speaker identity preservation (93.8% Top-1 accuracy). The proposed method demonstrates that signal-level prosody adaptation can produce perceptually meaningful emotional transformations without requiring parallel data or model training. This framework provides an interpretable, data-efficient step toward bridging the gap between neutral and emotional speech.
