Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling

Logo

A research website showcasing speech emotion conversion audio examples, based on the paper titled "Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling" and presented at ASRU 2025. Includes audio samples demonstrating emotion transformation in speech, project details, and related resources.

Link to the research paper

Speech Emotion Conversion: Audio Samples with Duration Modeling

Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic properties such as fundamental frequency, spectral envelope and energy, but often lack the ability to control the duration of sounds. To address this, we propose a duration modeling framework using resynthesis-based discrete content representations, enabling modification of speech duration to reflect target emotions and achieve controllable speech rates without using parallel data. Experimental results reveal that the inclusion of the proposed duration modeling framework significantly enhances emotional expressiveness, in the in-the-wild MSP-Podcast dataset. Analyses show that low-arousal emotions correlate with longer durations and slower speech rates, while high-arousal emotions produce shorter, faster speech.

Audio Examples

Select an audio file:   

Ground-truth Audio:

Emotion Converted Speech

Low Arousal: 1 :

Low Arousal: 4 :

Low Arousal: 7 :


Citation

@inproceedings{rajprabhu2025durmodel,
    title={Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling},
    author={Raj Prabhu, Navin and de Oliveira, Danilo and Lehmann-Willenbrock, Nale and Gerkmann, Timo},
    booktitle={IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)},
    year={2025}
}

This website is subsequent work to the following research:

  1. N. Raj Prabhu, B. Lay, S. Welker, N. Lehmann-Willenbrock and T. Gerkmann, "EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild Data," ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 11651-11655, doi:10.1109/ICASSP48485.2024.10447372.
  2. N. Raj Prabhu, N. Lehmann-Willenbrock and T. Gerkmann, "In-the-wild Speech Emotion Conversion Using Disentangled Self-Supervised Representations and Neural Vocoder-based Resynthesis," Speech Communication; 15th ITG Conference, Aachen, 2023, pp. 176-180, doi:10.30420/456164034.