Audio samples of "Improving Model Stability and Training Efficiency in Fast Speed High Quality Expressive Voice Conversion System"

Abstract: Voice conversion (VC) systems have made significant progress owing to advanced deep learning methods. Current research is not only concerned with high-quality and fast audio synthesis, but also richer expressiveness. The most popular VC system was constructed from the concatenation of an automatic speech recognition module with a text-to-speech module (ASR-TTS). Yet this system suffers from errors in recognition and pronunciation and it also requires a large amount of data for a pre-trained ASR mode l. We propose an approach to improve the model stability and training efficiency of a VC system. Firstly, a data redundancy reduction method is used to balance the distribution of vocabulary to avoid uncommon words being ignored during the training process; by adding connectionist temporal classification (CTC) loss, the word error rate (WER) of our system reduces to 3.02%, which is 5.63 percentage points lower than that of the ASR-TTS system (8.65%), and the inference speed (e.g., real-time rate 19.32) of our VC system is much higher than that of the baseline system (real-time rate 2.24). Finally, emotional embedding is added to the pre-trained VC system to generate expressive speech conversion. The results show that after fine-tuning on the multi-emotional dataset, the system can achieve high quality and expressive speech synthesis.

[pdf][slides]

 

MOS Test (Neutral)


心动不如行动,我不太擅长卖萌。

看着你高兴,我就觉得很幸福啦。

被宰后的我简直气愤的不得了!

是悲伤的痛苦,是不尽的思念。

Ground Truth
Our Proposed
Kaldi (CVTE M2) + MTTS

MOS Test (Happiness)


心动不如行动,我不太擅长卖萌。

看着你高兴,我就觉得很幸福啦。

被宰后的我简直气愤的不得了!

是悲伤的痛苦,是不尽的思念。

Ground Truth
Our Proposed
Kaldi (CVTE M2) + MTTS

MOS Test (Anger)


心动不如行动,我不太擅长卖萌。

看着你高兴,我就觉得很幸福啦。

被宰后的我简直气愤的不得了!

是悲伤的痛苦,是不尽的思念。

Ground Truth
Our Proposed
Kaldi (CVTE M2) + MTTS

MOS Test (Sadness)


心动不如行动,我不太擅长卖萌。

看着你高兴,我就觉得很幸福啦。

被宰后的我简直气愤的不得了!

是悲伤的痛苦,是不尽的思念。

Ground Truth
Our Proposed
Kaldi (CVTE M2) + MTTS


Ablation Study (CTC)


(Anger)
躺椅不错,适合卧着悦读。
是不是靠植物太近了。

(Happiness)
边伺候他,边窥伺动静。

(Neutral)
分析的靠谱!
大辉,一如既往的牛人!

(Sadness)
上校到校场找人校对材料。

Our Proposed (with CTC)
Our Proposed (without CTC)
Kaldi (CVTE M2) + MTTS


Applications


Few-Shot Cross Lingual Voice Conversion (5min)

Source Utterance
音像产业非常重要。

Source Content
There are many examples of this in France.

Synthesised Utterance
There are many examples of this in France.

Female (ZH)
Male (ZH)

Fine-Grained Prosody Control (5min)

Source Emotion (Neutral)
大苹果,喜欢我,我爱吃,大苹果。

Source Emotion (Expressive)
大苹果,喜欢我,我爱吃,大苹果。

Target Speaker (ZH)

Synthesised Emotion (Level 0)

Synthesised Emotion (Level 1)

Synthesised Emotion (Level 2)

Synthesised Audios