Demos for "ByteSing: A Chinese Singing Voice Synthesis System Using Duration Allocated Encoder-Decoder Acoustic Models and WaveRNN Vocoders"

Abstract

This paper presents ByteSing, a Chinese singing voice synthesis (SVS) system based on duration allocated Tacotron-like acoustic models and WaveRNN neural vocoders. Different from the conventional SVS models, the proposed ByteSing employs Tacotron-like encoder-decoder structures as the acoustic models, in which the CBHG models and recurrent neural networks (RNNs) are explored as encoders and decoders respectively. Meanwhile an auxiliary phoneme duration prediction model is utilized to expand the input sequence, which can enhance the model controllable capacity, model stability and tempo prediction accuracy. WaveRNN neural vocoders are also adopted as neural vocoders to further improve the voice quality of synthesized songs. Both objective and subjective experimental results prove that the SVS method proposed in this paper can produce quite natural, expressive and high-fidelity songs by improving the pitch and spectrogram prediction accuracy and the models using attention mechanism can achieve best performance.

Demos

Groundtruth songs

song1
song2

Synthesized songs (remixed with BGM)

song1
song2

Acappella (without BGM)

song1
wav1
wav2
wav3
song2
wav1
wav2
wav3
song3
wav1
wav2
wav3
song4
wav1
wav2
wav3
song5
wav1
wav2
wav3
song6
wav1
wav2
wav3
song7
wav1
wav2
wav3
song8
wav1
wav2
wav3

Credit cookies

Some synthesized demos of a new singer using an improved version of ByteSing (to be published soon)
Song1
Song2
Song3
Song4