详细信息

Automatic Speech Disentanglement for Voice Conversion using Rank Module and Speech Augmentation  ( EI收录)  

文献类型:期刊文献

英文题名:Automatic Speech Disentanglement for Voice Conversion using Rank Module and Speech Augmentation

作者:Liu, Zhonghua[1]; Wang, Shijun[2]; Chen, Ning[1]

机构:[1] East China University of Science and Technology, Shanghai, China; [2] University of St.Gallen, St.Gallen, Switzerland

年份:2023

卷号:2023-August

起止页码:2298

外文期刊名:Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH

收录:EI(收录号:20230226640)

语种:英文

外文关键词:Speech communication - Tuning

摘要:Voice Conversion (VC) converts the voice of a source speech to that of a target while maintaining the source's content. Speech can be mainly decomposed into four components: content, timbre, rhythm and pitch. Unfortunately, most related works only take into account content and timbre, which results in less natural speech. Some recent works are able to disentangle speech into several components, but they require laborious bottleneck tuning or various hand-crafted features, each assumed to contain disentangled speech information. In this paper, we propose a VC model that can automatically disentangle speech into four components using only two augmentation functions, without the requirement of multiple hand-crafted features or laborious bottleneck tuning. The proposed model is straightforward yet efficient, and the empirical results demonstrate that our model can achieve a better performance than the baseline, regarding disentanglement effectiveness and speech naturalness. ? 2023 International Speech Communication Association. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心