详细信息
文献类型:期刊文献
中文题名:改进深度信念网络在语音转换中的应用
英文题名:Improved deep belief network and its application in voice conversion
作者:王文浩[1];张筱[1];万永菁[1]
机构:[1]华东理工大学信息科学与工程学院
年份:2019
卷号:53
期号:12
起止页码:2372
中文期刊名:浙江大学学报(工学版)
外文期刊名:Journal of Zhejiang University:Engineering Science
收录:CSTPCD;;EI(收录号:20200107978706);Scopus;北大核心:【北大核心2017】;CSCD:【CSCD2019_2020】;
基金:国家自然科学基金资助项目(61872143)
语种:中文
中文关键词:深度信念网络(DBN);语音转换;区域融合谱特征;误差修正网络;谱失真度
外文关键词:deep belief network(DBN);voice conversion;regional fusion spectral feature;error correction network;spectral distortion
摘要:综合考虑语音帧间关系及后处理网络的效果,提出一种改进的基于深度信念网络(DBN)的语音转换方法.该方法利用线性预测分析-合成模型提取说话人线性预测谱的特征参数,构建基于区域融合谱特征参数的深度信念网络用以预训练模型,经过微调阶段后引入误差修正网络以实现细节谱特征的补偿.对比实验结果表明,随着训练语音帧数的增加,转换语音的谱失真呈下降趋势.同时,在训练语音帧数较少的情况下,改进方法在异性间转换的谱失真小于50%,在同性间转换的谱失真小于60%.实验结果表明,改进方法的谱失真度较传统方法降低约6.5%,且同性别间转换效果比异性间转换效果更为明显,转换后语音的自然度和可理解度明显提高.
An improved voice conversion method based on deep belief network(DBN) was proposed,comprehensively considering the relationship between the speech frames and the effect of post-processing network.The method utilized a linear predictive analysis-synthesis model to extract the feature parameters of a speaker’s linear predictive spectrum, and the regional fusion spectral feature parameters for DBN were constructed so as to pretrain the model. Finally, an error correction network for the feature compensation of a detailed spectrum was introduced after fine-tuning. The comparison results show that, the spectral distortion of the converted speech shows the tendency of decreasing as the number of speech frames increases. Meanwhile, when the number of training speech frames was small, the spectral distortion of the proposed method was less than 50% between genders and less than 60% within genders. The experimental results showed that the spectral distortion of the proposed method was6.5% lower than that of the traditional method. The proposed method significantly improves the naturalness and intelligibility of converted speech in view of two different subjective evaluations.
参考文献:
正在载入数据...
