详细信息
文献类型:期刊文献
中文题名:基于对抗网络的声纹识别域迁移算法
英文题名:GAN-Based Domain Adaptation Algorithm for Speaker Verification
作者:季敏飞[1];陈宁[1]
机构:[1]华东理工大学信息科学与工程学院,上海200237
年份:2022
卷号:48
期号:2
起止页码:231
中文期刊名:华东理工大学学报(自然科学版)
外文期刊名:Journal of East China University of Science and Technology
收录:Scopus;北大核心:【北大核心2020】;CSCD:【CSCD_E2021_2022】;
基金:国家自然科学基金面上项目(61771196)。
语种:中文
中文关键词:声纹识别;迁移学习;对抗网络
外文关键词:speaker verification;domain adaptation;adversarial network
摘要:针对声纹识别任务中常常出现的由于真实场景语音与模型训练语料在内部特征(情感、语言、说话风格、年龄)或外部特征(背景噪声、传输信号、麦克风、室内混响)等方面的差异所导致的模型识别率低的问题,提出了一种基于对抗网络的声纹识别域迁移算法。首先,利用源域语音对X-Vector的声纹识别模型进行训练;然后,采用域迁移方法将源域训练的XVector模型迁移至目标域训练数据;最后,在目标域测试数据上检测迁移后的模型性能,并将其与迁移前的模型性能进行对比。实验中采用AISHELL1作为源域,采用VoxCeleb1和CNCeleb分别作为目标域对算法性能进行测试。实验结果表明,采用本文方法进行迁移后,在VoxCeleb1和CN-Celeb的目标域测试集上的等错误率分别下降了21.46%和19.24%。
A key problem in speaker verification task is the condition mismatch between the training data and the testing data,which may significantly affect the verification performance.In most of the speaker recognition application scenarios,it is usually impossible to obtain enough samples to retrain the speaker recognition model.At the same time,the samples that is used to train the original model usually may be quite different from those obtained in real applications due to the variability caused by the intrinsic factors(e.g.,the changes in emotion,language,vocal effect,speaking style,and aging,etc.)or extrinsic ones(e.g.,background noise,transmission channel,microphone,room acoustics,and distance from the microphone,etc.).In this paper,an adversarial domain adaptation strategy is designed and applied to the X-Vector-based speaker verification scheme to enhance its domain adaptation ability.First,the X-Vector scheme is trained on the source dataset(AISHELL1).Then,the domain adaptation strategy is applied to the obtained X-Vector scheme for enabling it adapt to the target dataset(VoxCeleb1 or CN-Celeb).Finally,the performances of the X-Vector schemes obtained before and after adaptation are compared via the target dataset,from which it is demonstrated that the proposed adaptation strategy achieves 21.46%and 19.24%Equal Error Rate(EER)reduction on VoxCeleb1 and CN-Celeb dataset,respectively.
参考文献:
正在载入数据...
