详细信息

基于帧级特征的端到端说话人识别    

End-to-End Speaker Recognition Based on Frame-level Features

文献类型:期刊文献

中文题名:基于帧级特征的端到端说话人识别

英文题名:End-to-End Speaker Recognition Based on Frame-level Features

作者:花明[1];李冬冬[1];王喆[1];高大启[1]

机构:[1]华东理工大学信息科学与工程学院,上海200237

年份:2020

卷号:47

期号:10

起止页码:169

中文期刊名:计算机科学

外文期刊名:Computer Science

收录:CSTPCD;;北大核心:【北大核心2017】;CSCD:【CSCD_E2019_2020】;

基金:国家自然科学基金项目(61806078);国家重大新药开发科技专项(2019ZX0921004);上海市教育发展基金会和上海市教育委员会“曙光计划”(61725301)。

语种:中文

中文关键词:说话人识别;端到端;卷积神经网络;帧级特征;话语级语音

外文关键词:Speaker recognition;End-to-end;Convolutional Neural Networks;Frame-level features;Utterance-level speech

摘要:现有的说话人识别方法仍存在许多不足。基于话语级特征输入的端到端方法由于语音长短不一致需要将输入处理为同等大小,而特征训练加后验分类的两阶段方法使得识别系统过于复杂,这些因素都会影响模型的性能。文中提出了基于帧级特征的端到端说话人识别方法。模型采用帧级语音作为输入,同等大小的帧级特征有效解决了话语级语音输入长度不一致的问题,且帧级特征可保留更多的话者信息。与如今主流的两阶段法识别系统相比,端到端的识别方法将特征训练和分类打分一体化,简化了模型的复杂性。在训练阶段,每段语音被分帧成多个帧级语音输入到卷积神经网络(Convolutional Neural Networks,CNN)用于训练模型。在评估阶段,训练好的CNN模型对帧级语音进行分类,每段语音基于多个帧的预测得分计算该条语音数据的预测类别。每段语音的类别通过取各帧最多预测类别和各帧预测值平均的方法来计算。为了验证方法的有效性,使用普通话情感语音语料库(MASC)的语音数据进行训练和测试。实验结果表明,与现有方法相比,基于帧级特征的端到端识别方法的性能表现更佳。
There are still many shortcomings in the existing speaker recognition methods.The end-to-end method based on utte-rance-level features requires to process the input to be the same size due to the inconsistency of the speech length.The two-stage method of feature training with posterior classification makes the recognition system too complex.These factors affect the performance of the model.This paper proposed an end-to-end speaker recognition method based on frame-level features.The model uses frame-level speech as input,and the same size frame-level features effectively solve the problem of inconsistent speech-level speech input length,and the frame-level features can retain more speaker information.Compared with the mainstream two-stage identification system,the end-to-end identification method integrates feature training and classification,which simplifies the complexity of the model.During the training phase,each speech is segmented into multiple frame-level speech inputs to a Convolutional Neural Network(CNN)for training the model.In the evaluation phase,the trained CNN model classifies the frame-level speech,and each segment of speech calculates the prediction category of the speech data based on the prediction scores of multiple frames.The maximum predicted category of each frame and the average prediction value of each frame are adopted to calculate the class of each segment of speech respectively.In order to verify the validity of the work,the speech data of the Mandarin Emotio-nal Speech Corpus(MASC)were used for training and testing.The experimental results show that the end-to-end recognition method based on frame-level features achieves better performance than the existing methods.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心