详细信息

基于语音音素后验概率图关键特征提取的中文方言识别模型    

A Chinese Dialect Identification Model Based on Key Feature Extraction from Phonetic Posteriorgram

文献类型:期刊文献

中文题名:基于语音音素后验概率图关键特征提取的中文方言识别模型

英文题名:A Chinese Dialect Identification Model Based on Key Feature Extraction from Phonetic Posteriorgram

作者:冯罡[1];陈宁[1]

机构:[1]华东理工大学信息科学与工程学院,上海200237

年份:2023

卷号:49

期号:6

起止页码:900

中文期刊名:华东理工大学学报(自然科学版)

外文期刊名:Journal of East China University of Science and Technology

收录:Scopus;北大核心:【北大核心2020】;CSCD:【CSCD_E2023_2024】;

基金:国家自然科学基金面上项目(61771196)。

语种:中文

中文关键词:方言识别;音素特征;自注意力机制;ECAPA-TDNN;特征提取

外文关键词:dialect identification;phonetic feature;self-attention mechanism;ECAPA-TDNN;feature extractor

摘要:不同方言对相同字的发音往往有所不同,因此不同方言所包含音素的概率分布存在较大差异,这是方言差异性的重要体现。为了充分利用这一差异性,提出了基于音素后验概率图分析的方言识别模型,该模型引入Convolutional Block Attention Module(CBAM)的提取音素后验概率图关键特征,并利用Emphasized Channel Attention-Propagation and Aggregation in TDNN(ECAPA-TDNN)模型对其进行聚合和注意力池化得到句子级特征。为进一步提升类间距离,引入了Additive Angular Margin(AAM)损失。实验结果表明,该模型取得了比传统模型更高的分类准确率,并且以上改进均对准确率提升有所贡献。
There are relatively few existing dialect recognition models for phonemic features and different dialects have different pronunciations,all of which lead to large differences in the probability distribution of phonemes contained in different dialects.Aiming at the above issues,this paper proposes a dialect identification model based on the phonetic posteriorgram feature.For the single dimension of attention analysis,this model extracts key features of frame-level phonetic posteriorgram by using the self-attention mechanism of Convolutional Block Attention Module(CBAM).At the same time,in order to make full use of the information in the middle layers of the model and avoid the loss of dialect information,Emphasized Channel Attention-Propagation and Aggregation in TDNN(ECAPA-TDNN)model is used to extract long-range information of frame-level feature and obtain effective sentence-level features via feature aggregation and attention statistical pooling.Finally,in order to avoid the problem of single loss function,we introduce Additive Angular Margin loss based on cross-entropy loss and replace the decision boundary with decision region to maximize inter-class distance across dialects and optimize the classification decision.It is shown via experimental results on Aishell2 and Datatang-Dialect datasets that the proposed model can achieve higher performance than the traditional model.All the above improvements contribute to the improvement of model performance.Meanwhile,the results of the ablation experiments demonstrate that these improvements in phonetic posteriorgram features,convolutional block attention module,ECAPA-TDNN,and additive angular margin loss,contribute to the improvement of identification accuracy.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心