详细信息

基于QRNN-CTC的中文语音识别声学模型    

CHINESE SPEECH RECOGNITION ACOUSTIC MODEL BASED ON QRNN-CTC

文献类型:期刊文献

中文题名:基于QRNN-CTC的中文语音识别声学模型

英文题名:CHINESE SPEECH RECOGNITION ACOUSTIC MODEL BASED ON QRNN-CTC

作者:王先欢[1];孙自强[1]

机构:[1]华东理工大学化工过程先进控制和优化技术教育部重点实验室,上海200237

年份:2023

卷号:40

期号:12

起止页码:184

中文期刊名:计算机应用与软件

外文期刊名:Computer Applications and Software

收录:CSTPCD;;北大核心:【北大核心2020】;

基金:中央高校基本科研业务费专项资金资助项目(222201917006)。

语种:中文

中文关键词:深度学习;语音识别;声学模型;准循环神经网络;连接时序分类

外文关键词:Deep learning;Speech recognition;Acoustic model;QRNN;CTC

摘要:针对卷积神经网络(CNN)在语音识别中处理时序能力不足和循环神经网络(RNN)在语音识别中模型复杂度较高、训练慢的问题,提出一种新的基于准循环神经网络和连接时序主义(QRNN-CTC)的声学模型。该模型既降低了参数量,又保证了一定的时序间循环能力,利用CTC来实现输入序列和标签自动对齐,在训练时引入dropout防止过拟合。在Thchs-30数据集上的实验结果表明,QRNN-CTC比CNN-CTC相对错误率降低9.8%,最终词错误率为23.8%,训练时间为LSTM-CTC的一半。
Aimed at the problem of insufficient processing time sequence ability of convolutional neural network(CNN)in speech recognition and high model complexity and difficulty of training in recurrent neural network(RNN)in speech recognition,a new kind of quasi-recurrent neural network and connectionist temporal classification(QRNN-CTC)acoustic model is proposed.It not only reduced the numbers of parameters but also ensured a certain cycle capability between time series.CTC was used to realize automatic alignment of input sequence and label,and dropout was introduced to prevent overfitting during training.The experimental results on the Thchs-30 dataset show that QRNN-CTC has a relative error rate of 9.8% lower than that of CNN-CTC,and the final word error rate is 23.8%,and the training time is half of LSTM-CTC.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心