详细信息
Affect-insensitive speaker recognition systems via emotional speech clustering using prosodic features ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Affect-insensitive speaker recognition systems via emotional speech clustering using prosodic features
作者:Li, Dongdong[1,2];Yuan, Yubo[1];Wu, Zhaohui[3];Yang, Yingchun[3]
机构:[1]E China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[2]Georgia Inst Technol, Sch Elect & Comp Engn, Ctr Signal & Image Proc, Atlanta, GA 30332 USA;[3]Zhejiang Univ, Dept Comp Sci & Technol, Hangzhou 310027, Peoples R China
年份:2015
卷号:26
期号:2
起止页码:473
外文期刊名:NEURAL COMPUTING & APPLICATIONS
收录:;EI(收录号:20144200096523);WOS:【SCI-EXPANDED(收录号:WOS:000348451100022)】;
基金:The author would like to offer sincere thanks to reviewers. Their comments and suggestions are very important to improve the presentation and technical sounds. This research was supported by Nature Science Foundation of Shanghai Municipality, China (No. 11ZR1409600) and partly supported by the Natural Science Foundation of China (No. 61272198, No. 1272198, No. 90924013, No. 91324010), Innovation Program of Shanghai Municipal Education Commission (No. 14ZZ054). This work is also supported by the Fundamental Research Funds for the Central Universities of China.
语种:英文
外文关键词:Speaker recognition; Emotional speech clustering; Prosodic features
摘要:Voice-based biometric security systems involving only neutral speech have achieved promising performance. However, the speakers are very likely to fail the recognition when the test data exhibit multiple emotions. This paper aimed to address the mismatch of the emotional states between training and testing speech. We discuss different modeling strategies that incorporate the emotions (affects) of speakers into the training stage of a Mandarin-based speaker recognition system and propose an alternative approach, which could optimize the utilization of the limited affective speech. The training speeches are partitioned and clustered by the trends of the prosodic variations. Multiple models are built based on the clustered speech for a given speaker. The prosodic differences are characterized by a combination of features that describe the changes of the fundamental frequencies and energy contours. The experiments were carried out based on the Mandarin Affective Speech Corpus. The result shows 73.37 % improvement in recognition rate over that of the traditional speaker verification tasks relatively and also achieves 63.53 % higher in performance over the structural training-based systems relatively.
参考文献:
正在载入数据...
