详细信息

基于宽深学习网络的音乐情感识别    

Music Emotion Recognition Based on the Broad and Deep Learning Network

文献类型:期刊文献

中文题名:基于宽深学习网络的音乐情感识别

英文题名:Music Emotion Recognition Based on the Broad and Deep Learning Network

作者:王晶晶[1];黄如[1]

机构:[1]华东理工大学信息科学与工程学院,上海200237

年份:2022

卷号:48

期号:3

起止页码:373

中文期刊名:华东理工大学学报(自然科学版)

外文期刊名:Journal of East China University of Science and Technology

收录:Scopus;北大核心:【北大核心2020】;CSCD:【CSCD_E2021_2022】;

基金:国家自然科学基金(61673178,61922063);上海市自然科学基金(20ZR1413800)。

语种:中文

中文关键词:音乐情感识别;残差相位;宽度学习;深度学习;长短期记忆网络

外文关键词:music emotion recognition;residual phase;broad learning;deep learning;long short-term memory

摘要:将梅尔频率倒谱系数(Mel Frequency Cepstral Coefficient,MFCC)和残差相位(Residual Phase, RP)进行加权结合来提取音乐情感特征,提高了音乐情感特征的挖掘效率;同时为了提高音乐情感的分类精度,缩短模型训练时间,将长短期记忆网络(Long Short-Term Memory,LSTM)和宽度学习系统(Broad Learning System,BLS)相结合,使用LSTM作为BLS的特征映射节点,搭建了一种新型宽深学习网络(LSTM-BLS)进行音乐情感识别分类训练。在Emotion数据集上的实验结果表明,本文算法取得了比其他复杂网络更高的识别准确率,为音乐情感识别的发展提供了新的可行性思路。
With the development of artificial intelligence and digital audio technology, music information retrieval(MIR) has gradually become a research hotspot. Meanwhile, music emotion recognition(MER) is becoming an important research direction, due to its great research value for video soundtracks. Although some researchers combine Mel Frequency Cepstral coefficient(MFCC) and Residual Phase(RP) to extract music emotional features and improve classification accuracy, the training models in traditional deep learning takes longer time. In order to improve the efficiency of feature mining of music emotional features, MFCC and RP are weighted and combined in this work to extract music emotion features so that the mining efficiency of music emotion features can be effectively improved. At the same time, in order to improve the classification accuracy of music emotion and shorten the training time of the model, by integrating the Long Short-Term Memory(LSTM) and the Broad Learning System(BLS), a new wide and deep learning network(LSTM-BLS) is further built to train music emotion recognition and classification by using LSTM as the feature mapping node of BLS. The network structure of this model makes full use of the ability of BLS to quickly process complex data. Its advantages are simple structure and short model training time, thereby improving recognition efficiency, and LSTM has excellent performance in extracting time series features from time series data.The time sequence relationship of music can be extracted so that the emotional characteristics of the music can be preserved to the greatest extent. Finally, the experimental results on the emotion dataset show that the proposed algorithm can achieve higher recognition accuracy than other complex networks and provide new feasible ideas for the music emotion recognition.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心