详细信息

Audio-Visual Scene Classification Based on Multi-modal Graph Fusion  ( CPCI-S收录)  

文献类型:会议论文

英文题名:Audio-Visual Scene Classification Based on Multi-modal Graph Fusion

作者:Lei, Han[1];Chen, Ning[1]

机构:[1]East China Univ Sci & Technol, Shanghai, Peoples R China

会议论文集:Interspeech Conference

会议日期:SEP 18-22, 2022

会议地点:Incheon, SOUTH KOREA

语种:英文

外文关键词:Audio-visual scene classification; Graph convolutional network; Similarity fusion

摘要:Audio-Visual Scene Classification (AVSC) task tries to achieve scene classification through joint analysis of the audio and video modalities. Most of the existing AVSC models are based on feature-level or decision-level fusion. The possible problems are: i) Due to the distribution difference of the corresponding features in different modalities is large, the direct concatenation of them in the feature-level fusion may not result in good performance. ii) The decision-level fusion cannot take full advantage of the common as well as complementary properties between the features and corresponding similarities of different modalities. To solve these problems, Graph Convolutional Network (GCN)-based multi-modal fusion algorithm is proposed for AVSC task. First, the Deep Neural Network (DNN) is trained to extract essential feature from each modality. Then, the Sample-to-Sample Cross Similarity Graph (SSCSG) is constructed based on each modality features. Finally, the DynaMic GCN (DM-GCN) and the ATtention GCN (AT-GCN) are introduced respectively to realize both feature-level and similarity-level fusion to ensure the classification accuracy. Experimental results on TAU Audio-Visual Urban Scenes 2021 development dataset demonstrate that the proposed scheme, called AVSC-MGCN achieves higher classification accuracy and lower computational complexity than state-of-the-art schemes.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心