详细信息
基于可信细粒度对齐的多模态方面级情感分析
Aspect-based Multimodal Sentiment Analysis Based on Trusted Fine-grained Alignment
文献类型:期刊文献
中文题名:基于可信细粒度对齐的多模态方面级情感分析
英文题名:Aspect-based Multimodal Sentiment Analysis Based on Trusted Fine-grained Alignment
作者:范东旭[1];过弋[1,2,3]
机构:[1]华东理工大学信息科学与技术学院,上海200237;[2]大数据流通与交易技术国家工程实验室-商业智能与可视化技术研究中心,上海200436;[3]上海大数据与互联网受众工程技术研究中心,上海200072
年份:2023
卷号:50
期号:12
起止页码:246
中文期刊名:计算机科学
外文期刊名:Computer Science
收录:CSTPCD;;北大核心:【北大核心2020】;CSCD:【CSCD_E2023_2024】;
基金:上海市科学技术委员会科技计划项目(22DZ204903,22511104800)。
语种:中文
中文关键词:方面级情感分析;多模态;细粒度对齐;情感分析;自然语言处理
外文关键词:Aspect-based sentiment analysis;Multimodal;Fine-grained alignment;Sentiment analysis;Natural language processing
摘要:基于方面的多模态情感分析任务(Multimodal Aspect-Based Sentiment Analysis,MABSA),旨在根据文本和图像信息识别出文本中某特定方面词的情感极性。然而,目前主流的模型并没有充分利用不同模态之间的细粒度语义对齐,而是采用整个图像的视觉特征与文本中的每一个单词进行信息融合,忽略了图像视觉区域和方面词之间的强对应关系,这将导致图片中的噪声信息也被融合进最终的多模态表征中,因此提出了一个可信细粒度对齐模型TFGA(MABSA Based on Trusted Fine-grained Alignment)。具体来说,使用FasterRCNN捕获到图像中包含的视觉目标后,分别计算其与方面词之间的相关性,为了避免视觉区域与方面词的局部语义相似性在图像文本的全局角度不一致的情况,使用置信度对局部语义相似性进行加权约束,过滤掉不可靠的匹配对,使得模型重点关注图片中与方面词相关性最高且最可信的视觉局域信息,降低图片中多余噪声信息的影响;接着提出细粒度特征融合机制,将聚焦到的视觉信息与文本信息进行充分融合,以得到最终的情感分类结果。在Twitter数据集上进行实验,结果表明,文本与视觉的细粒度对齐对方面级情感分析是有利的。
Aspect based multimodal sentiment analysis task(MABSA)aims to identify the sentiment polarity of a specific aspect word in a text based on text and image information.However,the current mainstream model does not make full use of the fine-grained semantic alignment between different modes.Instead,it uses the image features of the entire image to fuse information with each word in the text,ignoring the strong correspondence between the local image information and aspect words,which will lead to the noise information in the image being integrated into the final multimodal representation,Therefore,this paper proposes a trusted fine-grained alignment model TFGA(MABSA based on trusted fine-grained alignment).Specifically,we use FasterRCNN to capture the visual objects contained in the image,and then calculate the correlation between them and aspect words respectively.To avoid the inconsistency of the local semantic similarity between the visual object and aspect words in the global perspective of the image-text,confidence is used to weight the local semantic similarity and filter out the unreliable matching pairs,then the model can focuse on the most reliable and highest visual local information related to aspect words in the image to reduce the impact of redundant noise information in the image.Then a fine-grained feature fusion mechanism is proposed to fully fuse the focused local image information with the text information to obtain the final sentiment classification result.Experiments on Twitter datasets show that fine-grained alignment of text and vision is beneficial to aspect based sentiment analysis.
参考文献:
正在载入数据...
