详细信息

MMpedia: A Large-Scale Multi-modal Knowledge Graph  ( EI收录)  

文献类型:期刊文献

英文题名:MMpedia: A Large-Scale Multi-modal Knowledge Graph

作者:Wu, Yinan[1]; Wu, Xiaowei[1]; Li, Junwen[1]; Zhang, Yue[1]; Wang, Haofen[2]; Du, Wen[3]; He, Zhidong[3]; Liu, Jingping[1]; Ruan, Tong[1]

机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] College of Design and Innovation, Tongji University, Shanghai, China; [3] DS Information Technology, Shanghai, China

年份:2023

卷号:14266 LNCS

起止页码:18

外文期刊名:Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)

收录:EI(收录号:20234715098365)

语种:英文

外文关键词:Computational linguistics - Graphic methods - HTTP - Image enhancement - Natural language processing systems - Search engines

摘要:Knowledge graphs serve as crucial resources for various applications. However, most existing knowledge graphs present symbolic knowledge in the form of natural language, lacking other modal information, e.g., images. Previous multi-modal knowledge graphs have encountered challenges with scaling and image quality. Therefore, this paper proposes a highly-scalable and high-quality multi-modal knowledge graph using a novel pipeline method. Summarily, we first retrieve images from a search engine and build a new Recurrent Gate Multi-modal model to filter out the non-visual entities. Then, we utilize entities’ textual and type information to remove noisy images of the remaining entities. Through this method, we construct a large-scale multi-modal knowledge graph named MMpedia, containing 2,661,941 entity nodes and 19,489,074 images. As we know, MMpedia has the largest collection of images among existing multi-modal knowledge graphs. Furthermore, we employ human evaluation and downstream tasks to verify the usefulness of images in MMpedia. The experimental result shows that both the state-of-the-art method and multi-modal large language model (e.g., VisualChatGPT) achieve about a 4% improvement on Hit@1 in the entity prediction task by incorporating our collected images. We also find that the multi-modal large language model is hard to ground entities to images. The dataset (https://zenodo.org/record/7816711 ) and source code of this paper are available at https://github.com/Delicate2000/MMpedia. ? The Author(s), under exclusive license to Springer Nature Switzerland AG 2023.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心