详细信息
MMpedia: A Large-Scale Multi-modal Knowledge Graph ( EI收录)
文献类型:期刊文献
英文题名:MMpedia: A Large-Scale Multi-modal Knowledge Graph
作者:Wu, Yinan[1]; Wu, Xiaowei[1]; Li, Junwen[1]; Zhang, Yue[1]; Wang, Haofen[2]; Du, Wen[3]; He, Zhidong[3]; Liu, Jingping[1]; Ruan, Tong[1]
机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] College of Design and Innovation, Tongji University, Shanghai, China; [3] DS Information Technology, Shanghai, China
年份:2023
卷号:14266 LNCS
起止页码:18
外文期刊名:Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
收录:EI(收录号:20234715098365)
语种:英文
外文关键词:Computational linguistics - Graphic methods - HTTP - Image enhancement - Natural language processing systems - Search engines
摘要:Knowledge graphs serve as crucial resources for various applications. However, most existing knowledge graphs present symbolic knowledge in the form of natural language, lacking other modal information, e.g., images. Previous multi-modal knowledge graphs have encountered challenges with scaling and image quality. Therefore, this paper proposes a highly-scalable and high-quality multi-modal knowledge graph using a novel pipeline method. Summarily, we first retrieve images from a search engine and build a new Recurrent Gate Multi-modal model to filter out the non-visual entities. Then, we utilize entities’ textual and type information to remove noisy images of the remaining entities. Through this method, we construct a large-scale multi-modal knowledge graph named MMpedia, containing 2,661,941 entity nodes and 19,489,074 images. As we know, MMpedia has the largest collection of images among existing multi-modal knowledge graphs. Furthermore, we employ human evaluation and downstream tasks to verify the usefulness of images in MMpedia. The experimental result shows that both the state-of-the-art method and multi-modal large language model (e.g., VisualChatGPT) achieve about a 4% improvement on Hit@1 in the entity prediction task by incorporating our collected images. We also find that the multi-modal large language model is hard to ground entities to images. The dataset (https://zenodo.org/record/7816711 ) and source code of this paper are available at https://github.com/Delicate2000/MMpedia. ? The Author(s), under exclusive license to Springer Nature Switzerland AG 2023.
参考文献:
正在载入数据...
