详细信息

CONSTRUCTURE: Benchmarking CONcept STRUCTUre REasoning for Multimodal Large Language Models  ( EI收录)  

文献类型:期刊文献

英文题名:CONSTRUCTURE: Benchmarking CONcept STRUCTUre REasoning for Multimodal Large Language Models

作者:Zha, Zhiwei[1,3]; Zhu, Xiangru[1]; Xu, Yuanyi[1]; Huang, Chenghua[1]; Liu, Jingping[2]; Li, Zhixu[4,5]; Wang, Xuwu[5]; Xiao, Yanghua[1]; Yang, Bei[3]; Xu, XiaoXiao[3]

机构:[1] Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University, China; [2] School of Information Science and Engineering, East China University of Science and Technology, China; [3] Alibaba Group, China; [4] School of Information, Renmin University of China, China; [5] Suzhou Key Laboratory of Artificial Intelligence and Social Governance Technologies, International College [Suzhou Research Institute], Renmin University of China, China

年份:2024

起止页码:4954

外文期刊名:EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024

收录:EI(收录号:20250717872781)

语种:英文

外文关键词:Benchmarking - Computational linguistics - Visual languages

摘要:Multimodal Large Language Models (MLLMs) have shown promising results in various tasks, but their ability to perceive the visual world with deep, hierarchical understanding similar to humans remains uncertain. To address this gap, we introduce CONSTRUCTURE, a novel concept-level benchmark to assess MLLMs' hierarchical concept understanding and reasoning abilities. Our goal is to evaluate MLLMs across four key aspects: 1) Understanding atomic concepts at different levels of abstraction; 2) Performing upward abstraction reasoning across concepts; 3) Achieving downward concretization reasoning across concepts; and 4) Conducting multi-hop reasoning between sibling or common ancestor concepts. Our findings indicate that even state-of-the-art multimodal models struggle with concept structure reasoning (e.g., GPT-4o averages a score of 62.1%). We summarize key findings of MLLMs in concept structure reasoning evaluation. Morever, we provide key insights from experiments using CoT prompting and fine-tuning to enhance their abilities. ? 2024 Association for Computational Linguistics.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心