详细信息

Enhancing Multimodal Video Summarization viaTemporal andSemantic Alignment  ( EI收录)  

文献类型:期刊文献

英文题名:Enhancing Multimodal Video Summarization viaTemporal andSemantic Alignment

作者:Ying, Fangli[1]; Luo, Ziyue[2]; Phaphuangwittayakul, Aniwat[3]

机构:[1] Department of Computer Science and Engineering, State Key Laboratory of Bioreactor Engineering, East China University of Science and Technology, Shanghai, China; [2] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China; [3] International College of Digital Innovation, Chiang Mai University, Chiang Mai, Thailand

年份:2025

卷号:2290 CCIS

起止页码:17

外文期刊名:Communications in Computer and Information Science

收录:EI(收录号:20252818764209)

语种:英文

外文关键词:Alignment - Contrastive Learning - Dynamics - Learning systems - Semantic Web - Semantics - Text processing - Unsupervised learning - Video analysis

摘要:Video summarization condenses video content into concise and informative summaries. To enhance the quality of video summaries, emerging multimodal approaches integrate visual and textual information, thereby outperforming conventional methods that rely solely on visual cues. However, the existing methods fail to effectively model the temporal dynamics within each modality and ensure semantic alignment across modalities simultaneously. In this paper, we introduce an unsupervised method for learning unified representations of video and text to address these challenges. Specifically, we propose a novel dual attentive network framework, which enhances the video representation through conditional text-derived information at a local scale and models long-term cross-modal dependencies at a global scale to leverage the temporal information across different scales of data. To further refine this model, we incorporate a hard negatives loss function within a contrastive learning framework, which learns to identify the irrelevant visual-textual representation pairs that closely resemble the relevant ones. Additionally, we propose a Dynamic Time Warping-based temporal alignment loss to maintain coherent sequential constraints over time within the same modality, addressing intra-modal dynamics. To evaluate our approach, we validate extensive experiments on standard video summarization datasets. The experimental results not only highlight the superiority of our approach but also emphasize its potential for practical applications in various domains. ? The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心