详细信息
Enhancing Multimodal Video Summarization viaTemporal andSemantic Alignment ( EI收录)
文献类型:期刊文献
英文题名:Enhancing Multimodal Video Summarization viaTemporal andSemantic Alignment
作者:Ying, Fangli[1]; Luo, Ziyue[2]; Phaphuangwittayakul, Aniwat[3]
机构:[1] Department of Computer Science and Engineering, State Key Laboratory of Bioreactor Engineering, East China University of Science and Technology, Shanghai, China; [2] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China; [3] International College of Digital Innovation, Chiang Mai University, Chiang Mai, Thailand
年份:2025
卷号:2290 CCIS
起止页码:17
外文期刊名:Communications in Computer and Information Science
收录:EI(收录号:20252818764209)
语种:英文
外文关键词:Alignment - Contrastive Learning - Dynamics - Learning systems - Semantic Web - Semantics - Text processing - Unsupervised learning - Video analysis
摘要:Video summarization condenses video content into concise and informative summaries. To enhance the quality of video summaries, emerging multimodal approaches integrate visual and textual information, thereby outperforming conventional methods that rely solely on visual cues. However, the existing methods fail to effectively model the temporal dynamics within each modality and ensure semantic alignment across modalities simultaneously. In this paper, we introduce an unsupervised method for learning unified representations of video and text to address these challenges. Specifically, we propose a novel dual attentive network framework, which enhances the video representation through conditional text-derived information at a local scale and models long-term cross-modal dependencies at a global scale to leverage the temporal information across different scales of data. To further refine this model, we incorporate a hard negatives loss function within a contrastive learning framework, which learns to identify the irrelevant visual-textual representation pairs that closely resemble the relevant ones. Additionally, we propose a Dynamic Time Warping-based temporal alignment loss to maintain coherent sequential constraints over time within the same modality, addressing intra-modal dynamics. To evaluate our approach, we validate extensive experiments on standard video summarization datasets. The experimental results not only highlight the superiority of our approach but also emphasize its potential for practical applications in various domains. ? The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
参考文献:
正在载入数据...
