详细信息

基于对齐查询的跨语言信息检索方法    

Cross-lingual Information Retrieval Based on Aligned Query

文献类型:期刊文献

中文题名:基于对齐查询的跨语言信息检索方法

英文题名:Cross-lingual Information Retrieval Based on Aligned Query

作者:李俊文[1];宋雨秋[2];张维彦[2];阮彤[2];刘井平[2];朱焱[1]

机构:[1]华东理工大学数学学院,上海200237;[2]华东理工大学信息科学与工程学院,上海200237

年份:2025

卷号:52

期号:8

起止页码:259

中文期刊名:计算机科学

外文期刊名:Computer Science

收录:;北大核心:【北大核心2023】;

语种:中文

中文关键词:跨语言信息检索;对齐查询;自指导;自适应层级系数

外文关键词:Cross-lingual Information Retrieval;Aligned query;Self-teaching;Adaptive layer-wise coefficient

摘要:跨语言信息检索是自然语言处理中一项重要的信息获取任务。最近,基于大语言模型的检索方法在这一任务中获得了广泛关注并取得了显著的进展。然而,现有基于提示大语言模型的无监督检索方法在效果和效率上仍有不足。对此,提出了一种全新的基于对齐查询的跨语言信息检索方法。具体而言,采用“预训练-微调”范式,基于预训练多语言模型提出了一种自适应的自指导编码器,通过同一语言内的检索学习指导跨语言检索学习。该方法引入与文档语种相同的语义对齐的查询,并设计了一种自适应的自指导机制,利用不同语种视角下的单语言检索结果的概率分布来指导跨语言检索。在22对语言组合上进行了广泛的实验来评估所提模型的有效性和效率,结果表明,所提方法的MRR指标达到了当前最先进水平。具体而言,其在高资源语种组合上相较于次优基线的平均MRR提高了15.45%,在低资源语种组合上相较于次优基线提高了18.9%。此外,相比基于大语言模型的方法,该方法在训练时间和推理时间上均更短,并且显著提升了收敛性能。相关代码已公开1)。
Cross-lingual Information Retrieval(CLIR)is an important information acquisition task in natural language proces-sing.Recently,LLM-based retrieval methods have gained attention and demonstrated remarkable progress in this task.However,existing unsupervised retrieval methods based on prompting large language models still insufficient in effectiveness and efficiency.To solve this problem,this paper introduces a novel CLIR method based on aligned query.Specifically,this paper adopts the“pretrain-finetune”paradigm and proposes an adaptive self-teaching encoder based on a pretrained multilingual model to guide cross-lingual retrieval learning by mono-lingual retrieval learning.This method introduces semantically aligned queries in the same language as the documents and designs an adaptive self-teaching mechanism to guide cross-lingual retrieval by leveraging the probability distribution of mono-lingual retrieval results from different linguistic perspectives.To evaluate the effectiveness and efficiency of this method,this paper conducts extensive experiments on 22 language pairs.The results demonstrate that the proposed method achieves SOTA performance in terms of MRR.In particular,this method improves average MRR by 15.45%over the sub-optimal baseline in high-resource language pairs and 18.9%over the sub-optimal baseline in low-resource language pairs.Furthermore,the method reduces training and inference times compared to LLM-based approaches and exhibits faster convergence with enhanced stability.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心