详细信息

Structure-Aware Protein Language Model with Multi-Branch Ensemble for Nanobody-Antigen Interaction Prediction  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Structure-Aware Protein Language Model with Multi-Branch Ensemble for Nanobody-Antigen Interaction Prediction

作者:Ying, Fangli[1];Li, Zilong[1];Mao, Junjie[1];Sun, Lihua[1];Phaphuangwittayakul, Aniwat[2];Dhuny, Riyad[3]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, State Key Lab Bioreactor Engn, Shanghai 200237, Peoples R China;[2]Chiang Mai Univ, Int Coll Digital Innovat, Chiang Mai 50200, Thailand;[3]Univ Technol, Dept Creat Arts Film & Media Technol, Pt Louis 11134, Mauritius

年份:2026

卷号:16

期号:10

外文期刊名:APPLIED SCIENCES-BASEL

收录:;EI(收录号:20262220794072);WOS:【SCI-EXPANDED(收录号:WOS:001774204800001)】;

基金:This research was funded by National Major Scientific Instruments and Equipments Development Project of National Natural Science Foundation of China, No. 32327801. This research was partially funded by National Key Research and Development Program of China, No. 2020YFA0907800. This research was partially funded by Research and Development Plan in Shandong Province No. 2022CXGC020206. This research was partially funded by the Key R & D Program of Shandong Province, China, grant number 2022SFGC0104.

语种:英文

外文关键词:Nanobody-Antigen Interaction; Protein Language Models; Protein-Protein Interaction; transfer learning

摘要:Nanobodies have emerged as highly valuable biotherapeutic and diagnostic reagents due to their high specificity, low immunogenicity, and superior tissue penetration. However, traditional nanobody discovery methods rely on camelid immunization and phage display techniques, which are time-consuming and labor-intensive. Meanwhile, existing computational prediction methods for Nanobody-Antigen Interaction (NAI) suffer from several limitations: first, general Protein-Protein Interaction (PPI) models cannot adapt to the specific binding patterns of NAI; second, sequence-based models struggle to capture critical binding features, resulting in unsatisfactory prediction accuracy. To address these challenges, we propose an NAI prediction method based on Protein Language Models (PLMs). Specifically, a structure-aware PLM is first fine-tuned on PPI datasets to learn universal protein binding patterns. This model performs joint encoding of amino acid sequences and protein local structure-related sequences. It can implicitly learn spatial structural priors solely from sequence inputs, which alleviates the limitation of conventional sequence-based models in capturing structural binding characteristics. Subsequently, we use mean, max and min pooling to extract complementary global sequence features that a single pooling method cannot fully capture. We then apply voting fusion to reduce prediction bias and improve model robustness under class imbalance and small-sample scenarios. Evaluated on the NAI benchmark dataset constructed from SAbDab-nano, the proposed model outperforms the best baseline methods in key metrics including Accuracy, Recall, F1-score, AUC-ROC, and AUPR. It exhibits robust performance under class imbalance and small-sample scenarios, validating the effectiveness of the framework.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心