详细信息

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models

作者:Xiong, Luolin[1,2];Wang, Haofen[3];Chen, Xi[4];Sheng, Lu[5];Xiong, Yun[4];Liu, Jingping[6];Xiao, Yanghua[4];Chen, Huajun[7];Han, Qing-Long[8];Tang, Yang[1,2]

机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Engn Res Ctr Proc Syst Engn, Minist Educ, Shanghai 200237, Peoples R China;[3]Tongji Univ, Coll Design & Innovat, Shanghai 200092, Peoples R China;[4]Fudan Univ, Sch Comp Sci, Shanghai Key Lab Data Sci, Shanghai 200433, Peoples R China;[5]Beihang Univ, Sch Software, Beijing 100191, Peoples R China;[6]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[7]Zhejiang Univ, Coll Comp Sci & Technol, AZFT Joint Lab Knowledge Engine Hangzhou Innovat C, Hangzhou 310058, Peoples R China;[8]Swinburne Univ Technol, Sch Sci Comp & Engn Technol, Melbourne, VIC 3122, Australia

年份:2025

卷号:12

期号:5

起止页码:841

外文期刊名:IEEE-CAA JOURNAL OF AUTOMATICA SINICA

收录:;EI(收录号:20250340395);WOS:【SCI-EXPANDED(收录号:WOS:001490407500010)】;

基金:This work was supported by the National Natural Science Foundation of China (62233005, 62293502, U2441245, 62176185, U23B2057, 62306112), the STCSM Science and Technology Innovation Action Plan Computational Biology Program (24JS2830400), the State Key Laboratory of Industrial Control Technology, China (ICT2024A22), the Shanghai Sailing Program (23YF1409400), and the National Science and Technology Major Project (2024ZD0532403). Recommended by Associate Editor Hui Yu. (L. Xiong, H. Wang, X. Chen, L. Sheng, Y. Xiong, J. Liu, Y. Xiao, and H. Chen contributed equally to this work.

语种:英文

外文关键词:DeepSeek; large AI models; reasoning capability; reinforcement learning; test-time scaling

摘要:DeepSeek, a Chinese artificial intelligence (AI) startup, has released their V3 and R1 series models, which attracted global attention due to their low cost, high performance, and open-source advantages. This paper begins by reviewing the evolution of large AI models focusing on paradigm shifts, the mainstream large language model (LLM) paradigm, and the DeepSeek paradigm. Subsequently, the paper highlights novel algorithms introduced by DeepSeek, including multi-head latent attention (MLA), mixture-of-experts (MoE), multi-token prediction (MTP), and group relative policy optimization (GRPO). The paper then explores DeepSeek's engineering breakthroughs in LLM scaling, training, inference, and system-level optimization architecture. Moreover, the impact of DeepSeek models on the competitive AI landscape is analyzed, comparing them to mainstream LLMs across various fields. Finally, the paper reflects on the insights gained from DeepSeek's innovations and discusses future trends in the technical and engineering development of large AI models, particularly in data, training, and reasoning.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心