详细信息

Investigation of Model Ensemble for Fine-Grained Air Quality Prediction  ( SCI-EXPANDED收录)  

文献类型:期刊文献

英文题名:Investigation of Model Ensemble for Fine-Grained Air Quality Prediction

作者:Zheng, Hong[1];Cheng, Yunhui[1];Li, Haibin[1]

机构:[1]East China Univ Sci & Technol, Informat Engn & Comp Sci Coll, Shanghai 200237, Peoples R China

年份:2020

卷号:17

期号:7

起止页码:207

外文期刊名:CHINA COMMUNICATIONS

收录:;WOS:【SCI-EXPANDED(收录号:WOS:000551546900016)】;

基金:We are pleased to acknowledge the National Natural Science Foundation of China under Grant 61103115; the National Natural Science Foundation of China under Grant 61103172; and suggestions.the National Natural Science Youth Foundation of China under Grant 61602175; the special fund for Software and Integrated Circuit Industry Development of Shanghai under Grant 150809; the "Action Plan for Innovation on Science and Technology" Projects of Shanghai (project No: 16511101000).The authors are also grateful to the anonymous referees for their insightful and valuable comments

语种:英文

外文关键词:air quality prediction; machine learning; model ensemble

摘要:Air pollution which is detrimental to people's health is a wide spread problem across many countries around the world. Developing better air quality prediction approaches is an important research issue. Existing methods often focus on the prediction of air pollution concentrations, which is not as intuitive to the public as the air quality levels. In this paper, near future fine-grained air quality level prediction task is explored with a series of machine learning ensemble methods. Included ensemble methods are majority voting, averaging, weighted averaging and 16 different stacking tactics. To investigate the performances of these ensemble methods, comprehensive comparative experiments are conducted. Included contrast models are classical Autoregressive Integrated Moving Average (ARIMA), popular deep learning model Long Short-Term Memory (LSTM) neural network, and nine of the base-level models such as Support Vector Machine (SVM), Random Forest (RF), Logistic Regression (LR) and several boosting models. Datasets acquired from a coastal city Hong Kong and an inland city Beijing are used to train and validate all the models. Experiments show that performances of the ensemble methods outperform most of the individual models, especially when stacking with probability distributions together with engineered original features, which demonstrates the best performance.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心