详细信息
Investigation of Model Ensemble for Fine-Grained Air Quality Prediction ( SCI-EXPANDED收录)
文献类型:期刊文献
英文题名:Investigation of Model Ensemble for Fine-Grained Air Quality Prediction
作者:Zheng, Hong[1];Cheng, Yunhui[1];Li, Haibin[1]
机构:[1]East China Univ Sci & Technol, Informat Engn & Comp Sci Coll, Shanghai 200237, Peoples R China
年份:2020
卷号:17
期号:7
起止页码:207
外文期刊名:CHINA COMMUNICATIONS
收录:;WOS:【SCI-EXPANDED(收录号:WOS:000551546900016)】;
基金:We are pleased to acknowledge the National Natural Science Foundation of China under Grant 61103115; the National Natural Science Foundation of China under Grant 61103172; and suggestions.the National Natural Science Youth Foundation of China under Grant 61602175; the special fund for Software and Integrated Circuit Industry Development of Shanghai under Grant 150809; the "Action Plan for Innovation on Science and Technology" Projects of Shanghai (project No: 16511101000).The authors are also grateful to the anonymous referees for their insightful and valuable comments
语种:英文
外文关键词:air quality prediction; machine learning; model ensemble
摘要:Air pollution which is detrimental to people's health is a wide spread problem across many countries around the world. Developing better air quality prediction approaches is an important research issue. Existing methods often focus on the prediction of air pollution concentrations, which is not as intuitive to the public as the air quality levels. In this paper, near future fine-grained air quality level prediction task is explored with a series of machine learning ensemble methods. Included ensemble methods are majority voting, averaging, weighted averaging and 16 different stacking tactics. To investigate the performances of these ensemble methods, comprehensive comparative experiments are conducted. Included contrast models are classical Autoregressive Integrated Moving Average (ARIMA), popular deep learning model Long Short-Term Memory (LSTM) neural network, and nine of the base-level models such as Support Vector Machine (SVM), Random Forest (RF), Logistic Regression (LR) and several boosting models. Datasets acquired from a coastal city Hong Kong and an inland city Beijing are used to train and validate all the models. Experiments show that performances of the ensemble methods outperform most of the individual models, especially when stacking with probability distributions together with engineered original features, which demonstrates the best performance.
参考文献:
正在载入数据...
