详细信息

Automatic Web Information Extraction Based on Rules  ( CPCI-S收录)  

文献类型:会议论文

英文题名:Automatic Web Information Extraction Based on Rules

作者:Hu, Fanghuai[1];Ruan, Tong[1];Shao, Zhiqing[1];Ding, Jun[1]

机构:[1]E China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China

会议论文集:12th International Conference on Web Information Systems Engineering (WISE 2011)

会议日期:OCT 13-14, 2011

会议地点:Sydney, AUSTRALIA

语种:英文

外文关键词:Web Information Extraction; Probabilistic Model; Rule

摘要:Web Information Extraction is the initial step of effective web mining. In this article a few heuristic rules which describe the characteristics of the main content of web pages are summarized. The rules are constructed by some pre-defined terms and metrics, which can be considered as reusable and extensible for different kinds of HTML pages. Afterwards, a probabilistic model which utilizes the rules and metrics is suggested and the corresponding algorithm is implemented. The algorithm is tested on 1000 randomly selected web pages. The experiment shows that the algorithm is more precise and more applicable to the diverse structure of different web sites than other algorithms.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心