详细信息
A method to improve the performance for storing massive small files in hadoop ( EI收录)
文献类型:会议论文
英文题名:A method to improve the performance for storing massive small files in hadoop
作者:Zheng, Tong[1]; Guo, Weibin[1]; Fan, Guisheng[1]
机构:[1] East China University of Science and Technology, Shanghai, 200237, China
会议论文集:7th International Conference on Computer Engineering and Networks, CENet 2017
会议日期:July 22, 2017 - July 23, 2017
会议地点:Shanghai, China
语种:英文
外文关键词:Mapping - Digital storage - Fault tolerance
摘要:As a new open source project, Hadoop provides a new way to store massive data. Because of high scalability, low cost, good flexibility, high speed and strong fault tolerance performance, it has been widely adopted by the internet companies. However, the performance of Hadoop will reduce significantly once it is used to handle massive small files. As a result, this paper proposes a new scheme to merge small files, which occupymuch memory in NameNode, into large files and establish the mapping relationship between small files and large files, and then store the mapping information in HBase. In order to improve the reading performance, the scheme provides a prefetching mechanism by analyzing the access logs and putting the metadata frequently accessed merge files in the client's memory. The experiment results show that this scheme can efficiently optimize small files storage in HDFS, thus reduce the overload of NameNode and improve the performance of file access. ? Copyright owned by the author(s) under the terms of the Creative Commons.
参考文献:
正在载入数据...
