International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1701310

1701310PublishedVol 2 · Issue 12

Merging Small Files For Cloud Storage Using Agglomerative Hierarchical Clustering

Htu Ra

Subject area: Science,Engineering and Technology  ·  Area of research: Computer Science

Abstract

Hadoop distributed file system (HDFS) was originally designed for large files. HDFS stores each small file as one separate block although the size of several small files is lesser than the size of block size.Therefore, a large number of blocks are created with massive small files. When the large number of small files is accessed, Name Node often becomes the bottleneck. The problem of storing and accessing large number of small files is named as small file problem. In order to solve this issue in HDFS, an approach of merging small files on HDFS is proposed. In this paper, small files are merged into a larger file based on the agglomeration hierarchical clustering mechanism to reduce Name Node memory consumption. This approach will provide small files for cloud storage.

How to cite this paper

Htu Ra "Merging Small Files For Cloud Storage Using Agglomerative Hierarchical Clustering" Iconic Research And Engineering Journals Volume 2 Issue 12 2019 Page 180-186
Htu Ra "Merging Small Files For Cloud Storage Using Agglomerative Hierarchical Clustering" Iconic Research And Engineering Journals, vol. 2, no. 12, Jun. 2019
Htu Ra (2019). Merging Small Files For Cloud Storage Using Agglomerative Hierarchical Clustering. Iconic Research And Engineering Journals, 2(12).
Htu Ra "Merging Small Files For Cloud Storage Using Agglomerative Hierarchical Clustering" Iconic Research And Engineering Journals, vol. 2, no. 12, Jun. 2019.
@article{1701310,
      author = {Htu Ra},
      title = {Merging Small Files For Cloud Storage Using Agglomerative Hierarchical Clustering},
      journal = {Iconic Research And Engineering Journals},
      year = {2019},
      volume = {2},
      number = {12},
      pages = {180-186},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1701310.pdf},
      abstract = {Hadoop distributed file system (HDFS) was originally designed for large files. HDFS stores each small file as one separate block although the size of several small files is lesser than the size of block size.Therefore, a large number of blocks are created with massive small files. When the large number of small files is accessed, Name Node often becomes the bottleneck. The problem of storing and accessing large number of small files is named as small file problem. In order to solve this issue in HDFS, an approach of merging small files on HDFS is proposed. In this paper, small files are merged into a larger file based on the agglomeration hierarchical clustering mechanism to reduce Name Node memory consumption. This approach will provide small files for cloud storage.},
      month = {June},
  }