International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1707468

1707468 Vol 6 · Issue 9 Download Paper

Scalable Metadata Management in Data Lakes Using Machine Learning

Shishir Tewari

Subject area: Science,Engineering and Technology  ·  Area of research: Artificial Intelligence and machine learning

Abstract

The quick growth of big data triggered big data lakes to become the scalable storage choice for massive handled and raw data. The preservation of effective metadata management in data lakes remains a major challenge because of inconsistencies that affect metadata together with retrieval difficulties and scalability problems. Manual tagging methods along with rule-based approaches struggle to manage rising data volumes so they produce governance problems and make data discovery difficult. Machine learning provides an effective solution to these challenges through automated processes of metadata extraction as well as metadata classification and retrieval. Numerous machine learning models provide solutions to improve scalable metadata management of data lake configurations. Metadata tagging effectiveness stands to benefit from supervised learning whereas unsupervised learning demonstrates value for pattern detection in metadata. Deep learning models which implement NLP techniques help organizations improve semantic metadata processing for data classification and retrieval purposes. Data management benefits from reinforcement learning approaches which make continuous user interaction observations to refine search efficiency as a result. The evaluation process for machine learning in metadata management utilizes a case study analysis between conventional systems and smart learning systems. The evaluation shows that better metadata accuracy and faster retrieval as well as improved scalability now exists. Through this research organizations can learn how to employ artificial intelligence technology for smarter metadata system development that leads to improved data lake governance and accessibility together with better decision capabilities

Keywords

Scalable Metadata Management, Machine Learning for Data Lakes, Automated Metadata Tagging, Big Data Governance, Metadata Optimization in Large-Scale Systems

References

[1] Avilés-González, A., Piernas, J., & González-Férez, P. (2014). Scalable metadata management through OSD+ devices. International Journal of Parallel Programming, 42(1), 4-29. https://doi.org/10.1007/s10766-012-0207-8

[2] Al-Badi, A., Tarhini, A., & Khan, A. I. (2018). Exploring big data governance frameworks. Procedia computer science, 141, 271-277.https://doi.org/10.1016/j.procs.2018.10.181

[3] Bhattacharya, A. A., Hong, D., Culler, D., Ortiz, J., Whitehouse, K., & Wu, E. (2015, November). Automated metadata construction to support portable building applications. In Proceedings of the 2nd ACM International Conference on Embedded Systems for Energy-Efficient Built Environments (pp. 3-12). https://doi.org/10.1145/2821650.2821667

[4] Balaji, B., Bhattacharya, A., Fierro, G., Gao, J., Gluck, J., Hong, D., ... & Whitehouse, K. (2016, November). Brick: Towards a unified metadata schema for buildings. In Proceedings of the 3rd ACM International Conference on Systems for Energy-Efficient Built Environments (pp. 41-50). https://doi.org/10.1145/2993422.2993577

[5] Blanas, S., & Byna, S. (2015). Towards exascale scientific metadata management. arXiv preprint arXiv:1503.08482.

[6] Bruno, N., Jain, S., & Zhou, J. (2013). Continuous cloud-scale query optimization and processing. Proceedings of the VLDB Endowment, 6(11), 961-972.https://doi.org/10.14778/2536222.2536223

[7] Bogatu, A., Fernandes, A. A., Paton, N. W., & Konstantinou, N. (2020, April). Dataset discovery in data lakes. In 2020 ieee 36th international conference on data engineering (icde) (pp. 709-720). IEEE.https://doi.org/10.1109/ICDE48307.2020.00067

[8] Cha, M. H., Lee, S. M., Kim, H. Y., & Kim, Y. K. (2019). Effective metadata management in exascale file system. The Journal of Supercomputing, 75, 7665-7689.https://doi.org/10.1007/s11227-019-02974-8

[9] Castro, A., Villagra, V. A., Garcia, P., Rivera, D., & Toledo, D. (2021). An ontological-based model to data governance for big data. IEEE Access, 9, 109943-109959.https://doi.org/10.1109/ACCESS.2021.3101938

[10] Chen, Y., Li, C., Lv, M., Shao, X., Li, Y., & Xu, Y. (2019). Explicit data correlations-directed metadata prefetching method in distributed file systems. IEEE Transactions on Parallel and Distributed Systems, 30(12), 2692-2705.https://doi.org/10.1109/TPDS.2019.2921760

[11] Dai, H., Wang, Y., Kent, K. B., Zeng, L., & Xu, C. (2022). The state of the art of metadata managements in large-scale distributed file systems—scalability, performance and availability. IEEE Transactions on Parallel and Distributed Systems, 33(12), 3850-3869.https://doi.org/10.1109/TPDS.2022.3170574

[12] Gao, J., Ploennigs, J., & Berges, M. (2015, November). A data-driven meta-data inference framework for building automation systems. In Proceedings of the 2nd ACM International Conference on Embedded Systems for Energy-Efficient Built Environments (pp. 23-32). https://doi.org/10.1145/2821650.2821670

[13] Hua, Y., Zhu, Y., Jiang, H., Feng, D., & Tian, L. (2010). Supporting scalable and adaptive metadata management in ultralarge-scale file systems. IEEE Transactions on Parallel and Distributed Systems, 22(4), 580-593.https://doi.org/10.1109/TPDS.2010.116

[14] Hua, Y., Jiang, H., Zhu, Y., Feng, D., & Tian, L. (2011). Semantic-aware metadata organization paradigm in next-generation file systems. IEEE Transactions on Parallel and Distributed Systems, 23(2), 337-344.https://doi.org/10.1109/TPDS.2011.169

[15] Janssen, M., Brous, P., Estevez, E., Barbosa, L. S., & Janowski, T. (2020). Data governance: Organizing data for trustworthy Artificial Intelligence. Government information quarterly, 37(3), 101493.

[16] Jiang, L., Li, B., & Song, M. (2010, October). THE optimization of HDFS based on small files. In 2010 3Rd IEEE international conference on broadband network and multimedia technology (IC-BNMT) (pp. 912-915). IEEE. https://doi.org/10.1109/ICBNMT.2010.5705223

[17] Khine, P. P., & Wang, Z. S. (2018). Data lake: a new ideology in big data era. In ITM web of conferences (Vol. 17, p. 03025). EDP Sciences. https://doi.org/10.1051/itmconf/20181703025

[18] Kim, H. Y., & Cho, J. S. (2017, June). Data governance framework for big data implementation with a case of Korea. In 2017 IEEE International Congress on Big Data (BigData Congress) (pp. 384-391). IEEE. https://doi.org/10.1109/BigDataCongress.2017.56

[19] Lawson, M., & Lofstead, J. (2018, November). Using a robust metadata management system to accelerate scientific discovery at extreme scales. In 2018 IEEE/ACM 3rd International Workshop on Parallel Data Storage & Data Intensive Scalable Computing Systems (PDSW-DISCS) (pp. 13-23). IEEE.https://doi.org/10.1109/PDSW-DISCS.2018.00004

[20] Miloslavskaya, N., & Tolstoy, A. (2016). Big data, fast data and data lake concepts. Procedia Computer Science, 88, 300-305.https://doi.org/10.1109/ICDE51399.2021.00046

[21] Miloslavskaya, N., & Tolstoy, A. (2016). Big data, fast data and data lake concepts. Procedia Computer Science, 88, 300-305.https://doi.org/10.1016/j.procs.2016.07.439

[22] Mehmood, H., Gilman, E., Cortes, M., Kostakos, P., Byrne, A., Valta, K., ... & Riekki, J. (2019, April). Implementing big data lake for heterogeneous data sources. In 2019 ieee 35th international conference on data engineering workshops (icdew) (pp. 37-44). IEEE.https://doi.org/10.1109/ICDEW.2019.00-37

[23] Nambiar, A., & Mundra, D. (2022). An overview of data warehouse and data lake in modern enterprise data management. Big data and cognitive computing, 6(4), 132.https://doi.org/10.3390/bdcc6040132

[24] Niazi, S., Ismail, M., Haridi, S., Dowling, J., Grohsschmiedt, S., & Ronström, M. (2017). {HopsFS}: Scaling hierarchical file system metadata using {NewSQL} databases. In 15th USENIX Conference on File and Storage Technologies (FAST 17) (pp. 89-104).

[25] Neumaier, S., Umbrich, J., & Polleres, A. (2016). Automated quality assessment of metadata across open data portals. Journal of Data and Information Quality (JDIQ), 8(1), 1-29. https://doi.org/10.1145/2964909

[26] O'Leary, D. E. (2014). Embedding AI and crowdsourcing in the big data lake. IEEE Intelligent Systems, 29(5), 70-73. https://doi.org/10.1109/MIS.2014.82

[27] Pallickara, S. L., Pallickara, S., Zupanski, M., & Sullivan, S. (2010, November). Efficient metadata generation to enable interactive data discovery over large-scale scientific data collections. In 2010 IEEE Second International Conference on Cloud Computing Technology and Science (pp. 573-580). IEEE. https://doi.org/10.1109/CloudCom.2010.99

[28] Riley, J. (2017). Understanding metadata. Washington DC, United States: National Information Standards Organization (http://www. niso. org/publications/press/UnderstandingMetadata. pdf), 23, 7-10.

[29] Rupprecht, L., Zhang, R., Owen, B., Pietzuch, P., & Hildebrand, D. (2017, April). SwiftAnalytics: Optimizing object storage for big data analytics. In 2017 IEEE International Conference on Cloud Engineering (IC2E) (pp. 245-251). IEEE. https://doi.org/10.1109/IC2E.2017.19

[30] Ren, K., Zheng, Q., Patil, S., & Gibson, G. (2014, November). IndexFS: Scaling file system metadata performance with stateless caching and bulk insertion. In SC'14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 237-248). IEEE. https://doi.org/10.1109/SC.2014.25

[31] Ravat, F., & Zhao, Y. (2019). Data lakes: Trends and perspectives. In Database and Expert Systems Applications: 30th International Conference, DEXA 2019, Linz, Austria, August 26–29, 2019, Proceedings, Part I 30 (pp. 304-313). Springer International Publishing. https://doi.org/10.1007/978-3-030-27615-7_23

[32] Singh, H. J., & Bawa, S. (2018). Scalable metadata management techniques for ultra-large distributed storage systems--A systematic review. ACM Computing Surveys (CSUR), 51(4), 1-37. https://doi.org/10.1145/3212686

[33] Sleimi, A., Sannier, N., Sabetzadeh, M., Briand, L., Ceci, M., & Dann, J. (2021). An automated framework for the extraction of semantic legal metadata from legal texts. Empirical Software Engineering, 26, 1-50. https://doi.org/10.1007/s10664-020-09933-5

[34] Schelter, S., Boese, J. H., Kirschnick, J., Klein, T., & Seufert, S. (2017). Automatically tracking metadata and provenance of machine learning experiments.

[35] Sleimi, A., Sannier, N., Sabetzadeh, M., Briand, L., & Dann, J. (2018, August). Automated extraction of semantic legal metadata using natural language processing. In 2018 IEEE 26th International Requirements Engineering Conference (RE) (pp. 124-135). IEEE. https://doi.org/10.1109/RE.2018.00022

[36] Tang, H., Byna, S., Dong, B., Liu, J., & Koziol, Q. (2017, September). Someta: Scalable object-centric metadata management for high performance computing. In 2017 IEEE International Conference on Cluster Computing (CLUSTER) (pp. 359-369). IEEE. https://doi.org/10.1109/CLUSTER.2017.53

[37] Tang, H., Byna, S., Dong, B., Liu, J., & Koziol, Q. (2017, September). Someta: Scalable object-centric metadata management for high performance computing. In 2017 IEEE International Conference on Cluster Computing (CLUSTER) (pp. 359-369). IEEE. https://doi.org/10.1109/CLUSTER.2017.53

[38] Tallon, P. P. (2013). Corporate governance of big data: Perspectives on value, risk, and cost. Computer, 46(6), 32-38.https://doi.org/10.1109/MC.2013.155

[39] Tuarob, S., Pouchard, L. C., & Giles, C. L. (2013, July). Automatic tag recommendation for metadata annotation using probabilistic topic modeling. In Proceedings of the 13th ACM/IEEE-CS joint conference on Digital libraries (pp. 239-248).https://doi.org/10.1145/2467696.2467706

[40] Trom, L., & Cronje, J. (2020). Analysis of data governance implications on big data. In Advances in Information and Communication: Proceedings of the 2019 Future of Information and Communication Conference (FICC), Volume 1 (pp. 645-654). Springer International Publishing.https://doi.org/10.1007/978-3-030-12388-8_45

[41] Tse, D., Chow, C. K., Ly, T. P., Tong, C. Y., & Tam, K. W. (2018, August). The challenges of big data governance in healthcare. In 2018 17th IEEE International Conference On Trust, Security And Privacy In Computing And Communications/12th IEEE International Conference On Big Data Science And Engineering (TrustCom/BigDataSE) (pp. 1632-1636). IEEE.https://doi.org/10.1109/TrustCom/BigDataSE.2018.00240

[42] Thomson, A., & Abadi, D. J. (2015). {CalvinFS}: Consistent {WAN} Replication and Scalable Metadata Management for Distributed File Systems. In 13th USENIX Conference on File and Storage Technologies (FAST 15) (pp. 1-14).

[43] Wimmer, J., Towsey, M., Planitz, B., Williamson, I., & Roe, P. (2013). Analysing environmental acoustic data through collaboration and automation. Future Generation Computer Systems, 29(2), 560-568.https://doi.org/10.1016/j.future.2012.03.004

[44] Winter, J. S., & Davidson, E. (2019). Big data governance of personal health information and challenges to contextual integrity. The Information Society, 35(1), 36-51. https://doi.org/10.1080/01972243.2018.1542648

[45] Xu, Q., Arumugam, R. V., Yong, K. L., & Mahadevan, S. (2013). Efficient and scalable metadata management in EB-scale file systems. IEEE Transactions on Parallel and Distributed Systems, 25(11), 2840-2850.https://doi.org/10.1109/TPDS.2013.293

[46] Xiong, J., Hu, Y., Li, G., Tang, R., & Fan, Z. (2010). Metadata distribution and consistency techniques for large-scale cluster file systems. IEEE Transactions on Parallel and Distributed Systems, 22(5), 803-816.https://doi.org/10.1109/TPDS.2010.154

[47] Zhang, Y., & Ives, Z. G. (2020, June). Finding related tables in data lakes for interactive data science. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data (pp. 1951-1966). https://doi.org/10.1145/3318464.3389726

[48] Zhu, S., Hrnjica, B., Ptak, M., Choiński, A., & Sivakumar, B. (2020). Forecasting of water level in multiple temperate lakes using machine learning models. Journal of Hydrology, 585, 124819.https://doi.org/10.1016/j.jhydrol.2020.124819

[49] Zhu, C. (2019). Big data as a governance mechanism

[50] The Review of Financial Studies, 32(5), 2021-2061.

How to cite this paper

Shishir Tewari "Scalable Metadata Management in Data Lakes Using Machine Learning" Iconic Research And Engineering Journals Volume 6 Issue 9 2023 Page 425-441
Shishir Tewari "Scalable Metadata Management in Data Lakes Using Machine Learning" Iconic Research And Engineering Journals, vol. 6, no. 9, Mar. 2023
Shishir Tewari (2023). Scalable Metadata Management in Data Lakes Using Machine Learning. Iconic Research And Engineering Journals, 6(9).
Shishir Tewari "Scalable Metadata Management in Data Lakes Using Machine Learning" Iconic Research And Engineering Journals, vol. 6, no. 9, Mar. 2023.
@article{1707468,
      author = {Shishir Tewari},
      title = {Scalable Metadata Management in Data Lakes Using Machine Learning},
      journal = {Iconic Research And Engineering Journals},
      year = {2023},
      volume = {6},
      number = {9},
      pages = {425-441},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1707468.pdf},
      abstract = {The quick growth of big data triggered big data lakes to become the scalable storage choice for massive handled and raw data. The preservation of effective metadata management in data lakes remains a major challenge because of inconsistencies that affect metadata together with retrieval difficulties and scalability problems. Manual tagging methods along with rule-based approaches struggle to manage rising data volumes so they produce governance problems and make data discovery difficult. Machine learning provides an effective solution to these challenges through automated processes of metadata extraction as well as metadata classification and retrieval. Numerous machine learning models provide solutions to improve scalable metadata management of data lake configurations. Metadata tagging effectiveness stands to benefit from supervised learning whereas unsupervised learning demonstrates value for pattern detection in metadata. Deep learning models which implement NLP techniques help organizations improve semantic metadata processing for data classification and retrieval purposes. Data management benefits from reinforcement learning approaches which make continuous user interaction observations to refine search efficiency as a result. The evaluation process for machine learning in metadata management utilizes a case study analysis between conventional systems and smart learning systems. The evaluation shows that better metadata accuracy and faster retrieval as well as improved scalability now exists. Through this research organizations can learn how to employ artificial intelligence technology for smarter metadata system development that leads to improved data lake governance and accessibility together with better decision capabilities},
      keywords = {Scalable Metadata Management, Machine Learning for Data Lakes, Automated Metadata Tagging, Big Data Governance, Metadata Optimization in Large-Scale Systems},
      month = {March},
  }