Home / Current Issue / Paper 1708527
Scalable Distributed ETL Architecture for Big Data Storage and Processing
Subject area: Science,Engineering and Technology · Area of research: Big Data Analytics
Abstract
The emergence of more massive volumes, faster speed, and an increasingly diverse range of big data has led to the creation of more sophisticated techniques for big data processing and archiving. ETL has evolved from extracted, transformed, load processes performed centrally and restricted in scalability to distributed structures that include cloud computing and big data technology. These modern distributed designs in ETL architecture present the processing of tasks by distributing them across the nodes or clusters, thus promoting scalability, performance, and fault tolerance. Hence, by utilizing parallelism in ETL processes by distributed systems, it is easier and faster to handle large datasets and reconcile time data processing together with data integration and transformation. In this paper, we provide a discussion of one distributed ETL architecture relevant to large-scale data processing. Here, we discuss its key factors of data extraction, data transformation and loading, all of which are intended for distributed high-performance environments. In this paper, through the case study on the described architecture and its performance analysis, we explain how this design contributes to the reduction of processing time and improvement of system scalability. The findings also reveal that distributed ETL frameworks play a crucial function in various applications of big data handling and analysis and prove how effective they can be in meeting the growing need for data management in the current society.
Keywords
Distributed ETL Architecture, Big Data Analytics, Scalable Data Processing, Cloud-Based Data Management, High-Performance Computing, Fault-Tolerant Systems, Parallel Data Integration.
References
[1] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified Data Processing on Large Clusters. Communications of the ACM, 51(1), 107-113.
[2] Zaharia, M., Chowdhury, M., Das, T., Dave, A., Ma, J., McCauley, M., & Stoica, I. (2012). Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing. In Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (pp. 15-28).
[3] Dhamotharan Seenivasan, "ETL vs ELT: Choosing the right approach for your data warehouse", International Journal for Research Trends and Innovation (www.ijrti.org), ISSN:2456-3315, Vol.7, Issue 2, page no.110 - 122, February-2022, https://www.ijrti.org/papers/IJRTI2202018.pdf
[4] AWS Glue. https://aws.amazon.com/glue/
[5] Google Cloud Dataflow. https://cloud.google.com/dataflow
[6] Azure Data Factory. https://azure.microsoft.com/en-us/services/data-factory/
[7] Dhamotharan Seenivasan, "ETL in a World of Unstructured Data: Advanced Techniques for Data Integration",International Journal of Management, IT and Engineering (IJMIE), Vol. 11, Issue 1, January 2021, pp. 127-145, https://www.ijmra.us/2021ijmie_january.php
[8] https://link.springer.com/article/10.1007/s10115-022-01757-7
[9] https://www.astera.com/knowledge-center/scalable-etl-architectures/
[10] https://www.scs.stanford.edu/17au-cs244b/labs/projects/wang.pdf
[11] https://www.rudderstack.com/learn/etl/etl-architecture/
[12] https://medium.com/@diehardankush/etl-pipelines-for-big-data-challenges-and-solutions-dafcfebaf3ff
[13] https://www.integrate.io/blog/big-data-architect-etl/
[14] Dhamotharan Seenivasan, "Optimizing Cloud Data Warehousing: A Deep Dive into Snowflakes Architecture and Performance", International Journal of Advanced Research in Engineering and Technology, 12(3), 2021, pp.951-962, https://iaeme.com/Home/article_id/IJARET_12_03_089
[15] https://www.sprinkledata.com/blogs/the-architecture-of-etl-processes
[16] https://blog.coupler.io/etl-architecture/
[17] https://www.dataversity.net/distributed-data-architecture-patterns-explained/
How to cite this paper
@article{1708527,
author = {Vandana Kollati},
title = {Scalable Distributed ETL Architecture for Big Data Storage and Processing},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {6},
number = {12},
pages = {1605-1612},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1708527.pdf},
abstract = {The emergence of more massive volumes, faster speed, and an increasingly diverse range of big data has led to the creation of more sophisticated techniques for big data processing and archiving. ETL has evolved from extracted, transformed, load processes performed centrally and restricted in scalability to distributed structures that include cloud computing and big data technology. These modern distributed designs in ETL architecture present the processing of tasks by distributing them across the nodes or clusters, thus promoting scalability, performance, and fault tolerance. Hence, by utilizing parallelism in ETL processes by distributed systems, it is easier and faster to handle large datasets and reconcile time data processing together with data integration and transformation. In this paper, we provide a discussion of one distributed ETL architecture relevant to large-scale data processing. Here, we discuss its key factors of data extraction, data transformation and loading, all of which are intended for distributed high-performance environments. In this paper, through the case study on the described architecture and its performance analysis, we explain how this design contributes to the reduction of processing time and improvement of system scalability. The findings also reveal that distributed ETL frameworks play a crucial function in various applications of big data handling and analysis and prove how effective they can be in meeting the growing need for data management in the current society.},
keywords = {Distributed ETL Architecture, Big Data Analytics, Scalable Data Processing, Cloud-Based Data Management, High-Performance Computing, Fault-Tolerant Systems, Parallel Data Integration.},
month = {June},
}