Home / Current Issue / Paper 1705069
Automating ETL Workflows with CI/CD Pipelines for Machine Learning Applications
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
Abstract
In today's fast-paced data-driven landscape, automating Extract, Transform, Load (ETL) workflows is crucial for enhancing the efficiency of Machine Learning (ML) applications. This research explores how Continuous Integration and Continuous Deployment (CI/CD) pipelines can automate and streamline ETL processes, reducing the time and manual intervention required for data preparation and deployment. Integrating CI/CD pipelines into ETL workflows ensures that the entire data lifecycle?from extraction, transformation, loading, to model training and deployment?operates seamlessly. The automation of these processes enables rapid iteration, minimizes errors, and accelerates time-to-market for ML models. This paper investigates key strategies for leveraging modern tools such as Jenkins, Apache Airflow, and Kubernetes to build scalable and efficient automated workflows. It also examines real-world case studies where CI/CD pipelines have optimized ML workflows, leading to enhanced productivity, accuracy, and cost savings. By adopting such automation techniques, organizations can better manage large-scale data pipelines, ensure model accuracy, and reduce operational complexities in machine learning projects.
Keywords
Automated ETL workflows, CI/CD pipelines, Machine Learning applications, data lifecycle automation, Jenkins, Apache Airflow, Kubernetes, model deployment.
How to cite this paper
@article{1705069,
author = {Antony Satya Vivek Vardhan Akisetty, Ashish Kumar, Murali Mohana Krishna Dandu, Prof. (Dr) Punit Goel, Prof. (Dr.) Arpit Jain; Er. Aman Shrivastav},
title = {Automating ETL Workflows with CI/CD Pipelines for Machine Learning Applications},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {7},
number = {3},
pages = {478-497},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1705069.pdf},
abstract = {In today's fast-paced data-driven landscape, automating Extract, Transform, Load (ETL) workflows is crucial for enhancing the efficiency of Machine Learning (ML) applications. This research explores how Continuous Integration and Continuous Deployment (CI/CD) pipelines can automate and streamline ETL processes, reducing the time and manual intervention required for data preparation and deployment. Integrating CI/CD pipelines into ETL workflows ensures that the entire data lifecycle?from extraction, transformation, loading, to model training and deployment?operates seamlessly. The automation of these processes enables rapid iteration, minimizes errors, and accelerates time-to-market for ML models. This paper investigates key strategies for leveraging modern tools such as Jenkins, Apache Airflow, and Kubernetes to build scalable and efficient automated workflows. It also examines real-world case studies where CI/CD pipelines have optimized ML workflows, leading to enhanced productivity, accuracy, and cost savings. By adopting such automation techniques, organizations can better manage large-scale data pipelines, ensure model accuracy, and reduce operational complexities in machine learning projects.},
keywords = {Automated ETL workflows, CI/CD pipelines, Machine Learning applications, data lifecycle automation, Jenkins, Apache Airflow, Kubernetes, model deployment.},
month = {September},
}