Home / Current Issue / Paper 1707514
Ensuring Data Quality and Integrity in Machine Learning Pipelines: Strategies for Data Engineers
Subject area: Science,Engineering and Technology · Area of research: Data engineering and machine learning
Abstract
Data quality and integrity are critical factors in ensuring the reliability and accuracy of machine learning (ML) models. Poor data quality?caused by missing values, inconsistencies, duplicate records, and biases?can lead to inaccurate predictions and unreliable insights. This paper explores key strategies that data engineers can implement to enhance data quality in ML pipelines. It covers data validation, data cleaning, automated anomaly detection, schema enforcement, and data governance frameworks. Additionally, it examines modern tools and frameworks, such as Great Expectations, TensorFlow Data Validation (TFDV), and Apache Deequ, which assist in maintaining high data integrity. The paper also highlights best practices for designing scalable and automated data quality monitoring systems to support real-time and batch ML workflows. By implementing these strategies, data engineers can ensure that ML models are trained on high-quality, trustworthy data, leading to more accurate and fair outcomes.
How to cite this paper
@article{1707514,
author = { Bhanu Prakash Reddy Rella},
title = {Ensuring Data Quality and Integrity in Machine Learning Pipelines: Strategies for Data Engineers},
journal = {Iconic Research And Engineering Journals},
year = {2022},
volume = {6},
number = {2},
pages = {331-339},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1707514.pdf},
abstract = {Data quality and integrity are critical factors in ensuring the reliability and accuracy of machine learning (ML) models. Poor data quality?caused by missing values, inconsistencies, duplicate records, and biases?can lead to inaccurate predictions and unreliable insights. This paper explores key strategies that data engineers can implement to enhance data quality in ML pipelines. It covers data validation, data cleaning, automated anomaly detection, schema enforcement, and data governance frameworks. Additionally, it examines modern tools and frameworks, such as Great Expectations, TensorFlow Data Validation (TFDV), and Apache Deequ, which assist in maintaining high data integrity. The paper also highlights best practices for designing scalable and automated data quality monitoring systems to support real-time and batch ML workflows. By implementing these strategies, data engineers can ensure that ML models are trained on high-quality, trustworthy data, leading to more accurate and fair outcomes.},
month = {August},
}