Home / Current Issue / Paper 1716725
Automated Data Quality Scoring for Analytics Readiness Using Integrated Profiling and Validation Frameworks
Subject area: Science,Engineering and Technology · Area of research: Analytics Readiness
Abstract
The quality of datasets is a critical factor in identifying the stability and accuracy of the present-day analytics systems. In most data-intensive contexts, data quality, including missing values, data duplication, inconsistency in the schema, and drift in distribution, can cause a substantial impact on the results of analytical processes and result in unreliable insights. Initial data quality assessment has pointed to the necessity of systematic mechanisms in measuring the reliability, completeness, and consistency of datasets prior to their application in analytical decision-making and the significance of structured data quality assessment mechanisms in complex information systems (Christopher S. Carson, 2001). In spite of these developments, a number of existing analytics pipelines do not have automated processes that can measure the reliability of datasets in a single and scalable way. The current paper will present an automated framework of dataset trust scoring aimed at assessing analytics readiness by fusing messages of data profiling with rule-based validation procedures. The suggested solution consists of the combination of indicators of profiling, which include missingness, duplicate records, distribution drift, and invalid categorical values, and validation rules implemented by automated data integrity checks. The emergence of recent automated data profiling and quality scoring tools has confirmed the efficacy of the algorithmic assessment techniques in identifying data anomalies and enhancing the predictive analytics integrity (Hugo Moura et al., 2024). Based on such improvements, the suggested framework consists of a structured scoring model that combines various data quality indicators into a single dataset trust score. The framework also assesses the possibility of using dataset trust scores as predictors of downstream analytics stability and errors in analytical processes. The adaptive data quality scoring models have demonstrated their potential in industrial contexts in which drift-sensitive monitoring systems are utilized to ensure the stability of data-driven systems in the long term (Fatih Bayram et al., 2024). The proposed framework builds a scalable and automated data readiness evaluation system before the analytical processing by extrapolating these concepts. The research also adds to the expanding area of automated data governance by introducing an effective model of continuous quality assessment of the dataset, which would help organizations increase the reliability of analytics, minimize the spread of errors, and increase the confidence in the decision-making systems it is based on.
Keywords
Automated Data Quality Scoring; Dataset Trust Score; Data Profiling and Validation; Analytics Readiness Assessment; Data Drift and Integrity Monitoring; Data Quality Automation; AI-Driven Data Governance.
References
[1] (2025). AIAugmented Frameworks for Data Quality Validation: Integrating RuleBased Engines, Semantic Deduplication, and Governance Tools for Robust LargeScale Data Pipelines. International Journal of Advanced Artificial Intelligence Research, 2(08), 9–15. https://aimjournals.com/index.php/ijaair/article/view/382
[2] Bauer, J. C., John, E., Wood, C. L., Plass, D., & Richardson, D. (2020). Data Entry Automation Improves Cost, Quality, Performance, and Job Satisfaction in a Hospital Nursing Unit. Journal of Nursing Administration, 50(1), 34–39. https://doi.org/10.1097/NNA.0000000000000836
[3] Bayram, F., Ahmed, B. S., & Hallin, E. (2024). Adaptive Data Quality Scoring Operations Framework Using DriftAware Mechanism for Industrial Applications. Journal of Systems and Software, 217. https://doi.org/10.1016/j.jss.2024.112184
[4] Bevilacqua, M., Oketch, K., Qin, R., Stamey, W., Zhang, X., Gan, Y., … Abbasi, A. (2025). When Automated Assessment Meets Automated Content Generation: Examining Text Quality in the Era of GPTs. ACM Transactions on Information Systems, 43(2). https://doi.org/10.1145/3702639
[5] Buchanan, E. M., & Scofield, J. E. (2018). Methods to detect low quality data and its implication for psychological research. Behavior Research Methods, 50(6), 2586–2596. https://doi.org/10.3758/s13428-018-1035-6
[6] Bui, N. M., & Barrot, J. S. (2025). ChatGPT as an automated essay scoring tool in the writing classrooms: how it compares with human scoring. Education and Information Technologies, 30(2), 2041–2058. https://doi.org/10.1007/s10639-024-12891-w
[7] Carson, C. S. (2001).Toward a framework for assessing data quality. IMF Working Paper.https://doi.org/10.5089/9781451844269.001
[8] Devi, C., Inampudi, R. K., & Vijayaboopathy, V. (2025). Federated DataMesh Quality Scoring with Great Expectations and Apache Atlas Lineage. Journal of Knowledge Learning and Science Technology, 4(2), 92–101. https://doi.org/10.60087/jklst.v4.n2.008
[9] Doris, L., & Potter, K. (2024). Continuous Monitoring and Improvement: Implement continuous monitoring of AI models to detect and correct issues in real-time. I-Manager s Journal on Artificial Intelligence & Machine Learning.
[10] Elragal, A., & Elgendy, N. (2024). A data-driven decision-making readiness assessment model: The case of a Swedish food manufacturer. Decision Analytics Journal, 10. https://doi.org/10.1016/j.dajour.2024.100405
[11] Ezerins, M. E., Ludwig, T. D., O’Neil, T., Foreman, A. M., & Açıkgöz, Y. (2022). Advancing safety analytics: A diagnostic framework for assessing system readiness within occupational safety and health. Safety Science, 146. https://doi.org/10.1016/j.ssci.2021.105569
[12] Fariha, A., Tiwari, A., Radhakrishna, A., Gulwani, S., & Meliou, A. (2021). Conformance Constraint Discovery: Measuring Trust in DataDriven Systems. Proceedings of the ACM SIGMOD International Conference on Management of Data. https://doi.org/10.1145/3448016.3452795
[13] Hasnain, M., Pasha, M. F., Ghani, I., Imran, M., Alzahrani, M. Y., & Budiarto, R. (2020). Evaluating Trust Prediction and Confusion Matrix Measures for Web Services Ranking. IEEE Access, 8, 90847–90861. https://doi.org/10.1109/ACCESS.2020.2994222
[14] Helbig, C., et al. (2019).Data quality assessment framework for critical raw materials. Resources, Conservation and Recycling.https://doi.org/10.1016/j.resconrec.2019.104564
[15] Kahn, M. G., et al. (2021).Facilitating harmonized data quality assessments: A data quality framework for observational research data collections.https://doi.org/10.1186/s12874-021-01252-7
[16] Luo, F., Ge, N., & Xu, J. (2023). Power Supply Reliability Analysis of Distribution Systems Considering Data Transmission Quality of Distribution Automation Terminals. Energies, 16(23). https://doi.org/10.3390/en16237826
[17] Martins, P., Cardoso, F., Váz, P., Silva, J., & Abbasi, M. (2025). Performance and Scalability of Data Cleaning and Preprocessing Tools: A Benchmark on Large RealWorld Datasets. Data, 10(5), 68. https://doi.org/10.3390/data10050068
[18] Moura, H., et al. (2024). Automated Data Profiling and Scoring Methods for Predictive Analytics. Information Processing & Management. https://doi.org/10.1016/j.ipm.2024.103903
[19] Nakamoto, R., Flanagan, B., Yamauchi, T., Dai, Y., Takami, K., & Ogata, H. (2023). Enhancing Automated Scoring of Math Self-Explanation Quality Using LLM-Generated Datasets: A Semi-Supervised Approach. Computers, 12(11). https://doi.org/10.3390/computers12110217
[20] Nalla, S. M. R. (2025). Building an AI Trust Score: A DataDriven Framework to Evaluate Dataset Fitness. International Journal of Computing and Engineering, 7(20), 54–63. https://doi.org/10.47941/ijce.3091
[21] Nicholson, N., Carvalho, R. N., & Štotl, I. (2025).A FAIR perspective on data quality frameworks. Data.https://doi.org/10.3390/data10090136
[22] Ozonze, O., Scott, P. J., & Hopgood, A. A. (2023, December 1). Automating Electronic Health Record Data Quality Assessment. Journal of Medical Systems. Springer. https://doi.org/10.1007/s10916-022-01892-2
[23] Prasad, N., & Paripati, L. K. (2025). AI-Driven Data Governance Framework For Cloud-Based Data Analytics. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.5052472
[24] Razali, A. (2024). Improving Data Reliability Assessment in ETL Processes Through Quality Scoring Techniques in Data Analytics. International Journal on Informatics Visualization. https://doi.org/10.1109/icst50505.2020.9732870
[25] Saini, H., Singh, G., Dalal, S., Moorthi, I., Aldossary, S. M., Nuristani, N., & Hashmi, A. (2024). A hybrid machine learning model with self-improved optimization algorithm for trust and privacy preservation in cloud environment. Journal of Cloud Computing, 13(1). https://doi.org/10.1186/s13677-024-00717-6
[26] Sarr, D. (2024).Towards explainable automated data quality enhancement without domain knowledge.https://doi.org/10.48550/arXiv.2409.10139Taleb, I., Serhani, M. A., & Bouhaddioui, C. (2021). Big Data Quality Framework: A Holistic Approach to Continuous Quality Management. Journal of Big Data, 8(76). https://doi.org/10.1186/s40537-021-00468-0
[27] Tariq, A., et al. (2025).A survey of data quality measurement and monitoring tools. Frontiers in Artificial Intelligence.https://doi.org/10.3389/frai.2025.1621514
[28] Tenneti, K. B., Pandula, S., & Pandula, S. (2024). Comparative Analysis of Traditional and AI-Driven Data Governance: A Systematic Review and Future Directions in IT. International Journal of Computer Trends and Technology, 72(11), 150–158. https://doi.org/10.14445/22312803/ijctt-v72i11p116
[29] Tute, E., Ganapathy, N., & Wulff, A. (2021). A DataDriven Learning Approach for the Assessment of Data Quality. BMC Medical Informatics and Decision Making, 21, 302. https://doi.org/10.1186/s12911-021-01656-x
[30] Venkatraman, S., & Sundarraj, R. (2023). Assessing organizational health-analytics readiness: artifacts based on elaborated action design method. Journal of Enterprise Information Management, 36(1), 123–150. https://doi.org/10.1108/JEIM-10-2020-0422
How to cite this paper
@article{1716725,
author = {Sai Lalitesh Pothukuchi},
title = {Automated Data Quality Scoring for Analytics Readiness Using Integrated Profiling and Validation Frameworks},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {7},
number = {6},
pages = {622-636},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1716725.pdf},
abstract = {The quality of datasets is a critical factor in identifying the stability and accuracy of the present-day analytics systems. In most data-intensive contexts, data quality, including missing values, data duplication, inconsistency in the schema, and drift in distribution, can cause a substantial impact on the results of analytical processes and result in unreliable insights. Initial data quality assessment has pointed to the necessity of systematic mechanisms in measuring the reliability, completeness, and consistency of datasets prior to their application in analytical decision-making and the significance of structured data quality assessment mechanisms in complex information systems (Christopher S. Carson, 2001). In spite of these developments, a number of existing analytics pipelines do not have automated processes that can measure the reliability of datasets in a single and scalable way. The current paper will present an automated framework of dataset trust scoring aimed at assessing analytics readiness by fusing messages of data profiling with rule-based validation procedures. The suggested solution consists of the combination of indicators of profiling, which include missingness, duplicate records, distribution drift, and invalid categorical values, and validation rules implemented by automated data integrity checks. The emergence of recent automated data profiling and quality scoring tools has confirmed the efficacy of the algorithmic assessment techniques in identifying data anomalies and enhancing the predictive analytics integrity (Hugo Moura et al., 2024). Based on such improvements, the suggested framework consists of a structured scoring model that combines various data quality indicators into a single dataset trust score. The framework also assesses the possibility of using dataset trust scores as predictors of downstream analytics stability and errors in analytical processes. The adaptive data quality scoring models have demonstrated their potential in industrial contexts in which drift-sensitive monitoring systems are utilized to ensure the stability of data-driven systems in the long term (Fatih Bayram et al., 2024). The proposed framework builds a scalable and automated data readiness evaluation system before the analytical processing by extrapolating these concepts. The research also adds to the expanding area of automated data governance by introducing an effective model of continuous quality assessment of the dataset, which would help organizations increase the reliability of analytics, minimize the spread of errors, and increase the confidence in the decision-making systems it is based on.},
keywords = {Automated Data Quality Scoring; Dataset Trust Score; Data Profiling and Validation; Analytics Readiness Assessment; Data Drift and Integrity Monitoring; Data Quality Automation; AI-Driven Data Governance.},
month = {December},
doi = {https://doi.org/10.64388/IREV7I6-1716725}
}