Home / Current Issue / Paper 1716726
Reliability of LLM-Assisted Data Cleaning in Pandas Pipelines: An Empirical Evaluation Framework for Detecting Silent Data Corruption
Subject area: Science,Engineering and Technology · Area of research: Pandas Pipelines
Abstract
Large Language Models (LLMs) are being used in data science pipelines in more and more cases to automate tabular data preprocessing in Pandas pipelines. Nevertheless, current evaluation standards are mostly focused on syntactic accuracy and unit-test accuracy, but not much on the semantic accuracy of the data transformations generated. Type casting, missing value imputation, outlier, encoding, and normalisation operations of data cleaning may silently corrupt statistical distributions and undercut event validity of downstream analytics, without inducing execution errors. The current paper is a reliably conducted systematic cross-domain empirical assessment of data cleaning using LLM on healthcare, financial, e-commerce, and sensor data. Our evaluation rubric is multi-dimensional in that it covers the structural correctness, logical validity, statistical soundness, preservation of data integrity, and reproducibility on a scale of 0 to 3. In 5,150 cleaning operations, transformations generated by LLM were highly structurally correct (>90%), but semantically more reliable when compared by task category. Missing value processing and outlier detection had a high harm rate (10-15) and a silent error rate as high as 7%. In order to address those risks, we suggest an automated validation system that includes schema validation, distribution shift, distribution shift detection (Kolmogorov-Smirnov testing and variance analysis), tracking the null propagation, and constraint-based integrity checks. The framework minimised silent errors by about 60 per cent with a precision level of 0.91 and a recall of 0.88. These results indicate that syntax-based metrics cannot be used to assess AI-aided preprocessing and propose the need to address semantic stability metrics and automated protection of responsible usage of LLMs in production data pipelines.
Keywords
Large Language Models; LLM-assisted Programming; Data Cleaning; Pandas Pipelines; Silent Errors; Data Integrity; AI Reliability; Semantic Code Evaluation; Automated Validation; Data Preprocessing
References
[1] Benoit, A., Cavelan, A., Cappello, F., Raghavan, P., Robert, Y., & Sun, H. (2018). Coping with silent and fail-stop errors at scale by combining replication and checkpointing. Journal of Parallel and Distributed Computing, 122, 209–225. https://doi.org/10.1016/j.jpdc.2018.08.002
[2] Bibi, N., Maqbool, A., Rana, T., Afzal, F., Akgul, A., & Eldin, S. M. (2023). Enhancing Semantic Code Search With Deep Graph Matching. IEEE Access, 11, 52392–52411. https://doi.org/10.1109/ACCESS.2023.3263878
[3] Bilal, M., Ali, G., Iqbal, M. W., Anwar, M., Malik, M. S. A., & Kadir, R. A. (2022). Auto-Prep: Efficient and Automated Data Preprocessing Pipeline. IEEE Access, 10, 107764–107784. https://doi.org/10.1109/ACCESS.2022.3198662
[4] Blüthgen, C. (2025). Technical foundations of large language models. Radiologie, 65(4), 227–234. https://doi.org/10.1007/s00117-025-01427-z
[5] Carson, E. C., & Hercík, J. (2025). The detection and correction of silent errors in pipelined Krylov subspace methods. Numerical Algorithms. https://doi.org/10.1007/s11075-025-02037-5
[6] Casanova, H., Herrmann, J., & Robert, Y. (2018). Computing the expected makespan of task graphs in the presence of silent errors. Parallel Computing, 75, 41–60. https://doi.org/10.1016/j.parco.2018.03.004
[7] Chang, C., Li, M., Guo, C., Ding, Y., Xu, K., Han, M., … Zhu, Y. (2019). PANDA: A comprehensive and flexible tool for quantitative proteomics data analysis. Bioinformatics, 35(5), 898–900. https://doi.org/10.1093/bioinformatics/bty727
[8] Choi, H. S., Song, J. Y., Shin, K. H., Chang, J. H., & Jang, B. S. (2023). Developing prompts from a large language model for extracting clinical information from pathology and ultrasound reports in breast cancer. Radiation Oncology Journal, 41(3), 209–216. https://doi.org/10.3857/roj.2023.00633
[9] Cui, Z., Zhong, S., Xu, P., He, Y., & Gong, G. (2013). PANDA: A pipeline toolbox for analysing brain diffusion images. Frontiers in Human Neuroscience, (FEB). https://doi.org/10.3389/fnhum.2013.00042
[10] Fan, C., Chen, M., Wang, X., Wang, J., & Huang, B. (2021, March 29). A Review on Data Preprocessing Techniques Toward Efficient and Reliable Knowledge Discovery From Building Operational Data. Frontiers in Energy Research. Frontiers Media S.A. https://doi.org/10.3389/fenrg.2021.652801
[11] Haque, S., Eberhart, Z., Bansal, A., & McMillan, C. (2022). Semantic Similarity Metrics for Evaluating Source Code Summarisation. In IEEE International Conference on Program Comprehension (Vol. 2022-March, pp. 36–47). IEEE Computer Society. https://doi.org/10.1145/3524610.3527909
[12] Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., … Liu, T. (2025). A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Transactions on Information Systems, 43(2). https://doi.org/10.1145/3703155
[13] Jorgensen, S., Nadizar, G., Pietropolli, G., Manzoni, L., Medvet, E., O’Reilly, U.-M., & Hemberg, E. (2025). Policy Search through Genetic Programming and LLM-assisted Curriculum Learning. ACM Transactions on Evolutionary Learning and Optimisation. https://doi.org/10.1145/3772718
[14] Koukaras, P., & Tjortjis, C. (2025, October 1). Data Preprocessing and Feature Engineering for Data Mining: Techniques, Tools, and Best Practices. AI (Switzerland). Multidisciplinary Digital Publishing Institute (MDPI). https://doi.org/10.3390/ai6100257
[15] Krajnc, D., Spielvogel, C. P., Ecsedi, B., Ritter, Z., Alizadeh, H., Hacker, M., & Papp, L. (2025). Clinician-driven automated data preprocessing in nuclear medicine AI environments. European Journal of Nuclear Medicine and Molecular Imaging, 52(9), 3444–3454. https://doi.org/10.1007/s00259-025-07183-5
[16] Li, L., Znati, T., & Melhem, R. (2023). diffReplication ̶ An Energy-Aware Fault Tolerance Model for Silent Error Detection and Mitigation in Heterogeneous Extreme-scale Computing Environment. Journal of Universal Computer Science, 29(8), 892–910. https://doi.org/10.3897/jucs.94462
[17] Li, Y., Wang, T., Yu, L., & Pan, Z. (2025). Fus: Combining Semantic and Structural Graph Information for Binary Code Similarity Detection. Electronics (Switzerland), 14(19). https://doi.org/10.3390/electronics14193781
[18] Martins, P., Cardoso, F., Váz, P., Silva, J., & Abbasi, M. (2025). Performance and Scalability of Data Cleaning and Preprocessing Tools: A Benchmark on Large Real-World Datasets. Data, 10(5). https://doi.org/10.3390/data10050068
[19] Meurant, G. (2023). Detection and correction of silent errors in the conjugate gradient algorithm. Numerical Algorithms, 92(1), 869–891. https://doi.org/10.1007/s11075-022-01380-1
[20] Murray, B., Kerfoot, E., Chen, L., Deng, J., Graham, M. S., Sudre, C. H., … Ourselin, S. (2021). Accessible data curation and analytics for international-scale citizen science datasets. Scientific Data, 8(1). https://doi.org/10.1038/s41597-021-01071-x
[21] Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., … Mian, A. (2025). A Comprehensive Overview of Large Language Models. ACM Transactions on Intelligent Systems and Technology, 16(5). https://doi.org/10.1145/3744746
[22] Sin, C. K., & Kung, S. W. (2025). Implementation and development experience of an AI-assisted rostering system in a Hong Kong emergency department. Hong Kong Journal of Emergency Medicine, 32(6). https://doi.org/10.1002/hkj2.70061
[23] Tawakuli, A., Havers, B., Gulisano, V., Kaiser, D., & Engel, T. (2025). Survey: Time-series data preprocessing: A survey and an empirical analysis. Journal of Engineering Research (Kuwait), 13(2), 674–711. https://doi.org/10.1016/j.jer.2024.02.018
[24] Thomas, G. F., Martin, N. F., Fattahi, A., Ibata, R. A., Helly, J., McConnachie, A. W., … Pakmor, R. (2021). Observing the Stellar Halo of Andromeda in Cosmological Simulations: The AURIGA2PANDAS Pipeline. The Astrophysical Journal, 910(2), 92. https://doi.org/10.3847/1538-4357/abdfd2
[25] Wong, M. F., & Tan, C. W. (2024). Aligning Crowd-Sourced Human Feedback for Reinforcement Learning on Code Generation by Large Language Models. IEEE Transactions on Big Data. https://doi.org/10.1109/TBDATA.2024.3524104
[26] Xu, R., Jung, H., Choueiry, F., Zhang, S., Pearlman, R., Hampel, H., … Zhu, J. (2025). Novel machine‐learning bioinformatics reveal distinct metabolic alterations for enhanced colorectal cancer diagnosis and monitoring. IMetaOmics, 2(2). https://doi.org/10.1002/imo2.70003
[27] Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., & Chen, E. (2024, December 1). A survey on multimodal large language models. National Science Review. Oxford University Press. https://doi.org/10.1093/nsr/nwae403
[28] Yu, H., Hu, X., Li, G., Li, Y., Wang, Q., & Xie, T. (2022). Assessing and Improving an Evaluation Dataset for Detecting Semantic Code Clones via Deep Learning. ACM Transactions on Software Engineering and Methodology, 31(4). https://doi.org/10.1145/3502852
[29] Zhang, X., Lin, Z., Hu, X., Wang, J., Lu, W., & Zhou, D. Y. (2025). SECON: Maintaining Semantic Consistency in Data Augmentation for Code Search. ACM Transactions on Information Systems, 43(2). https://doi.org/10.1145/3686151
[30] Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., … Du, M. (2024). Explainability for Large Language Models: A Survey. ACM Transactions on Intelligent Systems and Technology, 15(2). https://doi.org/10.1145/3639372
How to cite this paper
@article{1716726,
author = {Sai Lalitesh Pothukuchi},
title = {Reliability of LLM-Assisted Data Cleaning in Pandas Pipelines: An Empirical Evaluation Framework for Detecting Silent Data Corruption},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {8},
number = {8},
pages = {1172-1182},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1716726.pdf},
abstract = {Large Language Models (LLMs) are being used in data science pipelines in more and more cases to automate tabular data preprocessing in Pandas pipelines. Nevertheless, current evaluation standards are mostly focused on syntactic accuracy and unit-test accuracy, but not much on the semantic accuracy of the data transformations generated. Type casting, missing value imputation, outlier, encoding, and normalisation operations of data cleaning may silently corrupt statistical distributions and undercut event validity of downstream analytics, without inducing execution errors. The current paper is a reliably conducted systematic cross-domain empirical assessment of data cleaning using LLM on healthcare, financial, e-commerce, and sensor data. Our evaluation rubric is multi-dimensional in that it covers the structural correctness, logical validity, statistical soundness, preservation of data integrity, and reproducibility on a scale of 0 to 3. In 5,150 cleaning operations, transformations generated by LLM were highly structurally correct (>90%), but semantically more reliable when compared by task category. Missing value processing and outlier detection had a high harm rate (10-15) and a silent error rate as high as 7%. In order to address those risks, we suggest an automated validation system that includes schema validation, distribution shift, distribution shift detection (Kolmogorov-Smirnov testing and variance analysis), tracking the null propagation, and constraint-based integrity checks. The framework minimised silent errors by about 60 per cent with a precision level of 0.91 and a recall of 0.88. These results indicate that syntax-based metrics cannot be used to assess AI-aided preprocessing and propose the need to address semantic stability metrics and automated protection of responsible usage of LLMs in production data pipelines.},
keywords = {Large Language Models; LLM-assisted Programming; Data Cleaning; Pandas Pipelines; Silent Errors; Data Integrity; AI Reliability; Semantic Code Evaluation; Automated Validation; Data Preprocessing},
month = {February},
doi = {https://doi.org/10.64388/IREV8I8-1716726}
}