Home / Current Issue / Paper 1705652
Data Anonymization in AI and ML Engineering: Balancing Privacy and Model Performance Using Presidio
Subject area: Science,Engineering and Technology · Area of research: Data Anonymization
Abstract
Data anonymization plays a pivotal role in Artificial Intelligence (AI) & Machine Learning (ML) to ensure individual privacy even as it allows data-led insight. Large datasets in healthcare, finance, and many other industries carry highly sensitive information; consequently, privacy is a sensitive issue. Microsoft's open-source tool, Presidio, detects and eliminates personally identifiable information (PII) in structured and unstructured data. In this article, we examine the tradeoff between privacy and model performance, where Presidio's machine learning-friendly PII anonymization techniques for tokens, masks, and data perturbation allow good data to be protected while still providing utility. Moreover, the article looks at how Presidio helps an organization meet the requirements of privacy laws like GDPR, HIPAA, and CCPA and how it does it responsibly. Data anonymization has key challenges, such as loss of model accuracy and re-identification risks, which are discussed with insight into how Presidio helps mitigate them. Presidio takes this further by utilizing effective data anonymization to allow organizations to develop privacy-compliant, high-performing AI systems that respect privacy and run on data responsibly.
Keywords
Data Anonymization, AI Privacy, Machine Learning, Presidio, Personally Identifiable Information (PII)
References
[1] Chen, L. (2022, October 7). PII anonymization made easy by Presidio - Towards Data Science. Medium. https://towardsdatascience.com/building-a-customized-pii-anonymizer-with-microsoft-presidio-b5c2ddfe523b
[2] Benchmarking Advanced Text Anonymisation Methods: A comparative study on novel and traditional approaches. (n.d.). Retrieved from https://arxiv.org/html/2404.14465v1
[3] Latanya Sweeney. “k-anonymity: a model for protecting privacy”. In: International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems 10.5 (Oct. 1, 2002), pp. 557–570. ISSN: 0218-4885. DOI: 10.1142/S0218488502001648. URL: https://doi.org/10.1142/S0218488502001648 (visited on 11/21/2022) (cit. on p. 3).
[4] Vpnreports. (2022b, July 18). Is Anonymous Data Really Anonymous? VPN Reports. https://www.vpnreports.com/is-anonymous-data-really-anonymous/
[5] A. Machanavajjhala et al. “L-diversity: privacy beyond k-anonymity”. In: 22nd International Conference on Data Engineering (ICDE’06). 22nd International Conference on Data Engineering (ICDE’06). ISSN: 2375-026X. Apr. 2006, pp. 24–24. DOI: 10.1109/ICDE.2006.1 (cit. on p. 3).
[6] Ninghui Li, Tiancheng Li, and Suresh Venkatasubramanian. “t-Closeness: Privacy Beyond k-Anonymity and l-Diversity”. In: 2007 IEEE 23rd International Conference on Data Engineering. 2007 IEEE 23rd International Conference on Data Engineering. ISSN: 2375-026X. Apr. 2007, pp. 106–115. DOI: 10.1109/ICDE.2007.367856 (cit. on p. 3).
[7] Cynthia Dwork. “Differential Privacy: A Survey of Results”. In: Theory and Applications of Models of Computation. Ed. by Manindra Agrawal et al. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer, 2008, pp. 1–19. ISBN: 978-3-540- 79228-4. DOI: 10.1007/978-3-540-79228-4_1 (cit. on p. 3).
[8] Brijesh Mehta et al. “Towards privacy preserving unstructured big data publishing”. In: Journal of Intelligent & Fuzzy Systems 36.4 (Jan. 1, 2019). Publisher: IOS Press, pp. 3471–3482. ISSN: 1064-1246. DOI: 10.3233/JIFS- 181231. URL: https: / / content . iospress . com / articles / journal - of - intelligent - and - fuzzy - systems/ifs181231 (visited on 11/04/2022) (cit. on p. 3).
[9] Art. 4 GDPR – Definitions. General Data Protection Regulation (GDPR). URL: https: //gdpr-info.eu/art-4-gdpr/ (visited on 11/04/2022) (cit. on p. 5).
[10] Batet, Montserrat and David Sánchez. 2018. Semantic disclosure control: Semantics meets data privacy. Online Information Review, 42(3):290–303. https://doi.org/10.1108/OIR-03-2017-0090
[11] Batet, Montserrat and David Sánchez. 2020. Leveraging synonymy and polysemy to improve semantic similarity assessments based on intrinsic information content. Artificial Intelligence Review, 53(3):2023–2041. https://doi.org/10.1007/s10462-019-09725-4
[12] Mendels, Omri. 2020. Custom NLP approaches to data anonymization. Towards Data Science. https://towardsdatascience.com/nlp-approaches-to-data-anonymization-1fb5bde6b929. Accessed: 2022-06-06.
[13] Lison, P., Pilán, I., Sánchez, D., Batet, M., & Øvrelid, L. (2021). Anonymisation Models for Text Data: State of the Art, Challenges and Future Directions. In Association for Computational Linguistics & International Joint Conference on Natural Language Processing, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (pp. 4188–4203). https://aclanthology.org/2021.acl-long.323.pdf
[14] You, L.L. & Pollack, K.T. & Long, Darrell. (2005). Deep Store: an archival storage system architecture. Proceedings - International Conference on Data Engineering. 804- 815. 10.1109/ICDE.2005.47.
[15] Vpnreports. (2022, July 18). Is Anonymous Data Really Anonymous? VPN Reports. https://www.vpnreports.com/is-anonymous-data-really-anonymous/
[16] Pierre Lison, Ildikó Pilán, David Sánchez, Montserrat Batet, and Lilja Øvrelid, Anonymisation Models for Text Data: State of the Art, Challenges and Future Directions (2021). Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing
[17] Pierangela Samarati and Latanya Sweeney, Protecting Privacy when Disclosing Information: k-Anonymity and its Enforcement through Generalization and Suppression (1998). Technical report, SRI International
[18] Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith, Calibrating Noise to Sensitivity in Private Data Analysis (2006). Theory of Cryptography
[19] Franck Dernoncourt, Ji Young Lee, Ozlem Uzuner, and Peter Szolovits, De-identification of patient notes with recurrent neural networks (2017). Journal of the American Medical Informatics Association
[20] Alistair EW Johnson, Lucas Bulgarelli, and Tom J Pollard, De-identification of free-text medical records using pre-trained bidirectional transformers (2020). Proceedings of the ACM Conference on Health, Inference, and Learning
[21] Y. Pei, Y. Liu, N. Ling, Y. Ren and L. Liu, "An End-to-End Deep Generative Network for Low Bitrate Image Coding," 2023 IEEE International Symposium on Circuits and Systems (ISCAS), Monterey, CA, USA, 2023, pp. 1-5, doi: 10.1109/ISCAS46773.2023.10182028.
[22] Zhu, Y. (2023). Beyond Labels: A Comprehensive Review of Self-Supervised Learning and Intrinsic Data Properties. Journal of Science & Technology, 4(4), 65-84.
[23] B. Edwards, “Openai introduces gpt-4 turbo: Larger memory, lower cost, new knowledge,” Ars Technica, November 2023. [Online]. Available: https://arstechnica.com/information-technology/2023/11/ openai-introduces-gpt-4-turbo-larger-memory-lower-cost-new-knowledge/
[24] R. Khawaja, “Best large language models (llms) in 2024,” may 2024. [Online]. Available: https://datasciencedojo.com/blog/best-large-language-models/
[25] H. N. B, “Confusion matrix, accuracy, precision, recall, f1 score,” Medium, 2019. [Online]. Available: https://medium.com/analytics-vidhya/ confusion-matrix-accuracy-precision-recall-f1-score-ade299cf63cd
[26] D. R. Almeida, “Synthetic data generation (part 1),” Apr 2024. [Online]. Available: https://cookbook.openai.com/examples/sdg1
[27] “Managing environments,” 2017. [Online]. Available: https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html
[28] A. Stam and B. Kleiner, “Data anonymization: Legal, ethical, and strategic considerations,” FORS Guide No. 11, Version 1.1, Lausanne, 2020, last update January 2022.
[29] Krishna, K. (2022). Optimizing query performance in distributed NoSQL databases through adaptive indexing and data partitioning techniques. International Journal of Creative Research Thoughts (IJCRT). https://ijcrt. org/viewfulltext. php.
[30] Krishna, K., & Thakur, D. (2021). Automated Machine Learning (AutoML) for Real-Time Data Streams: Challenges and Innovations in Online Learning Algorithms. Journal of Emerging Technologies and Innovative Research (JETIR), 8(12).
[31] Murthy, P., & Thakur, D. (2022). Cross-Layer Optimization Techniques for Enhancing Consistency and Performance in Distributed NoSQL Database. International Journal of Enhanced Research in Management & Computer Applications, 35.
[32] Murthy, P., & Mehra, A. (2021). Exploring Neuromorphic Computing for Ultra-Low Latency Transaction Processing in Edge Database Architectures. Journal of Emerging Technologies and Innovative Research, 8(1), 25-26.
[33] Mehra, A. (2024). HYBRID AI MODELS: INTEGRATING SYMBOLIC REASONING WITH DEEP LEARNING FOR COMPLEX DECISION-MAKING. Journal of Emerging Technologies and Innovative Research (JETIR), Journal of Emerging Technologies and Innovative Research (JETIR), 11(8), f693-f695.
[34] Thakur, D. (2021). Federated Learning and Privacy-Preserving AI: Challenges and Solutions in Distributed Machine Learning. International Journal of All Research Education and Scientific Methods (IJARESM), 9(6), 3763-3764.
How to cite this paper
@article{1705652,
author = {Surya Gangadhar Patchipala},
title = {Data Anonymization in AI and ML Engineering: Balancing Privacy and Model Performance Using Presidio},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {6},
number = {10},
pages = {992-1004},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1705652.pdf},
abstract = {Data anonymization plays a pivotal role in Artificial Intelligence (AI) & Machine Learning (ML) to ensure individual privacy even as it allows data-led insight. Large datasets in healthcare, finance, and many other industries carry highly sensitive information; consequently, privacy is a sensitive issue. Microsoft's open-source tool, Presidio, detects and eliminates personally identifiable information (PII) in structured and unstructured data. In this article, we examine the tradeoff between privacy and model performance, where Presidio's machine learning-friendly PII anonymization techniques for tokens, masks, and data perturbation allow good data to be protected while still providing utility. Moreover, the article looks at how Presidio helps an organization meet the requirements of privacy laws like GDPR, HIPAA, and CCPA and how it does it responsibly. Data anonymization has key challenges, such as loss of model accuracy and re-identification risks, which are discussed with insight into how Presidio helps mitigate them. Presidio takes this further by utilizing effective data anonymization to allow organizations to develop privacy-compliant, high-performing AI systems that respect privacy and run on data responsibly.},
keywords = {Data Anonymization, AI Privacy, Machine Learning, Presidio, Personally Identifiable Information (PII)},
month = {April},
}