Home / Current Issue / Paper 1723188
Evaluating the Performance of Serverless Data Analytics Platforms for Large-Scale Data Processing
Subject area: Science,Engineering and Technology · Area of research: Serverless Data Analytics
Abstract
The rapid growth of data-intensive applications has increased the demand for scalable and efficient cloud analytics platforms. This study evaluates the performance of SQL-on-Lake and traditional cloud data warehouse architectures for analytical workloads. The evaluation compares selected platforms based on query execution time, throughput, response latency, scalability, cost efficiency, and performance stability. A benchmark-driven experimental approach was adopted using standardized analytical workloads, and the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) method was applied for overall performance ranking. The results show that Google BigQuery achieved the best overall performance, recording the lowest query execution time, highest throughput, superior scalability, and improved cost efficiency, followed by Snowflake and Amazon Athena. The findings indicate that cloud data warehouses provide better performance for intensive analytical workloads, while SQL-on-Lake architectures remain suitable for flexible and cost-effective data management. The study recommends selecting cloud analytics platforms based on workload requirements, optimizing SQL-on-Lake environments through effective data management techniques, and adopting hybrid approaches where both flexibility and performance are required.
Keywords
Cloud Computing, Data Lakehouse, SQL-on-Lake, Cloud Data Warehouse, Big Data Analytics, Performance Evaluation, TOPSIS.
References
[1] Abadi, D. J. (2012). Consistency tradeoffs in modern distributed database system design: CAP is only part of the story. Computer, 45(2), 37–42. Crossref
[2] Armbrust, M., Fox, A., Griffith, R., Joseph, A. D., Katz, R., Konwinski, A., Lee, G., Patterson, D. A., Rabkin, A., Stoica, I., & Zaharia, M. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50–58. ACM
[3] Armbrust, M., Ghodsi, A., Xin, R., Zaharia, M., Franklin, M. J., Stoica, I., & Gonzalez, J. E. (2021). Lakehouse: A new generation of open platforms that unify data warehousing and advanced analytics. In Proceedings of the 11th Biennial Conference on Innovative Data Systems Research (CIDR 2021). CIDR
[4] Baldini, I., Castro, P., Chang, K., Cheng, P., Fink, S., Ishakian, V., Mitchell, N., Muthusamy, V., Rabbah, R., Slominski, A., & Suter, P. (2017). Serverless computing: Current trends and open problems. In L. Bougé, C. de Laat, J. M. Menaud, & J. Tourancheau (Eds.), Research advances in cloud computing (pp. 1–20). Springer. Springer
[5] Castro, P., Ishakian, V., Muthusamy, V., & Slominski, A. (2019). The rise of serverless computing. Communications of the ACM, 62(12), 44–54. ACM
[6] Eizaguirre, G. T., & Sánchez-Artigas, M. (2024). A Seer knows best: Auto-tuned object storage shuffling for serverless analytics. Journal of Parallel and Distributed Computing, 183, 104763. ScienceDirect
[7] Hashem, I. A. T., Yaqoob, I., Anuar, N. B., Mokhtar, S., Gani, A., & Khan, S. U. (2015). The rise of big data on cloud computing: Review and open research issues. Information Systems, 47, 98–115. ScienceDirect
[8] Hwang, C. L., & Yoon, K. (1981). Multiple attribute decision making: Methods and applications: A state-of-the-art survey. Springer. Springer
[9] Jonas, E., Schleier-Smith, J., Sreekanti, V., Tsai, C.-C., Khandelwal, A., Pu, Q., Shankar, V., Carreira, J., Krauth, K., Yadwadkar, N., Gonzalez, J. E., Popa, R. A., Stoica, I., & Patterson, D. A. (2019). Cloud programming simplified: A Berkeley view on serverless computing (Technical Report No. UCB/EECS-2019-3). University of California, Berkeley. UC Berkeley
[10] Khezr, S., Benatallah, B., & Dustdar, S. (2024). Serverless computing: Recent advances, challenges, and future directions. ACM Computing Surveys, 56(8), Article 214. ACM
[11] Shojaee Rad, Z., & Ghobaei-Arani, M. (2024). Data pipeline approaches in serverless computing: A taxonomy, review, and research trends. Journal of Big Data, 11, Article 82. Springer
[12] Li, Y., Lin, Y., Wang, Y., Ye, K., & Xu, C. (2023). Serverless computing: State-of-the-art, challenges and opportunities. IEEE Transactions on Services Computing, 16(2), 1522–1539. Crossref
[13] van Renen, A., & Leis, V. (2023). Cloud Analytics Benchmark. Proceedings of the VLDB Endowment, 16(6), 1413–1425. VLDB
[14] Wen, J., Chen, Z., Jin, X., & Liu, X. (2023). Rise of the planet of serverless computing: A systematic review. ACM Transactions on Software Engineering and Methodology, 32(5), Article 131, 1–61. ACM
[15] Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J. E., Shenker, S., & Stoica, I. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65. ACM
How to cite this paper
@article{1723188,
author = {Mission Franklin, Ngozi Kingsley-Opara},
title = {Evaluating the Performance of Serverless Data Analytics Platforms for Large-Scale Data Processing},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {3},
pages = {2188-2199},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1723188.pdf},
abstract = {The rapid growth of data-intensive applications has increased the demand for scalable and efficient cloud analytics platforms. This study evaluates the performance of SQL-on-Lake and traditional cloud data warehouse architectures for analytical workloads. The evaluation compares selected platforms based on query execution time, throughput, response latency, scalability, cost efficiency, and performance stability. A benchmark-driven experimental approach was adopted using standardized analytical workloads, and the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) method was applied for overall performance ranking. The results show that Google BigQuery achieved the best overall performance, recording the lowest query execution time, highest throughput, superior scalability, and improved cost efficiency, followed by Snowflake and Amazon Athena. The findings indicate that cloud data warehouses provide better performance for intensive analytical workloads, while SQL-on-Lake architectures remain suitable for flexible and cost-effective data management. The study recommends selecting cloud analytics platforms based on workload requirements, optimizing SQL-on-Lake environments through effective data management techniques, and adopting hybrid approaches where both flexibility and performance are required.},
keywords = {Cloud Computing, Data Lakehouse, SQL-on-Lake, Cloud Data Warehouse, Big Data Analytics, Performance Evaluation, TOPSIS.},
month = {September},
}