Home / Current Issue / Paper 1707512
Real-Time Data Processing for Machine Learning: Streaming Architectures, Challenges, and Use Cases
Subject area: Science,Engineering and Technology · Area of research: Data engineering and machine learning
Abstract
Real-time data processing is a critical component of modern machine learning (ML) applications, enabling rapid insights, decision-making, and automation across various industries. This paper explores the architectures, challenges, and practical use cases of real-time data streaming for ML workflows. It delves into key streaming frameworks such as Apache Kafka, Apache Flink, and Spark Streaming, highlighting their roles in efficient data ingestion, transformation, and model deployment. The challenges of real-time ML, including data latency, scalability, fault tolerance, and data quality, are analyzed with potential solutions. Furthermore, real-world applications such as fraud detection, recommendation systems, predictive maintenance, and anomaly detection are discussed to showcase the impact of real-time streaming on AI-driven systems. The paper concludes by addressing future trends in real-time ML, including edge computing, federated learning, and cloud-native streaming solutions, emphasizing their growing importance in handling dynamic and large-scale data environments.
Keywords
Real-time Data Processing, Machine Learning (ML), Data Streaming, Apache Kafka, Apache Flink.
References
[1] Marz, Nathan, and Warren, James. Big Data: Principles and Best Practices of Scalable Realtime Data Systems. Manning Publications, 2013.
[2] Gama, João, Žliobaitė, Indrė, Bifet, Albert, Pechenizkiy, Mykola, and Bouchachia, Anna. "A survey on concept drift adaptation." ACM Computing Surveys (CSUR) 46.4 (2014): 1-37.
[3] Alexandrov, Alexander, et al. "The Stratosphere platform for big data analytics." The VLDB Journal 23.6 (2014): 939-964.
[4] Zaharia, Matei, et al. "Structured Streaming in Apache Spark: A new high-level API for streaming." (2016).
[5] Bifet, Albert, Gavaldà, Ricard, Holmes, Geoff, and Pfahringer, Bernhard. Machine Learning for Data Streams with Practical Examples in MOA. MIT Press, 2018.
How to cite this paper
@article{1707512,
author = {Bhanu Prakash Reddy Rella},
title = {Real-Time Data Processing for Machine Learning: Streaming Architectures, Challenges, and Use Cases},
journal = {Iconic Research And Engineering Journals},
year = {2021},
volume = {5},
number = {4},
pages = {230-234},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1707512.pdf},
abstract = {Real-time data processing is a critical component of modern machine learning (ML) applications, enabling rapid insights, decision-making, and automation across various industries. This paper explores the architectures, challenges, and practical use cases of real-time data streaming for ML workflows. It delves into key streaming frameworks such as Apache Kafka, Apache Flink, and Spark Streaming, highlighting their roles in efficient data ingestion, transformation, and model deployment. The challenges of real-time ML, including data latency, scalability, fault tolerance, and data quality, are analyzed with potential solutions. Furthermore, real-world applications such as fraud detection, recommendation systems, predictive maintenance, and anomaly detection are discussed to showcase the impact of real-time streaming on AI-driven systems. The paper concludes by addressing future trends in real-time ML, including edge computing, federated learning, and cloud-native streaming solutions, emphasizing their growing importance in handling dynamic and large-scale data environments.},
keywords = {Real-time Data Processing, Machine Learning (ML), Data Streaming, Apache Kafka, Apache Flink.},
month = {October},
}