International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1702915

1702915PublishedVol 5 · Issue 4

Optimizing Cloud-Based Data Pipelines Using AWS, Kafka, and Postgres

Akash Balaji Mali Rahul Arulkumaran Ravi Kiran Pagidi Dr S P Singh Prof. (Dr) Sandeep Kumar Shalu Jain

Subject area: Science,Engineering and Technology  ·  Area of research: Data Pipelines

Abstract

The increasing reliance on data-driven insights has made the optimization of cloud-based data pipelines a critical aspect of modern business operations. This paper explores the use of AWS, Apache Kafka, and PostgreSQL as key components in designing efficient, scalable, and fault-tolerant data pipelines. AWS provides cloud infrastructure for storage, compute, and orchestration, ensuring high availability and scalability. Kafka, as a distributed messaging system, enables real-time data streaming with low latency, supporting event-driven architectures. PostgreSQL serves as the relational database for structured data storage, offering robust querying capabilities and transaction management. The study focuses on best practices for integrating these technologies to address challenges such as data latency, reliability, and performance bottlenecks. It further examines automation techniques using AWS Lambda and Glue for seamless data transformation and orchestration. The proposed framework enhances data pipeline efficiency by leveraging parallel processing and stream management techniques while ensuring secure data flows. This research provides actionable insights into creating agile and cost-effective data pipelines, supporting both real-time analytics and long-term data retention.

Keywords

Cloud-based data pipelines, AWS, Apache Kafka, PostgreSQL, real-time data streaming, data integration, scalability, fault tolerance, low-latency processing, data transformation, pipeline optimization, event-driven architecture, cloud automation, parallel processing, secure data flow.

How to cite this paper

Akash Balaji Mali, Rahul Arulkumaran, Ravi Kiran Pagidi, Dr S P Singh, Prof. (Dr) Sandeep Kumar; Shalu Jain "Optimizing Cloud-Based Data Pipelines Using AWS, Kafka, and Postgres" Iconic Research And Engineering Journals Volume 5 Issue 4 2021 Page 153-178
Akash Balaji Mali, Rahul Arulkumaran, Ravi Kiran Pagidi, Dr S P Singh, Prof. (Dr) Sandeep Kumar; Shalu Jain "Optimizing Cloud-Based Data Pipelines Using AWS, Kafka, and Postgres" Iconic Research And Engineering Journals, vol. 5, no. 4, Oct. 2021
Akash Balaji Mali, Rahul Arulkumaran, Ravi Kiran Pagidi, Dr S P Singh, Prof. (Dr) Sandeep Kumar; Shalu Jain (2021). Optimizing Cloud-Based Data Pipelines Using AWS, Kafka, and Postgres. Iconic Research And Engineering Journals, 5(4).
Akash Balaji Mali, Rahul Arulkumaran, Ravi Kiran Pagidi, Dr S P Singh, Prof. (Dr) Sandeep Kumar; Shalu Jain "Optimizing Cloud-Based Data Pipelines Using AWS, Kafka, and Postgres" Iconic Research And Engineering Journals, vol. 5, no. 4, Oct. 2021.
@article{1702915,
      author = {Akash Balaji Mali, Rahul Arulkumaran, Ravi Kiran Pagidi, Dr S P Singh, Prof. (Dr) Sandeep Kumar; Shalu Jain},
      title = {Optimizing Cloud-Based Data Pipelines Using AWS, Kafka, and Postgres},
      journal = {Iconic Research And Engineering Journals},
      year = {2021},
      volume = {5},
      number = {4},
      pages = {153-178},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1702915.pdf},
      abstract = {The increasing reliance on data-driven insights has made the optimization of cloud-based data pipelines a critical aspect of modern business operations. This paper explores the use of AWS, Apache Kafka, and PostgreSQL as key components in designing efficient, scalable, and fault-tolerant data pipelines. AWS provides cloud infrastructure for storage, compute, and orchestration, ensuring high availability and scalability. Kafka, as a distributed messaging system, enables real-time data streaming with low latency, supporting event-driven architectures. PostgreSQL serves as the relational database for structured data storage, offering robust querying capabilities and transaction management. The study focuses on best practices for integrating these technologies to address challenges such as data latency, reliability, and performance bottlenecks. It further examines automation techniques using AWS Lambda and Glue for seamless data transformation and orchestration. The proposed framework enhances data pipeline efficiency by leveraging parallel processing and stream management techniques while ensuring secure data flows. This research provides actionable insights into creating agile and cost-effective data pipelines, supporting both real-time analytics and long-term data retention.},
      keywords = {Cloud-based data pipelines, AWS, Apache Kafka, PostgreSQL, real-time data streaming, data integration, scalability, fault tolerance, low-latency processing, data transformation, pipeline optimization, event-driven architecture, cloud automation, parallel processing, secure data flow.},
      month = {October},
  }