{"repo":"HamzaG737/data-engineering-project","free":true,"listed":false,"github":"https://github.com/HamzaG737/data-engineering-project","clone":"git clone https://github.com/HamzaG737/data-engineering-project.git","description":"End to end data engineering project with kafka, airflow, spark, postgres and docker.","language":"Python","stars":118,"topics":["api","data-engineering","docker","kafka","postgresql","spark"],"license":"MIT","category":"data-pipelines","readme_excerpt":"Building a simple End-to-End Data Engineering System This project uses different tools such as kafka, airflow, spark, postgres and docker. A step by step guide to run this pipeline: https://medium.com/@hamzagharbi 19502/end-to-end-data-engineering-system-on-real-data-with-kafka-spark-airflow-postgres-and-docker-a70e18df4090 Overview 1. Data Streaming: Initially, data is streamed from the API into a Kafka topic. 2. Data Processing: A Spark job then takes over, consuming the data from the Kafka topic and transferring it to a PostgreSQL database. 3. Scheduling with Airflow: Both the streaming task and the Spark job are orchestrated using Airflow. While in a real-world scenario, the Kafka producer would constantly listen to the API, for demonstration purposes, we'll schedule the Kafka streaming task to run daily. Once the streaming is complete, the Spark job processes the data, making it ready for use by the LLM application. All of these tools will be built and run using docker, and more specifically docker-compose.","default_branch":null,"files":null,"tree":[],"storefront":"/r/HamzaG737","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/HamzaG737/data-engineering-project/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}