{"repo":"alero-awani/batch-data-engineering-project","free":true,"listed":false,"github":"https://github.com/alero-awani/batch-data-engineering-project","clone":"git clone https://github.com/alero-awani/batch-data-engineering-project.git","description":"A batch Data Pipeline that retrieves data from a user purchase table and a movie review table and is transformed to form a user behaviour metric table.","language":"HCL","stars":18,"topics":["airflow","aws-s3","data-engineering-pipeline","docker","pipeline","pyspark","sql","terraform"],"license":null,"category":"deployment-docker-iac","readme_excerpt":"Batch-Data-Engineering-Project The task is to build a data pipeline to populate the user behavior metric table. The user behavior metric table is an OLAP table, meant to be used by analysts, dashboard software, etc. It is built from user purchase, an OLTP table with user purchase information and movie review.csv, data sent every day by an external data vendor. REFERENCE https://www.startdataengineering.com/post/data-engineering-project-for-beginners-batch-edition/ Architecture Table of contents 1. Pipeline Workflow 2. Terraform setup 3. Airflow/Airflow Configurations Pipeline Workflow User Purchase Data The user purchase data is extracted from an OLTP database and loaded into the Redshift data warehouse. AWS S3 is used as storage for use with AWS Redshift Spectrum(data lakehouse) With Redshift Spectrum the data can be queried directly from s3 on Redshift by creating an external schema with the help of AWS Glue. Movie Review Data The movie review data is loaded into a staging area in an s3 bucket where it can be directly accessed by AWS EMR. The data is loaded along side a spark script The spark script performs basic text classification on the data and loads it back to the s3 bucket User Behaviour Metric Table The transformed movie review data and the user purchase data are joined in Redshift to get the user behaviour metric table Terraform setup overview This pipeline requires us to setup Apache Airflow, AWS EMR,AWS Redshift, AWS Spectrum, and AWS S3, AWS EC2. The EC2 instanc","default_branch":null,"files":null,"tree":[],"storefront":"/r/alero-awani","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/alero-awani/batch-data-engineering-project/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}