{"repo":"MinhKhoi3104/aws-end-to-end-data-lakehouse-analytics-platform","free":true,"listed":false,"github":"https://github.com/MinhKhoi3104/aws-end-to-end-data-lakehouse-analytics-platform","clone":"git clone https://github.com/MinhKhoi3104/aws-end-to-end-data-lakehouse-analytics-platform.git","description":"An end-to-end AWS data lakehouse leveraging batch data processing, big data technologies, and advanced business analytics.","language":"Python","stars":11,"topics":["airflow","aws-cloud","data-modeling","data-warehouse","dbt","docker","iceberg","pyspark","python","terraform"],"license":"MIT","category":"data-pipelines","readme_excerpt":"🚀 AWS End-to-End Data Lakehouse Analytics Platform This project implements an AWS-based Data Lakehouse platform for processing and analyzing large-scale user search and behavior data across Internet TV, OTT, and online entertainment platforms . The system follows a Lakehouse architecture on Amazon S3 using Apache Iceberg as the table format, while data ingestion and transformation across the Bronze, Silver, and Gold layers are performed using Apache Spark (PySpark) . At the Gold layer , data is modeled into fact and dimension tables using a star schema , providing curated and analytics-ready datasets. The Gold-layer datasets are then replicated into PostgreSQL , which serves as the analytical serving layer for downstream consumption. Within PostgreSQL, dbt is used to apply business transformations, build datamarts , and define analytical metrics in a structured and version-controlled manner. These datamarts are subsequently consumed by BI and visualization tools (Apache Superset) , enabling efficient OLAP-style analysis and reporting. The data pipeline is orchestrated using Apache Airflow , with infrastructure provisioned via Terraform (Infrastructure as Code) on AWS to ensure scalable, reproducible environments, and system observability supported by Grafana and Prometheus . End-to-End Project Overview 📋 Table of Contents - 📁 Project Structure - 📚 Dataset - Customer search data log - Crawled Movie Dataset - 🌐 Architecture Overview - 1. AWS Configuration (Infrastructure a","default_branch":null,"files":null,"tree":[],"storefront":"/r/MinhKhoi3104","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/MinhKhoi3104/aws-end-to-end-data-lakehouse-analytics-platform/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}