{"repo":"Smars-Bin-Hu/azure-cloud-datapipeline-EDA","free":true,"listed":false,"github":"https://github.com/Smars-Bin-Hu/azure-cloud-datapipeline-EDA","clone":"git clone https://github.com/Smars-Bin-Hu/azure-cloud-datapipeline-EDA.git","description":"A cloud-native data pipeline and visualization project analyzing Formula 1 racing data using Azure, Databricks, Delta Lake, Tableau, and Python for insightful EDA and interactive dashboards.","language":"Jupyter Notebook","stars":98,"topics":["azure-data-factory","azure-data-lake-storage-gen2","azure-databricks","bi-dashboard","data-visualization","delta-lake","exploratory-data-analysis","lakehouse","matplotlib","medallion-architecture"],"license":"MIT","category":"data-pipelines","readme_excerpt":"👆 click the picture to see the presentation video! Cloud Native Data Pipeline on Azure Databricks for Exploratory Data Analysis This project presents an end-to-end data pipeline and analytics workflow centered around Formula 1 racing data, with a strong emphasis on exploratory data analysis (EDA) and visualization. It is structured into two core components: cloud-native data engineering and analytical data visualization. On the data engineering side, we leverage Azure Cloud services to build a scalable and automated data pipeline following the medallion architecture (bronze, silver, gold layers) . Raw data is ingested and stored in Azure Data Lake Storage Gen2 , processed and transformed using Azure Databricks with PySpark and SparkSQL , and managed through Delta Lake to ensure ACID transactions and schema enforcement. We incorporate Unity Catalog for data governance and access control, while Azure Data Factory orchestrates the workflow to achieve full automation. The architecture demonstrates cloud-native best practices such as decoupled storage and compute, batch-stream unification, and automated job triggering—effectively realizing a modern Lakehouse design. On the data analysis and visualization front, we utilize both Tableau and Python for different analytical tasks. Tableau connects directly to the Databricks-backed gold layer, enabling real-time, interactive BI dashboards that cover historical driver and team rankings, national-level aggregations, and top driver trend","default_branch":null,"files":null,"tree":[],"storefront":"/r/Smars-Bin-Hu","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Smars-Bin-Hu/azure-cloud-datapipeline-EDA/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}