{"repo":"shreyamalogi/Retail-Data-Engineering-Pipeline","free":true,"listed":false,"github":"https://github.com/shreyamalogi/Retail-Data-Engineering-Pipeline","clone":"git clone https://github.com/shreyamalogi/Retail-Data-Engineering-Pipeline.git","description":"Scalable ETL Pipeline: Processing 5M+ retail records with PySpark on GCP Dataproc. Automated the extraction of global business KPIs and consumer trends. Includes an Ethical Data Framework to ensure privacy and fairness at scale","language":"Python","stars":15,"topics":["python","bigdataanalytics","pyspark","big-data","data-engineering","distributed-computing","etl-pipeline","gcp","gcp-dataproc","google-cloud-storage"],"license":"MIT","category":"data-pipelines","readme_excerpt":"Retail-Data-Engineering-Pipeline: Engineering Insights from 5M+ Transactions High-Throughput ETL Distributed Cloud Computing Ethical Big Data 📖 The Narrative: Scaling Business Intelligence In enterprise retail, the challenge isn't just having data; it's the speed at which you can turn 5 million raw transactions into a roadmap for growth. This project tells the story of architecting a production-grade ETL pipeline on the Google Cloud Platform (GCP) , designed to handle high-velocity data and extract complex financial KPIs across global markets while maintaining a rigorous framework for data ethics. --- 🏗️ Chapter 1: The Ingestion Layer (GCS to Spark) Processing 5 million records requires more than a script; it requires a cloud-native ecosystem capable of horizontal scaling. Storage Orchestration : Managed the full storage lifecycle using Google Cloud Storage (GCS) , ensuring low-latency data availability for the compute cluster. Distributed Compute : Provisioned and managed GCP Dataproc clusters to perform parallelized transformations, significantly reducing processing time compared to local execution. Data Flow : Engineered the bridge from raw CSV assets in GCS to active Spark RDDs/DataFrames for high-speed manipulation. --- ⚙️ Chapter 2: Performance Engineering & ETL To extract value from a high-velocity dataset, I engineered a multi-stage transformation pipeline focusing on resource optimization and analytical depth . Aggregate Logic : Developed Spark jobs to calculate gl","default_branch":null,"files":null,"tree":[],"storefront":"/r/shreyamalogi","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/shreyamalogi/Retail-Data-Engineering-Pipeline/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}