{"repo":"Victor-Kipruto-Rop/cloud-etl-pipeline","free":true,"listed":false,"github":"https://github.com/Victor-Kipruto-Rop/cloud-etl-pipeline","clone":"git clone https://github.com/Victor-Kipruto-Rop/cloud-etl-pipeline.git","description":"Built a scalable cloud ETL pipeline using Python, PostgreSQL, Pandas, and Docker to efficiently process, clean, and load over 1.6 million records with optimized batch processing, fault tolerance, and automated data validation.","language":"Python","stars":27,"topics":["data-engineering","etl","aws","cloud","postresql","python","terraform","warehouse","grafana","playbook"],"license":"MIT","category":"data-pipelines","readme_excerpt":"ETL Pipeline A Python ETL repository that ingests Kaggle datasets, runs local extract/transform/load workflows, and optionally uploads raw and processed data to AWS S3. What this project contains - Local Kaggle ingestion scripts for e-commerce, healthcare, finance, sports, and climate domains under ingest/ - A modular ETL pipeline under src/ and etl/ - Root environment configuration templates in .env.example and reusable settings under config/ - AWS helper code for S3 upload and optional Redshift load under src/cloud/ - E-commerce analytics SQL in analytics/ecommerce queries.sql - Warehouse schema DDL in warehouse/schemas/ .sql - Monitoring examples in monitoring/ - A pytest-based test suite in tests/ Note: terraform/ contains a Terraform root configuration and AWS provider file, but the referenced Terraform module sources are not included in this repository. The supported workflow is local development with optional AWS helper support. Status - Local ETL and data ingestion are implemented in Python. - AWS S3 upload and optional Redshift helper methods exist, but full multi-service cloud provisioning is not available in this checkout. - dags/ and k8s/ provide deployment skeletons rather than a complete cloud production stack. - .env is a local configuration file that should not be committed. - Data directories under data/ are excluded from version control and should be created locally. - This repository is best used for local pipeline development, testing, and Kaggle ingestion","default_branch":null,"files":null,"tree":[],"storefront":"/r/Victor-Kipruto-Rop","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Victor-Kipruto-Rop/cloud-etl-pipeline/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}