{"repo":"josephmachado/efficient_data_processing_spark","free":true,"listed":false,"github":"https://github.com/josephmachado/efficient_data_processing_spark","clone":"git clone https://github.com/josephmachado/efficient_data_processing_spark.git","description":"Code for \"Efficient Data Processing in Spark\" Course","language":"Python","stars":394,"topics":["apache-spark","data-engineering","data-pipeline","minio","pyspark","pyspark-notebook"],"license":null,"category":"data-pipelines","readme_excerpt":"Code for my Efficient Data Processing in Spark course. Efficient Data Processing in Spark - Efficient Data Processing in Spark - Setup - Create aliases for long commands with a Makefile - Run a Jupyter notebook - Infrastructure Repository for examples and exercises from the \"Efficient Data Processing in Spark\" course (under data-processing-spark). The capstone project is also present in this repository (under capstone/rainforest). Setup In order to run the project you'll need to install the following: 1. git version = 2.37.1 2. Docker version = 20.10.17 and Docker compose v2 version = v2.10.2. Windows users : please setup WSL and a local Ubuntu Virtual machine following the instructions here . Install the above prerequisites on your ubuntu terminal; if you have trouble installing docker, follow the steps here (only Step 1 is necessary). Please install the make command with sudo apt install make -y (if its not already present). All the commands shown below are to be run via the terminal (use the Ubuntu terminal for WSL users). The make commands in this book should be run in the efficient data processing spark folder. We will use docker to set up our containers. Clone and move into the lab repository, as shown below. Note : If you are using mac M1 or later, please replace the \"FROM deltaio/delta-docker:latest\" in data-processing-spark/1-lab-setup/containers/spark/Dockerfile with \"FROM deltaio/delta-docker:latest arm64\" Create aliases for long commands with a Makefile Makefile l","default_branch":null,"files":null,"tree":[],"storefront":"/r/josephmachado","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/josephmachado/efficient_data_processing_spark/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}