{"repo":"cluster-apps-on-docker/spark-standalone-cluster-on-docker","free":true,"listed":false,"github":"https://github.com/cluster-apps-on-docker/spark-standalone-cluster-on-docker","clone":"git clone https://github.com/cluster-apps-on-docker/spark-standalone-cluster-on-docker.git","description":"Learn Apache Spark in Scala, Python (PySpark) and R (SparkR) by building your own cluster with a JupyterLab interface on Docker. :zap:","language":"Jupyter Notebook","stars":512,"topics":["spark","docker","python","scala","pyspark","jupyter","r","sparkr"],"license":"MIT","category":"deployment-docker-iac","readme_excerpt":"Apache Spark Standalone Cluster on Docker The project was featured on an article at MongoDB official tech blog! :scream: The project just got its own article at Towards Data Science Medium blog! :sparkles: Introduction This project gives you an Apache Spark cluster in standalone mode with a JupyterLab interface built on top of Docker . Learn Apache Spark through its Scala , Python (PySpark) and R (SparkR) API by running the Jupyter notebooks with examples on how to read, process and write data. TL;DR Contents - Quick Start - Tech Stack - Metrics - Contributing - Contributors - Support Quick Start Cluster overview Application URL Description ----------------- ------------------------------------------ ------------------------------------------------------------ JupyterLab localhost:8888 Cluster interface with built-in Jupyter notebooks Spark Driver localhost:4040 Spark Driver web ui Spark Master localhost:8080 Spark Master node Spark Worker I localhost:8081 Spark Worker node with 1 core and 512m of memory (default) Spark Worker II localhost:8082 Spark Worker node with 1 core and 512m of memory (default) Prerequisites - Install Docker and Docker Compose, check infra supported versions Download from Docker Hub (easier) 1. Download the docker compose file; 2. Edit the docker compose file with your favorite tech stack version, check apps supported versions; 3. Start the cluster; 4. Run Apache Spark code using the provided Jupyter notebooks with Scala, PySpark and SparkR examples; ","default_branch":null,"files":null,"tree":[],"storefront":"/r/cluster-apps-on-docker","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/cluster-apps-on-docker/spark-standalone-cluster-on-docker/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}