{"repo":"martandsingh/ApacheSpark","free":true,"listed":false,"github":"https://github.com/martandsingh/ApacheSpark","clone":"git clone https://github.com/martandsingh/ApacheSpark.git","description":"This repository will help you to learn about databricks concept with the help of examples. It will include all the important topics which we need in our real life experience as a data engineer. We will be using pyspark & sparksql for the development. At the end of the course we also cover few case studies.","language":"Python","stars":105,"topics":["apachespark","data-analysis","data-engineering","database","databricks","datalake","deltalake","etl-pipeline","hadoop","hive"],"license":"MIT","category":"databases-storage","readme_excerpt":"Data Engineering Using Azure Databricks Introduction This course include multiple sections. We are mainly focusing on Databricks Data Engineer certification exam. We have following tutorials: 1. Spark SQL ETL 2. Pyspark ETL DATASETS All the datasets used in the tutorials are available at: https://github.com/martandsingh/datasets HOW TO USE? follow below article to learn how to clone this repository to your databricks workspace. https://www.linkedin.com/pulse/databricks-clone-github-repo-martand-singh/ Spark SQL This course is the first installment of databricks data engineering course. In this course you will learn basic SQL concept which include: 1. Create, Select, Update, Delete tables 1. Create database 1. Filtering data 1. Group by & aggregation 1. Ordering 1. SQL joins 1. Common table expression (CTE) 1. External tables 1. Sub queries 1. Views & temp views 1. UNION, INTERSECT, EXCEPT keywords 1. Versioning, time travel & optimization PySpark ETL This course will teach you how to perform ETL pipelines using pyspark. ETL stands for Extract, Load & Transformation. We will see how to load data from various sources & process it and finally will load the process data to our destination. This course includes: 1. Read files 2. Schema handling 3. Handling JSON files 4. Write files 5. Basic transformations 6. partitioning 7. caching 8. joins 9. missing value handling 10. Data profiling 11. date time functions 12. string function 13. deduplication 14. grouping & aggregation 15. Use","default_branch":null,"files":null,"tree":[],"storefront":"/r/martandsingh","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/martandsingh/ApacheSpark/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}