{"repo":"sodadata/soda-spark","free":true,"listed":false,"github":"https://github.com/sodadata/soda-spark","clone":"git clone https://github.com/sodadata/soda-spark.git","description":"Soda Spark is a PySpark library that helps you with testing your data in Spark Dataframes","language":"Python","stars":64,"topics":["spark","pyspark","data-engineering","data-quality","data-observability","data-testing","soda-sql","python"],"license":"Apache-2.0","category":"data-pipelines","readme_excerpt":"Soda Spark Data testing, monitoring, and profiling for Spark Dataframes. Soda Spark is an extension of Soda SQL that allows you to run Soda SQL functionality programmatically on a Spark data frame. Soda SQL is an open-source command-line tool. It utilizes user-defined input to prepare SQL queries that run tests on tables in a data warehouse to find invalid, missing, or unexpected data. When tests fail, they surface \"bad\" data that you can fix to ensure that downstream analysts are using \"good\" data to make decisions. Requirements Soda Spark has the same requirements as soda-sql-spark . Install From your shell, execute the following command. Use From your Python prompt, execute the following commands. Or, use a scan YAML file See the scan result object for all attributes and methods. Or, return Spark data frames: See the to data frame functions in the scan.py to see how the conversion is done. Send results to Soda cloud Send the scan result to Soda cloud. Understand Under the hood soda-spark does the following. 1. Setup the scan Use the Spark dialect Use Spark session as warehouse connection 2. Create (or replace) global temporary view for the Spark data frame 3. Execute the scan on the temporary view","default_branch":null,"files":null,"tree":[],"storefront":"/r/sodadata","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/sodadata/soda-spark/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}