{"repo":"jorvik-io/jorvik","free":true,"listed":false,"github":"https://github.com/jorvik-io/jorvik","clone":"git clone https://github.com/jorvik-io/jorvik.git","description":"Jorvik is a collection of utilities for creating and managing ETL pipelines in pyspark.","language":"Python","stars":13,"topics":["data-engineering","databricks","etl","pyspark"],"license":"Apache-2.0","category":"data-pipelines","readme_excerpt":"Jorvik Jorvik is a collection of utilities for creating and managing ETL pipeline in Pyspark. Build from Data Engineers for Data Engineers. Contribute The Jorvik project welcomes your expertise and enthusiasm! Writing code isn’t the only way to contribute. You can also: - review pull requests - suggest improvements through issues - let us know your pain-points and repetitive tasks - help us stay on top of new and old issues - develop tutorials, videos, presentations, and other educational materials See How to Contribute for instructions on setting up your local machine and opening your first Pull Request. Getting Started. Jorvik is available in Pypi and can be installed with pip Packages: - Storage: Interact with the storage layer - Pipelines: Build and test etl pipelines with ease - Data Lineage: Track data lineage Examples: See the full power of jorvik when all the features come together in the examples bellow: Databricks - Transactions: A multi step pipeline that creates customer statistics from customers and transaction data.","default_branch":null,"files":null,"tree":[],"storefront":"/r/jorvik-io","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/jorvik-io/jorvik/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}