{"repo":"adidas/lakehouse-engine","free":true,"listed":false,"github":"https://github.com/adidas/lakehouse-engine","clone":"git clone https://github.com/adidas/lakehouse-engine.git","description":"The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.","language":"Python","stars":294,"topics":["big-data","configuration-driven","data-engineering","data-quality","databricks","delta-lake","framework","great-expectations","lakehouse","spark"],"license":"Apache-2.0","category":"data-pipelines","readme_excerpt":"Lakehouse Engine A configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products. --- Note: whenever you read Data Product or Data Product team, we want to refer to Teams and use cases, whose main focus is on leveraging the power of data, on a particular topic, end-to-end (ingestion, consumption...) to achieve insights, supporting faster and better decisions, which generate value for their businesses. These Teams should not be focusing on building reusable frameworks, but on re-using the existing frameworks to achieve their goals. --- Main Goals The goal of the Lakehouse Engine is to bring some advantages, such as: - offer cutting-edge, standard, governed and battle-tested foundations that several Data Product teams can benefit from; - avoid that Data Product teams develop siloed solutions, reducing technical debts and high operating costs (redundant developments across teams); - allow Data Product teams to focus mostly on data-related tasks, avoiding wasting time & resources on developing the same code for different use cases; - benefit from the fact that many teams are reusing the same code, which increases the likelihood that common issues are surfaced and solved faster; - decrease the dependency and learning curve to Spark and other technologies that the Lakehouse Engine abstracts; - speed up repetitive tasks; - reduced vendor lock-in. --- Note: even though","default_branch":null,"files":null,"tree":[],"storefront":"/r/adidas","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/adidas/lakehouse-engine/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}