{"repo":"AvinashThimmareddy/data-contract-execution-engine","free":true,"listed":false,"github":"https://github.com/AvinashThimmareddy/data-contract-execution-engine","clone":"git clone https://github.com/AvinashThimmareddy/data-contract-execution-engine.git","description":"DCEE is a lightweight Python framework for validating data against contracts and enforcing SLA rules. Built on pandas and boto3, it provides simple, fast data validation without heavy dependencies.","language":"Python","stars":19,"topics":["aws","aws-lambda","data-contracts","data-engineering","data-ingestion","data-lake","data-pipelines","data-validation","runtime-enforcement","schema-validation"],"license":"Apache-2.0","category":"data-pipelines","readme_excerpt":"Data Contract Execution Engine Overview Data Contract Execution Engine (DCEE) is a lightweight Python framework for validating data against contracts and enforcing SLA rules. Built on pandas and boto3, it provides simple, fast data validation without heavy dependencies. It reads YAML contract definitions that specify: - Schema requirements (column names, types, nullability) - SLA rules (min/max rows, data completeness) - Source and target S3 paths for data files The engine validates data on ingestion to ensure quality and consistency before writing to target destinations. Deploy as a Lambda function for serverless validation pipelines. --- Key Features - Simple Python implementation without heavy dependencies - YAML-based contract definitions for schemas and SLA rules - Data quality validation including schema checks and SLA enforcement - AWS Lambda deployment for serverless processing - S3 integration for reading and writing data - Pandas-based data manipulation --- Limitations - CSV files only - Currently supports CSV format - No support for TXT files with custom delimiters (tab, pipe, space, etc.) - No support for other formats (JSON, Parquet, Excel, etc.) - See CONTRIBUTING.md for how to add multi-format support - Single-machine processing - Best for datasets under 1GB - Loads entire dataset into memory - For larger files, consider using Spark for distributed processing --- Installation 1. Clone the repository: 2. Install dependencies: --- Quick Start See QUICKSTART.md fo","default_branch":null,"files":null,"tree":[],"storefront":"/r/AvinashThimmareddy","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/AvinashThimmareddy/data-contract-execution-engine/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}