{"repo":"aws-samples/aws-glue-streaming-etl-with-apache-iceberg","free":true,"listed":false,"github":"https://github.com/aws-samples/aws-glue-streaming-etl-with-apache-iceberg","clone":"git clone https://github.com/aws-samples/aws-glue-streaming-etl-with-apache-iceberg.git","description":"Streaming ETL job cases in AWS Glue to integrate Iceberg and creating an in-place updatable data lake on Amazon S3","language":"Python","stars":27,"topics":["aws-glue","apache-iceberg","aws-athena","apache-spark","aws-glue-streaming"],"license":"MIT-0","category":"data-pipelines","readme_excerpt":"AWS Glue Streaming ETL Job with Apace Iceberg CDK Python project! In this project, we create a streaming ETL job in AWS Glue to integrate Iceberg with a streaming use case and create an in-place updatable data lake on Amazon S3. After ingested to Amazon S3, you can query the data with Amazon Athena. This project can be deployed with AWS CDK Python. The cdk.json file tells the CDK Toolkit how to execute your app. This project is set up like a standard Python project. The initialization process also creates a virtualenv within this project, stored under the .venv directory. To create the virtualenv it assumes that there is a python3 (or python for Windows) executable in your path with access to the venv package. If for any reason the automatic creation of the virtualenv fails, you can create the virtualenv manually. To manually create a virtualenv on MacOS and Linux: After the init process completes and the virtualenv is created, you can use the following step to activate your virtualenv. If you are a Windows platform, you would activate the virtualenv like this: Once the virtualenv is activated, you can install the required dependencies. In case of AWS Glue 3.0 , before synthesizing the CloudFormation, you first set up Apache Iceberg connector for AWS Glue to use Apache Iceber with AWS Glue jobs. (For more information, see References (2)) Then you should set approperly the cdk context configuration file, cdk.context.json . For example: { \"kinesis stream name\": \"iceberg-demo-st","default_branch":null,"files":null,"tree":[],"storefront":"/r/aws-samples","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/aws-samples/aws-glue-streaming-etl-with-apache-iceberg/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}