{"repo":"spicyparrot/kafka_scrapy_connect","free":true,"listed":false,"github":"https://github.com/spicyparrot/kafka_scrapy_connect","clone":"git clone https://github.com/spicyparrot/kafka_scrapy_connect.git","description":"A custom library that integrates Scrapy with Kafka.","language":"Python","stars":12,"topics":["crawling-python","kafka","python","scraping-python","scrapy"],"license":"MIT","category":"data-pipelines","readme_excerpt":"🚀 Kafka Scrapy Connect Overview kafka scrapy connect is a custom Scrapy library that integrates Scrapy with Kafka. It consists of two main components: spiders and pipelines, which interact with Kafka for message consumption and item publishing. The library also comes with a custom extension that publishes log stats to a kafka topic at EoD, which allows the user to analyse offline how well the spider is performing! This project has been motivated by the great work undertaken in: https://github.com/dfdeshom/scrapy-kafka. kafka scrapy connect utilises Confluent's Kafka Python client under the hood, to provide high-level producer and consumer features. Features 1️⃣ Integration with Kafka 📈 - Enables communication between Scrapy spiders and Kafka topics for efficient data processing. - Through partitions and consumer groups message processing can be parallelised across multiple spiders! - Reduces overhead and improves throughput by giving the user the ability to consume messages in batches. 2️⃣ Customisable Settings 🛠️ - Provides flexibility through customisable configuration for both consumers and producers. 3️⃣ Error Handling 🚑 - Automatically handles network errors during crawling and publishes failed URLs to a designated output topic. 4️⃣ Serialisation Customisation 🧬 - Allows users to customize how Kafka messages are deserializsd by overriding the process kafka message method. Installation You can install kafka scrapy connect via pip: Example A full example using the kaf","default_branch":null,"files":null,"tree":[],"storefront":"/r/spicyparrot","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/spicyparrot/kafka_scrapy_connect/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}