{"repo":"alireza-heidarii/Real-Time-Data-Cleaning-Pipeline-for-Medical-and-Healthcare-Data","free":true,"listed":false,"github":"https://github.com/alireza-heidarii/Real-Time-Data-Cleaning-Pipeline-for-Medical-and-Healthcare-Data","clone":"git clone https://github.com/alireza-heidarii/Real-Time-Data-Cleaning-Pipeline-for-Medical-and-Healthcare-Data.git","description":"A real-time data cleaning pipeline for medical and healthcare data using Apache Spark, SparkNLP, Spark Streaming, and Kafka.","language":"Python","stars":13,"topics":["data-cleaning","data-pipelines","data-preprocessing","healthcare-datasets","kafka","medical-data-analysis","natural-language-processing","parquet","pyspark","python"],"license":"Apache-2.0","category":"data-pipelines","readme_excerpt":"Realtime Data Processing Pipeline using SparkNLP Spark-Streaming and Kafka Project Overview This project implements a real-time data processing pipeline using Apache Spark , Kafka , and Spark NLP . The pipeline ingests streaming data from Kafka, processes it using NLP techniques to extract meaningful insights, and stores the cleaned data in Parquet format for further analysis. Features - Real-time Data Ingestion : Consumes streaming data from Kafka topics. - Data Processing with Spark NLP : - Cleans HTML content - Extracts metadata (titles, paragraphs) - Detects Named Entities (NER) - Redacts sensitive information (PII) - Batch Processing and Storage : Saves processed data in Parquet format for later use. - Scalable Architecture : Built using Apache Spark , making it suitable for large-scale processing. Project Structure Technologies Used - Apache Spark (Structured Streaming) - Spark NLP (Text Processing & NER) - Apache Kafka (Streaming Data Source) - PySpark - Parquet (Data Storage) Setup & Installation 1. Clone the Repository 2. Install Dependencies Ensure you have Python installed and set up a virtual environment: 3. Run Kafka & Spark Ensure you have Kafka and Spark running. If using Docker, you can set up a docker-compose.yml file to start the services. Configuration Modify config.py to update settings like Kafka topics, Spark master URL, and storage paths. Example: Start the Pipeline How It Works 1. Kafka Producer sends raw HTML content to a Kafka topic ( data-pipeline )","default_branch":null,"files":null,"tree":[],"storefront":"/r/alireza-heidarii","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/alireza-heidarii/Real-Time-Data-Cleaning-Pipeline-for-Medical-and-Healthcare-Data/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}