{"repo":"vishwajeetdabholkar/eGet-Crawler-for-ai","free":true,"listed":false,"github":"https://github.com/vishwajeetdabholkar/eGet-Crawler-for-ai","clone":"git clone https://github.com/vishwajeetdabholkar/eGet-Crawler-for-ai.git","description":"Web scraping framework built for AI applications. Extract clean, structured content from any website with dynamic content handling, markdown conversion, and intelligent crawling capabilities. Perfect for RAG applications and AI training data pipelines. Features async processing, browser management, and Prometheus monitoring.","language":"Python","stars":57,"topics":["aitools","rag","web-scraping-api","web-scraping-python","knowledge-base","markdown","pdf"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"eGet - Advanced Web Scraping Framework for AI eGet is a high-performance, production-grade web scraping framework built with Python. It provides a robust API for extracting content from web pages with features like dynamic content handling, structured data extraction, and extensive customization options. eGet transforms complex websites into AI-ready content with a single API call, handling everything from JavaScript-rendered pages to dynamic content while delivering clean, structured markdown that's perfect for RAG applications. With its powerful crawling capabilities and intelligent content extraction, developers can effortlessly build comprehensive knowledge bases by turning any website into high-quality training data, making it an essential tool for teams building modern AI applications that need to understand and process web content at scale. 🚀 Features - Dynamic Content Handling : - Full JavaScript rendering support - Configurable wait conditions - Custom page interactions - Content Extraction : - Smart main content detection - Markdown conversion - HTML cleaning and formatting - Structured data extraction (JSON-LD, OpenGraph, Twitter Cards) - Performance & Reliability : - Browser resource pooling - Concurrent request handling - Rate limiting and retry mechanisms - Prometheus metrics integration - Additional Features : - Screenshot capture - Metadata extraction - Link discovery - Robots.txt compliance - Configurable crawl depth and scope 🛠️ Technology Stack - FastAPI ","default_branch":null,"files":null,"tree":[],"storefront":"/r/vishwajeetdabholkar","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/vishwajeetdabholkar/eGet-Crawler-for-ai/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}