{"repo":"GramosoftAI/GcrawlAI","free":true,"listed":false,"github":"https://github.com/GramosoftAI/GcrawlAI","clone":"git clone https://github.com/GramosoftAI/GcrawlAI.git","description":"Turn any website into clean, LLM-ready data. Open-source web crawler with stealth mode, distributed crawling, real-time WebSocket progress & Markdown output. Power your AI apps with GcrawlAI.","language":"Python","stars":38,"topics":["ai","celery","data-pipeline","document-extraction","fastapi","llm","llm-tools","markdown","nlp","open-source"],"license":"MIT","category":"ai-agents","readme_excerpt":"--- 🤔 Why GcrawlAI? GcrawlAI is a high-performance, enterprise-grade, distributed web crawler, scraper, and extraction platform. Designed to feed retrieval-augmented generation (RAG) pipelines, LLMs, and semantic search indexes, it converts complex, noisy web structures into clean Markdown, structured JSON metadata, and full-page screenshots. GcrawlAI automates browser steering, stealth obfuscation, anti-bot evasion, and distributed scaling so that you can focus on building AI features rather than managing crawling blockages. --- ✨ Features - 🥷 Fingerprint Hygiene & Stealth Browsing : Mask automated runtimes, WebGL signatures, canvas fingerprints, and automation leaks to seamlessly bypass aggressive anti-bot protections. - 🔀 Stepped Residential Proxy Rotation : Multi-tier automatic proxy escalation with geographic IP targeting matching the target site's local region. - ✨ Fit-Markdown Extraction : Converts pages to clean, LLM-ready markdown (pruning HTML boilerplate, menus, footers, and advertisements). - 💾 Offline HTML Bundle : Downloads full pages along with CSS, images, and other assets, packaging them into a single ZIP file for local offline rendering. - 📊 SEO Data Collection : Automatically extracts metadata, headers, titles, descriptions, open graph tags, and links structure from crawled pages. - 📸 High-Resolution Screenshotting & Document Parsing : Physics-based scrolling to capture lazy-loaded content correctly. - 🗺️ URL Mapping : /links endpoint discovers sitem","default_branch":null,"files":null,"tree":[],"storefront":"/r/GramosoftAI","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/GramosoftAI/GcrawlAI/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}