{"repo":"watercrawl/WaterCrawl","free":true,"listed":false,"github":"https://github.com/watercrawl/WaterCrawl","clone":"git clone https://github.com/watercrawl/WaterCrawl.git","description":"Transform Web Content into LLM-Ready Data","language":"TypeScript","stars":2010,"topics":["crawl4ai","crawler","crawling-python","html2markdown","llm-crawler","llm-scraper","scraper","aicrawler"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"🕷️ WaterCrawl is a powerful web application that uses Python, Django, Scrapy, and Celery to crawl web pages and extract relevant data. 🚀 Quick Start 1. 🐳 Quick start 2. 💻 Development (For Contributing) 🐳 Quick start To build and run WaterCrawl on Docker locally, please follow these steps: 1. Clone the repository: 2. Build and run the Docker containers: 3. Access the application with open http://localhost ⚠️ IMPORTANT : If you're deploying on a domain or IP address other than localhost, you MUST update the MinIO configuration in your .env file: Failure to update these settings will result in broken file uploads and downloads. For more details, see DEPLOYMENT.md. Important: Before deploying to production, ensure that you update the .env file with the appropriate configuration values. Additionally, make sure to set up and configure the database, MinIO, and any other required services. for more information, please read the Deployment Guide. 💻 Development (For Contributing) For local development and contribution, please follow our Contributing Guide 🤝 ✨ Features - 🕸️ Advanced Web Crawling & Scraping - Crawl websites with highly customizable options for depth, speed, and targeting specific content - 🔍 Powerful Search Engine - Find relevant content across the web with multiple search depths (basic, advanced, ultimate) - 🌐 Multi-language Support - Search and crawl content in different languages with country-specific targeting - ⚡ Asynchronous Processing - Monitor real-time ","default_branch":null,"files":null,"tree":[],"storefront":"/r/watercrawl","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/watercrawl/WaterCrawl/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}