{"repo":"pixlie/SmartCrawler","free":true,"listed":false,"github":"https://github.com/pixlie/SmartCrawler","clone":"git clone https://github.com/pixlie/SmartCrawler.git","description":"A smart web crawler built in Rust that uses Claude AI to select the most relevant URLs from website sitemaps based on crawling objectives.","language":"Rust","stars":20,"topics":["claude-api","claude-code","crawler","vibe-coding"],"license":"GPL-3.0","category":"scrapers-browser-automation","readme_excerpt":"SmartCrawler A web crawler that uses WebDriver to extract and parse HTML content from web pages with intelligent duplicate detection and template pattern recognition. ✨ Features - 🌐 Multi-URL Crawling : Crawl multiple URLs in a single session - 🔍 Intelligent Duplicate Detection : Automatically identifies and filters duplicate content patterns across domains - 📋 Template Pattern Recognition : Detects variable patterns in content (e.g., \"42 comments\" → \"{count} comments\") - 🌳 Structured HTML Tree : Provides filtered HTML tree view with duplicate marking - ⚡ WebDriver Integration : Uses WebDriver for dynamic content handling - 📊 Verbose Output : Detailed HTML tree analysis with filtering information 🚀 Quick Start 1. Install SmartCrawler - Download from releases or build from source 2. Set up WebDriver - Install Firefox/Chrome and corresponding WebDriver 3. Start crawling - Run SmartCrawler with your target URLs 📖 Documentation Getting Started Choose your operating system for detailed setup instructions: - Windows Setup - Complete Windows installation guide - macOS Setup - macOS installation and setup - Linux Setup - Linux installation for various distributions Usage - CLI Options - Complete command-line reference and examples Development - Development Guide - Setup, building, testing, and contributing instructions 🔧 System Requirements - Operating System : Windows 10+, macOS 10.15+, or Linux - Browser : Firefox (recommended) or Chrome - WebDriver : GeckoDriver (Firefox) ","default_branch":null,"files":null,"tree":[],"storefront":"/r/pixlie","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/pixlie/SmartCrawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}