{"repo":"roshanlam/Spider","free":true,"listed":false,"github":"https://github.com/roshanlam/Spider","clone":"git clone https://github.com/roshanlam/Spider.git","description":"Web Crawler built using asynchronous Python and distributed task management that extracts and saves web data for analysis.","language":"Python","stars":34,"topics":["hacktoberfest","hacktoberfest-accepted","spider","web-crawler","web-crawler-python"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"🕷️ Spider A modern, scalable, and extensible web crawler for efficient distributed crawling and data extraction Built with asynchronous I/O, plugin architecture, and distributed task processing --- ✨ Features Core Capabilities - 🚀 Asynchronous Crawling - Non-blocking I/O with aiohttp and asyncio for high performance - 🌐 Distributed Processing - Scale across multiple workers using Celery and Redis - 💾 Database Persistence - PostgreSQL storage with SQLAlchemy ORM - 🔌 Plugin Architecture - Extensible system for custom data processing - 📊 Robust Logging - Console, file, and database logging for diagnostics - 🔗 URL Normalization - Smart deduplication and link management Included Plugins - 🕸️ Web Scraper - Comprehensive webpage data extraction - 📝 Title Logger - Extract and store page titles - 🤖 Entity Extraction - NLP-based named entity recognition (spaCy) - 🎭 Dynamic Scraper - JavaScript-rendered pages (Playwright) - 📈 Real-time Metrics - Live crawl statistics via WebSocket --- 🚀 Quick Start Prerequisites - Python 3.11+ - PostgreSQL - Redis Installation Configuration Edit src/spider/config.yaml : Run the Crawler Simple way: Or using module: Query Scraped Data --- 📦 Project Structure --- 🔌 Plugin System Spider uses a powerful plugin architecture for extensibility. Using the Web Scraper Plugin The comprehensive web scraper extracts structured data from every page: What gets extracted: - Metadata (title, description, keywords, author, language) - Content structure (he","default_branch":null,"files":null,"tree":[],"storefront":"/r/roshanlam","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/roshanlam/Spider/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}