{"repo":"amerkurev/scrapper","free":true,"listed":false,"github":"https://github.com/amerkurev/scrapper","clone":"git clone https://github.com/amerkurev/scrapper.git","description":"Web scraper with a simple REST API living in Docker and using a Headless browser and Readability.js for parsing.","language":"Python","stars":327,"topics":["crawler","readability","scraper","web-parsers","crawling","web-scraping","web-parsing","headless","crawler-python","scraping"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Scrapper Scrapper is a web scraper tool designed to download web pages and extract articles in a structured format. The application combines functionality from several open-source projects to provide an effective solution for web content extraction. Quick start Start a Scrapper instance with: Scrapper will be available at http://localhost:3000/. For more details, see Usage Demo Watch a 30-second demo reel showcasing the web interface of Scrapper. https://user-images.githubusercontent.com/28217522/225941167-633576fa-c9e2-4c63-b1fd-879be2d137fa.mp4 Features Scrapper provides the following features: - Built-in headless browser - Integrates with Playwright to handle JavaScript-heavy websites, cookie consent forms, and other interactive elements. - Read mode parsing - Uses Mozilla's Readability.js library to extract article content similar to browser \"Reader View\" functionality. - Web interface - Provides a user-friendly interface for debugging queries and experimenting with parameters. Built with the Pico CSS framework with dark theme support. - Simple REST API - Features a straightforward API requiring minimal parameters for integration. - News link extraction - Identifies and extracts links to news articles from website main pages. Additional capabilities include: - Result caching - Caches parsing results to disk for faster retrieval. - Page screenshots - Captures visual representation of pages as seen by the parser. - Session management - Configurable incognito mode or persist","default_branch":null,"files":null,"tree":[],"storefront":"/r/amerkurev","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/amerkurev/scrapper/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}