{"repo":"goodreasonai/ScrapeServ","free":true,"listed":false,"github":"https://github.com/goodreasonai/ScrapeServ","clone":"git clone https://github.com/goodreasonai/ScrapeServ.git","description":"A self-hosted API that takes a URL and returns a file with browser screenshots.","language":"Python","stars":1180,"topics":[],"license":"MIT","category":"self-hosted-apps","readme_excerpt":"ScrapeServ: Simple URL to screenshots server You run the API as a web server on your machine, you send it a URL, and you get back the website data as a file plus screenshots of the site. Simple as. This project was made to support Abbey, an AI platform. Its author is Gordon Kamer. Please leave a star if you like the project! Some highlights: - Scrolls through the page and takes screenshots of different sections - Runs in a docker container - Browser-based (will run websites' Javascript) - Gives you the HTTP status code and headers from the first request - Automatically handles redirects - Handles download links properly - Tasks are processed in a queue with configurable memory allocation - Blocking API - Zero state or other complexity This web scraper is resource intensive but higher quality than many alternatives. Websites are scraped using Playwright, which launches a Firefox browser context for each job. Setup You should have Docker and docker compose installed. Easy (using pre-built image) A pre-built image is available for your use called usaiinc/scraper . You can use it with docker compose by creating a file called docker-compose.yml and putting the following inside it: Then you can run it by running docker compose up in the same directory as your file. See the Usage section below on how to interact with the server! Customizable (build from source) Another option is to clone the repo and build it yourself, which is also quite easy! This will also let you modify server s","default_branch":null,"files":null,"tree":[],"storefront":"/r/goodreasonai","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/goodreasonai/ScrapeServ/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}