{"repo":"ivan-sincek/scrapy-scraper","free":true,"listed":false,"github":"https://github.com/ivan-sincek/scrapy-scraper","clone":"git clone https://github.com/ivan-sincek/scrapy-scraper.git","description":"Web crawler and scraper based on Scrapy and Playwright's headless browser.","language":"Python","stars":17,"topics":["bug-bounty","crawler","crawling","headless-browser","offensive-security","python","scraper","scraping","scrapy","security"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Scrapy Scraper Web crawler and scraper based on Scrapy and Playwright's headless browser. To use the headless browser specify -p option. Browsers, unlike other standard web request libraries, have the ability to render JavaScript encoded HTML content. Future plans: check if Playwright's Chromium headless browser is installed, add option to stop on rate limiting. Resources: docs.scrapy.org - docs playwright.dev - docs scrapy/scrapy - GitHub scrapy-plugins/scrapy-playwright - GitHub Tested on Kali Linux v2024.2 (64-bit). Made for educational purposes. I hope it will help! Table of Contents How to Install Install Playwright and Chromium Standard Install Build and Install From the Source How to Run Usage Images How to Install Install Playwright and Chromium Make sure each time you upgrade your Playwright dependency to re-install Chromium; otherwise, you might get an error using the headless browser. Standard Install Build and Install From the Source How to Run Example, start in-scope crawling from https://example.com/home , download in-scope JavaScript files, and extract links: Example, start in-scope crawling from URLs specified in urls.txt , take a screenshot of only the start URLs, and extract links: Usage Images Figure 1 - Scraping","default_branch":null,"files":null,"tree":[],"storefront":"/r/ivan-sincek","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/ivan-sincek/scrapy-scraper/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}