{"repo":"jxlil/scrapy-impersonate","free":true,"listed":false,"github":"https://github.com/jxlil/scrapy-impersonate","clone":"git clone https://github.com/jxlil/scrapy-impersonate.git","description":"Scrapy download handler that can impersonate browser' TLS signatures or JA3 fingerprints.","language":"Python","stars":239,"topics":["scrapy","scrapy-plugin","scrapy-impersonate","ja3","ja3-fingerprint","tls-fingerprint"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"scrapy-impersonate scrapy-impersonate is a Scrapy download handler. This project integrates curl cffi to perform HTTP requests, so it can impersonate browsers' TLS signatures or JA3 fingerprints. Installation Activation To use this package, replace the default http and https Download Handlers by updating the DOWNLOAD HANDLERS setting: By setting USER AGENT = None , curl cffi will automatically choose the appropriate User-Agent based on the impersonated browser: Also, be sure to install the asyncio-based Twisted reactor for proper asynchronous execution: Usage Set the impersonate Request.meta key to download a request using curl cffi : impersonate-args You can pass any necessary arguments to curl cffi through impersonate args . For example: Some arguments worth knowing about: Argument Description --- --- http version Set to \"v3\" to use HTTP/3. Targets with an HTTP/3 fingerprint are marked in the table below doh url Resolve DNS over HTTPS instead of using the system resolver ( curl cffi = 0.16.0 ) interface Bind the request to a network interface or local source address extra fp Fine-tune fingerprint details on top of the impersonated browser [!WARNING] allow redirects is forced to False so that redirects are handled by Scrapy. Re-enabling it through impersonate args makes curl cffi follow redirects internally, which bypasses Scrapy's redirect and offsite middlewares and exposes you to redirect-based SSRF. curl-options For settings that have no curl cffi argument, impersonate c","default_branch":null,"files":null,"tree":[],"storefront":"/r/jxlil","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/jxlil/scrapy-impersonate/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}