{"repo":"MikeMeliz/TorCrawl.py","free":true,"listed":false,"github":"https://github.com/MikeMeliz/TorCrawl.py","clone":"git clone https://github.com/MikeMeliz/TorCrawl.py.git","description":"Crawl and extract (regular or onion) webpages through TOR network","language":"Python","stars":532,"topics":["crawler","onion","python","tor","extractor","osint"],"license":"GPL-3.0","category":"scrapers-browser-automation","readme_excerpt":"### TorCrawl.py is a Python script designed for anonymous web scraping via the Tor network. It combines ease of use with the robust privacy features of Tor, allowing for secure and untraceable data collection. Ideal for both novice and experienced programmers, this tool is essential for responsible data gathering in the digital age. [![Release][release-version-shield]][releases-link] [![Last Commit][last-commit-shield]][commit-link] ![Python][python-version-shield] [![Quality Gate Status][quality-gate-shield]][quality-gate-link] [![license][license-shield]][license-link] What makes it simple and easy to use? If you are a terminal maniac you know that things have to be simple and clear. Passing the output into other tools is necessary and accuracy is the key. With a single argument, you can read an .onion webpage or a regular one, through TOR Network and by using pipes you can pass the output at any other tool you prefer. If you want to crawl the links of a webpage use the -c and BAM you got on a file all the inside links. You can even use -d to crawl them and so on. You can also use the argument -p to wait some seconds before the next crawl. [!TIP] Crawling is not illegal, but violating copyright is . It’s always best to double-check a website’s T&C before start crawling them. Some websites set up what’s called robots.txt to tell crawlers not to visit those pages. This crawler will allow you to go around this, but we always recommend respecting robots.txt. Installation Easy I","default_branch":null,"files":null,"tree":[],"storefront":"/r/MikeMeliz","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/MikeMeliz/TorCrawl.py/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}