{"repo":"TUVIMEN/forumscraper","free":true,"listed":false,"github":"https://github.com/TUVIMEN/forumscraper","clone":"git clone https://github.com/TUVIMEN/forumscraper.git","description":"automatic and extensive scraper for forums","language":"Python","stars":53,"topics":["forum","hackernews","invision-power-board","json","phpbb","python","reliq","scraper","simple-machines-forum","stackexchange"],"license":"GPL-3.0","category":"scrapers-browser-automation","readme_excerpt":"forumscraper forumscraper aims to be an universal, automatic and extensive scraper for forums. Installation pip install forumscraper Supported forums - Invision Power Board (only 4.x and 5.x versions) - PhpBB (currently excluding 1.x version) - Simple Machines Forum - XenForo - XMB - Hacker News (has aggressive protection) - StackExchange - vBulletin (3.x and higher) Output examples Are created by create-format-examples script and contained in examples directory where they're grouped based on scraper and version. Files are in json format. Discover You can find forums discoverer script in discover. Usage CLI General Download any kind of supported forums from URL s into DIR , creating json files for threads named by their id's, same for users but beginning with m- e.g. 24 29 m-89 m-125 . forumscraper --directory DIR URL1 URL2 URL3 Above behaviour is set by default --names id , and can be changed with --names hash which names files by sha256 sum of their source urls. forumscraper --names hash --directory DIR URL By default if files to be created are found and are not empty, function exits not overwriting them. This can be changed using --force option. forumscraper output logging information to stdout (can be changed with --log FILE ) and information about failures to stderr (can be changed with --failed FILE ) Failures are generally ignored but setting --pedantic flag stops the execution if any failure is encountered. Download URL s into DIR using 8 threads and log failures into","default_branch":null,"files":null,"tree":[],"storefront":"/r/TUVIMEN","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/TUVIMEN/forumscraper/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}