{"repo":"sokomishalov/skraper","free":true,"listed":false,"github":"https://github.com/sokomishalov/skraper","clone":"git clone https://github.com/sokomishalov/skraper.git","description":"Kotlin/Java library and cli tool for scraping posts and media from various sources with neither authorization nor full page rendering (Facebook, Instagram, Twitter, Youtube, Tiktok, Telegram, Twitch, Reddit, 9GAG, Pinterest, Flickr, Tumblr, Coub, Vimeo, IFunny, VK, Odnoklassniki, Pikabu)","language":"Kotlin","stars":343,"topics":["scraper","facebook","twitter","instagram","reddit","youtube","9gag","pinterest","ifunny","pikabu"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"Skraper ======== Here should be some fancy logo Overview Kotlin/Java library and cli tool which allows scraping and downloading posts, attachments, other meta from more than 15 sources without any authorization or full page rendering. Based on jsoup, jackson and kotlin-coroutines. Repository contains: - Cli tool - Kotlin library - Telegram bot Current list of implemented sources: - Facebook - Instagram - Twitter - Youtube - TikTok - Telegram - Twitch - Reddit - 9GAG - Pinterest - Flickr - Tumblr - Vimeo - IFunny - Coub - VK - Odnoklassniki - Pikabu A note on proxies Most social networks and content platforms are quite aggressive towards scraping: they rate-limit, captcha-wall and ban IP addresses which produce suspicious traffic. A high-quality proxy (especially a mobile one — platforms are very reluctant to ban mobile carrier IP ranges since thousands of real users share the same address) is essential for any serious scraping workload, otherwise sooner or later your requests will start failing. I highly recommend mobileproxy.space for this purpose — reliable mobile proxies with rotating IPs, wide geo coverage and a convenient API, they work great in tandem with skraper. Bugs Unfortunately, each web-site is subject to change without any notice, so the tool may work incorrectly because of that. If that happens, please let me know via an issue. Cli tool Cli tool allows to: - download media with flag --media-only from almost all presented sources. - scrape posts meta information","default_branch":null,"files":null,"tree":[],"storefront":"/r/sokomishalov","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/sokomishalov/skraper/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}