{"repo":"google/corpuscrawler","free":true,"listed":false,"github":"https://github.com/google/corpuscrawler","clone":"git clone https://github.com/google/corpuscrawler.git","description":"Crawler for linguistic corpora","language":"Python","stars":218,"topics":["corpus-linguistics","corpus-builder","crawling","linguistics","minority-language"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"Corpus Crawler Corpus Crawler is a tool for Corpus Linguistics. Modern linguistic research works on language corpora, which are large samples of “real world” text. This crawler helps to build such corpora: it follows links to publicly accessible web pages known to be written in a certain language; it removes boilerplate and HTML markup; finally, it writes its output into plaintext files. The crawler implements the Robots Exclusion Standard, and it is intentionally slow so it does not cause much load on the crawled web sites. This is not an official Google product. But if you’re a linguistic researcher, or if you’re writing a spell checker (or similar language-processing software) for an “exotic” language, you might find Corpus Crawler useful. To build corpora for not-yet-supported languages, please read the contribution guidelines and send us GitHub pull requests. The crawled corpora have been used to compute word frequencies in Unicode’s Unilex project. Supported Languages IETF BCP47 Code Language Tokens¹ :------------------ :--------------------------- ----------------------------------------------------------------------------------: aai Arifama-Miniafia 181K 💾 aak Ankave 194K 💾 aau Abau 313K 💾 aaz Amarasi 308K 💾 abt Ambulas 297K 💾 aby Aneme Wake 233K 💾 acd Gikyode 323K 💾 ace Aceh/Acehnese 817K 💾 acf Saint Lucian Creole French 236K 💾 ach Acoli 178K 💾 acn Achang 232K 💾 acr Achi 239K 💾 acu Achuar-Shiwiar 174K 💾 ade Adele 267K 💾 adh Adhola 166K 💾 adj Adioukrou ","default_branch":null,"files":null,"tree":[],"storefront":"/r/google","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/google/corpuscrawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}