{"repo":"scriptin/kanji-frequency","free":true,"listed":false,"github":"https://github.com/scriptin/kanji-frequency","clone":"git clone https://github.com/scriptin/kanji-frequency.git","description":"Kanji usage frequency data collected from various sources","language":"Astro","stars":165,"topics":["kanji","data","japanese","japanese-language","data-visualization","kanji-frequency","frequency-lists","corpus","corpus-linguistics","cjk"],"license":"CC-BY-4.0","category":"scrapers-browser-automation","readme_excerpt":"Kanji usage frequency Datasets built from various Japanese language corpora - see this website for the dataset description. This readme describes only technical aspects. You can download the datasets here: Building the datasets You'll need Node.js 18 or later. See scripts section in package.json. Aozora: - aozora:download - use crawler/scraper to collect the data - aozora:gaiji:extract - extract gaiji notations data from scraped pages. Gaiji refers to kanji charasters which are replaced with images in the documents, because Shift-JIS encoding cannot represent them - aozora:gaiji:replacements - build gaiji replacements file - produces only partial results, which may need to be manually completed - aozora:clean - clean the scraped pages (apply gaiji replacements) - aozora:count - create the dataset Wikipedia: - wikipedia:fetch - fetch random pages using MediaWiki API - wikipedia:count - create the dataset News: - news:wikinews:fetch - fetch random pages from Wikinews using MediaWiki API - news:count - create the dataset - news:dates - create additional file with dates of articles Building the website See Astro docs and the scripts section in package.json.","default_branch":null,"files":null,"tree":[],"storefront":"/r/scriptin","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/scriptin/kanji-frequency/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}