{"repo":"MsnAmiri/DigiKlothes","free":true,"listed":false,"github":"https://github.com/MsnAmiri/DigiKlothes","clone":"git clone https://github.com/MsnAmiri/DigiKlothes.git","description":"A dataset of more than 55,000 clothing items in the digikala website and their current information, such as, name, item url, image url (+ current price, rating & discount).","language":"Jupyter Notebook","stars":16,"topics":["dataset","clothes-retrieval","data-crawling","image-retrieval","preprocessing","web-scraping","clothes","digikala","digikala-crawler","scraper"],"license":"GPL-3.0","category":"scrapers-browser-automation","readme_excerpt":"DigiKlothes A Digikala Clothes Image Dataset - Dataset is available to scholars and researchers upon request. Email me HERE. Introduction This dataset contains: - more than 55,000 clothing items from digikala.com - name, price, discount, rating, url and a list of the url for all of each item's images. Data What each folder contains: 1. raw Data: Contains everything except image urls. 2. clean Data: Contains everything. 3. clean reduced Data: Contains only names, urls, and image urls. Reproducing the data What each code file does: 1. 'category product scraper.py' searches the input category in the website and extracts the name, price, discount, rating and url of each item. 2. 'preprocess clean.ipynb' uses the csv files in the 'raw Data' folder and finds the image links for each record of the csv file. 3. 'Single product image scraper.py' extracts the image links of the input (have to be set by hand) link. 4. 'preprocess reduce.ipynb' uses the csv files in the 'clean Data' folder and drops the price, discount and rating columns.","default_branch":null,"files":null,"tree":[],"storefront":"/r/MsnAmiri","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/MsnAmiri/DigiKlothes/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}