{"owner":"ispras","github":"https://github.com/ispras","claimed":false,"inventory":[],"indexed":[{"repo":"ispras/dedoc","github":"https://github.com/ispras/dedoc","description":"Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser","language":"Python","stars":719,"topics":["doc","docx","odt","documents","excel","pdf","txt","ocr","scanned-documents","document-content-extraction"],"license":"Apache-2.0","category":"media-processing"},{"repo":"ispras/scrapy-puppeteer","github":"https://github.com/ispras/scrapy-puppeteer","description":"Library that helps use puppeteer in scrapy.","language":"Python","stars":51,"topics":["scrapy-puppeteer","middleware","spider","scrapy"],"license":"BSD-3-Clause","category":"scrapers-browser-automation"},{"repo":"ispras/scrapy-puppeteer-service","github":"https://github.com/ispras/scrapy-puppeteer-service","description":"A special service that runs puppeteer instances.","language":"JavaScript","stars":18,"topics":["scrapy","puppeteer","selector"],"license":"BSD-3-Clause","category":"scrapers-browser-automation"}],"how_to_buy":"GET /r/ispras/<repo> (Accept: application/json) for any listed repo here: tree, README, price and the checkout to pay (x402; rehearse first at its test twin, simulated money). Repos under 'indexed' are free: clone them from GitHub."}