{"repo":"dbashford/textract","free":true,"listed":false,"github":"https://github.com/dbashford/textract","clone":"git clone https://github.com/dbashford/textract.git","description":"node.js module for extracting text from html, pdf, doc, docx, xls, xlsx, csv, pptx, png, jpg, gif, rtf and more!","language":"HTML","stars":1694,"topics":["extract-text","extraction","nodejs"],"license":"MIT","category":"dev-tools","readme_excerpt":"textract ======== A text extraction node module. Currently Extracts... HTML, HTM ATOM, RSS Markdown EPUB XML, XSL PDF DOC, DOCX ODT, OTT (experimental, feedback needed!) RTF XLS, XLSX, XLSB, XLSM, XLTX CSV ODS, OTS PPTX, POTX ODP, OTP ODG, OTG PNG, JPG, GIF DXF application/javascript All text/ mime-types. In almost all cases above, what textract cares about is the mime type. So .html and .htm , both possessing the same mime type, will be extracted. Other extensions that share mime types with those above should also extract successfully. For example, application/vnd.ms-excel is the mime type for .xls , but also for 5 other file types. Does textract not extract from files of the type you need? Add an issue or submit a pull request. It many cases textract is already capable, it is just not paying attention to the mime type you may be interested in. Install Extraction Requirements Note, if any of the requirements below are missing, textract will run and extract all files for types it is capable. Not having these items installed does not prevent you from using textract, it just prevents you from extracting those specific files. PDF extraction requires pdftotext be installed, link DOC extraction requires antiword be installed, link, unless on OSX in which case textutil (installed by default) is used. RTF extraction requires unrtf be installed, link, unless on OSX in which case textutil (installed by default) is used. PNG , JPG and GIF require tesseract to be available, link. Images","default_branch":null,"files":null,"tree":[],"storefront":"/r/dbashford","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/dbashford/textract/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}