{"repo":"ispras/dedoc","free":true,"listed":false,"github":"https://github.com/ispras/dedoc","clone":"git clone https://github.com/ispras/dedoc.git","description":"Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser","language":"Python","stars":719,"topics":["doc","docx","odt","documents","excel","pdf","txt","ocr","scanned-documents","document-content-extraction"],"license":"Apache-2.0","category":"media-processing","readme_excerpt":"Dedoc Dedoc is an open universal system for converting documents to a unified output format. It extracts a document’s logical structure and content: tables, text formatting and metadata. The document’s content is represented as a tree storing headings and lists of any level. Dedoc can be integrated in a document contents and structure analysis system as a separate module. Workflow Workflow description is given here Features and advantages Dedoc is implemented in Python and works with semi-structured data formats (DOC/DOCX, ODT, XLS/XLSX, CSV, TXT, JSON) and unstructured data formats like images (PNG, JPG etc.), archives (ZIP, RAR etc.), PDF and HTML formats. Document structure extraction is fully automatic regardless of input data type. Metadata and text formatting are also extracted automatically. In 2022, the system won a grant to support the development of promising AI projects from the Innovation Assistance Foundation (Фонд содействия инновациям). Dedoc provides: Extensibility due to flexible addition of new document formats and easy change of an output data format. Support for extracting document structure out of nested documents having different formats. Extracting various text formatting features (indentation, font type, size, style etc.). Working with documents of various origin (statements of work, legal documents, technical reports, scientific papers) allowing flexible tuning for new domains. Working with PDF documents containing a textual layer: Support to automati","default_branch":null,"files":null,"tree":[],"storefront":"/r/ispras","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/ispras/dedoc/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}