{"repo":"ENDEVSOLS/LongParser","free":true,"listed":false,"github":"https://github.com/ENDEVSOLS/LongParser","clone":"git clone https://github.com/ENDEVSOLS/LongParser.git","description":"Privacy-first document intelligence engine — parse PDFs, DOCX, PPTX, XLSX & CSV into AI-ready chunks for RAG pipelines. Includes HITL review, 3-layer memory chat, and a production FastAPI server.","language":"Python","stars":30,"topics":["chunking","docling","document-intelligence","document-parsing","fastapi","human-in-the-loop","langchain","llm","ocr","openai"],"license":null,"category":"ai-agents","readme_excerpt":"Privacy-first document intelligence engine for production RAG pipelines. Parse PDFs, DOCX, PPTX, XLSX &amp; CSV → validated, AI-ready chunks with HITL review. --- Features Feature Detail --------- -------- Multi-format extraction PDF, DOCX, PPTX, XLSX, CSV via Docling & Marker Hybrid chunking Token-aware, heading-hierarchy-aware, table-aware Semantic chunking Embedding-based boundaries using all-MiniLM-L6-v2 Cross-referencing Deterministic linking of explicit and implicit charts/figures Quality scoring Zero-ML heuristic scoring with dictionary & fastText validation PII redaction Hybrid Regex + NER (spaCy) redaction with secure HITL preservation Summary chunks Async ARQ worker generating hierarchical LLM section summaries HITL review Human-in-the-Loop block & chunk editing before embedding LangGraph HITL approve / edit / reject workflow with LangGraph interrupt() and MongoDB checkpointer 3-layer memory Short-term turns + rolling summary + long-term facts Multi-provider LLM OpenAI, Gemini, Groq, OpenRouter Multi-backend vectors Chroma, FAISS, Qdrant Production-ready API FastAPI + Motor (MongoDB) + ARQ + Redis (Queue & Rate Limiting) Enterprise Security Tenant isolation, Role-Based Access Control (RBAC), and CORS LangChain adapters Drop-in BaseRetriever and LlamaIndex QueryEngine Privacy-first All processing runs locally; no data leaves your infra --- Installation Quick install (recommended) Includes everything — server, embeddings, vector DB, OCR, LangChain, LlamaIndex. Works o","default_branch":null,"files":null,"tree":[],"storefront":"/r/ENDEVSOLS","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/ENDEVSOLS/LongParser/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}