{"repo":"prajwal10001/semantic-chunker-langchain","free":true,"listed":false,"github":"https://github.com/prajwal10001/semantic-chunker-langchain","clone":"git clone https://github.com/prajwal10001/semantic-chunker-langchain.git","description":"Token-aware, LangChain-compatible semantic chunker with PDF, markdown, and layout support","language":"Python","stars":13,"topics":["ai","langchain","markdown","nlp","pdf","python","rag","semantic-chunking"],"license":"MIT","category":"ai-agents","readme_excerpt":"Semantic Chunker for LangChain Hitting limits on passing the larger context to your limited character token limit llm model not anymore this chunker solves the problem It is a token-aware , LangChain-compatible chunker that splits text (from PDF, markdown, or plain text) into semantically coherent chunks while respecting model token limits. --- 🚀 Features 🔍 Model-Aware Token Limits : Automatically adjusts chunking size for GPT-3.5, GPT-4, Claude, and others. 📄 Multi-format Input Support : PDF via pdfplumber Plain .txt Markdown (Extendable to .docx and .html ) 🔁 Overlapping Chunks : Smart overlap between paragraphs to preserve context. 🧠 Smart Merging : Merges chunks smaller than 300 tokens. 🧩 Retriever-Ready : Direct integration with LangChain retrievers via FAISS. 🔧 CLI Support : Run from terminal with one command. --- 📆 Installation Requires Python 3.9 - 3.12 --- 🛠️ Usage 🔸 Chunk a PDF and Save to JSON/TXT 🔸 From Code 🔸 Convert to Retriever --- 📊 Testing --- 👨‍💻 Authors Prajwal Shivaji Mandale Sudhnwa Ghorpade --- 📜 License This project is licensed under the MIT License.","default_branch":null,"files":null,"tree":[],"storefront":"/r/prajwal10001","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/prajwal10001/semantic-chunker-langchain/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}