{"repo":"BjornMelin/ai-docs-vector-db-hybrid-scraper","free":true,"listed":false,"github":"https://github.com/BjornMelin/ai-docs-vector-db-hybrid-scraper","clone":"git clone https://github.com/BjornMelin/ai-docs-vector-db-hybrid-scraper.git","description":"Retrieval-augmented docs ingestion stack: Firecrawl + Crawl4AI + Qdrant vector search with FastAPI and MCP interfaces for AI engineers.","language":"Python","stars":10,"topics":["ai","crawl4ai","fastapi","firecrawl","mcp","qdrant","rag","retrieval-augmented-generation","vector-search","documentation-ingestion"],"license":"MIT","category":"ai-agents","readme_excerpt":"AI Documentation Vector Database Hybrid Scraper AI-focused documentation ingestion and retrieval stack that combines Firecrawl and Crawl4AI powered scraping with a Qdrant vector database. The project exposes both FastAPI and MCP interfaces, offers mode-aware configuration (solo developer vs enterprise feature sets), and ships with tooling for embeddings, hybrid search, retrieval-augmented generation (RAG) workflows, and operational monitoring. Overview The system ingests documentation sources, generates dense and sparse embeddings, stores them in Qdrant, and serves hybrid search and RAG building blocks. It is built for AI engineers who need reliable documentation ingestion pipelines, reproducible retrieval quality, and integration points for agents or applications. Highlights - Multi-tier crawling orchestration ( src/services/browser/unified manager.py ) covering lightweight HTTP, Crawl4AI, browser-use, Playwright, and Firecrawl, plus a resumable bulk embedder CLI ( src/crawl4ai bulk embedder.py ). - Hybrid retrieval stack leveraging OpenAI and LangChain FastEmbed dense embeddings plus FastEmbedSparse BM25 signals, reranking, and HyDE augmentation through the modular Qdrant service ( src/services/vector db/ and src/services/hyde/ ). - Dual interfaces: REST endpoints in FastAPI ( src/api/routers/v1/ ) and a FastMCP server ( src/unified mcp server.py ) that registers search, document management, analytics, and content intelligence tools for Claude Desktop / Code. - Built-in API","default_branch":null,"files":null,"tree":[],"storefront":"/r/BjornMelin","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/BjornMelin/ai-docs-vector-db-hybrid-scraper/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}