{"repo":"dead8309/ai-rag-crawler","free":true,"listed":false,"github":"https://github.com/dead8309/ai-rag-crawler","clone":"git clone https://github.com/dead8309/ai-rag-crawler.git","description":"AI pipeline built with the honc and workers-ai. vector embeddings, web scraping and processing with Cloudflare Workflows (beta)","language":"TypeScript","stars":33,"topics":["drizzle-orm","ai","cloudflare-workers","cloudflare-workflows","hono","neon","vector-embeddings","web-scraping","workers-ai","honc"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"ai-rag-crawler --- An AI RAG pipeline built with the Hono stack and Workers AI. This project uses vectorization embeddings to enable semantic search, orchestrates web scraping and data processing with Cloudflare Workflows (beta), and generates context-aware responses using AI. [!NOTE] See How to use the frontend in live url 📝 Table of Contents - About - Api Flow - How to use the frontend in live url - Demo - Technology Stack - Setting up a local environment - Usage 🧐 About The ideal state is having a system that can effortlessly ingest documentation from any website, understand its content semantically, and provide accurate, context-aware answers to user questions. This would allow for easy exploration and utilization of vast amounts of documentation without manual effort. This project provides an automated RAG (Retrieval-Augmented Generation) pipeline built using serverless technologies. It takes a base URL of a documentation website, scrapes the site and all linked pages recursively, generates vector embeddings of the content, stores them in a database, and uses those embeddings to generate accurate, context-aware responses to questions. It uses Cloudflare Workflows(Beta) thus providing a more resilient and scalable solution to the problem. Api Flow How to use the frontend in live url Currently the database has 1 site processed properly which is So to use the frontend to ask questions: 1. Head over to https://ai-docs-rag.cjjdxhdjd.workers.dev/ 2. Enter the following url w","default_branch":null,"files":null,"tree":[],"storefront":"/r/dead8309","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/dead8309/ai-rag-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}