{"repo":"Sparkfetch/sparkfetch","free":true,"listed":false,"github":"https://github.com/Sparkfetch/sparkfetch","clone":"git clone https://github.com/Sparkfetch/sparkfetch.git","description":"🔥 Turn any URL into clean, structured, LLM-ready content. The open-source web fetching & extraction API.","language":"TypeScript","stars":46,"topics":["ai","api","content-extraction","data-extraction","html-to-markdown","llm","nodejs","open-source","rag","typescript"],"license":"MIT","category":"ai-agents","readme_excerpt":"⚡ SparkFetch Turn any URL into clean, structured, LLM-ready content. The open-source web fetching & content extraction API. --- What is SparkFetch? SparkFetch is a powerful, self-hostable API that crawls and extracts web content — converting messy HTML into clean Markdown, structured JSON, or plain text. Built for AI applications, RAG pipelines, research tools, and any workflow that needs reliable web data. Key capabilities: - 🌐 Scrape — Fetch any URL and get clean Markdown + metadata - 🕷️ Crawl — Recursively crawl entire websites with depth control - 🗺️ Map — Discover all URLs on a domain instantly - 🧹 Clean output — Strip nav, ads, and boilerplate automatically - 📦 Structured data — Returns JSON with title, description, links, and content - ⚡ Fast — Built on Node.js 24 + Express 5 --- Getting Started Prerequisites - Node.js 18+ - pnpm 9+ Installation The API will be available at http://localhost:5000/api . --- API Reference POST /api/v1/scrape Fetch a single URL and return its content as Markdown with metadata. Request: Response: --- POST /api/v1/crawl Crawl a website recursively and extract content from all pages. Request: Response: --- GET /api/v1/crawl/:jobId Check the status of a crawl job. Response: --- POST /api/v1/map Discover all accessible URLs on a domain. Request: Response: --- GET /api/healthz Health check endpoint. --- Use Cases Use Case How SparkFetch Helps ---------- --------------------- AI / RAG pipelines Feed clean Markdown into your LLM context Resea","default_branch":null,"files":null,"tree":[],"storefront":"/r/Sparkfetch","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Sparkfetch/sparkfetch/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}