{"repo":"xberg-io/crawlberg","free":true,"listed":false,"github":"https://github.com/xberg-io/crawlberg","clone":"git clone https://github.com/xberg-io/crawlberg.git","description":"High-performance web crawling engine with bindings for 11 languages","language":"Rust","stars":156,"topics":["crawling","csharp","elixir","ffi","golang","java","mcp","php","python","ruby"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Crawlberg Turn any website into clean, structured data. Point Crawlberg at a URL and get back Markdown, metadata, and links — from a single page or a whole site — in the language you already use. What and Why? You need data that lives on the web, and raw HTML is not it. Crawlberg does the crawling, scraping, and cleanup end-to-end: it fetches pages, follows links, converts each one to Markdown, and hands you structured metadata (titles, links, images, social-card and JSON-LD data) — so you skip the parsing and go straight to the content. It runs from a single Rust core with identical results across 14 language bindings, and it handles the awkward parts for you: JavaScript-heavy pages fall back to a real headless browser, bot filters are detected and worked around, robots and sitemaps are respected, requests are throttled per domain, and requests to private or internal addresses are refused by default. Drive it from your code, an AI agent, a REST service, or the CLI. Every part of the pipeline is a trait you can swap — the crawl frontier, rate limiter, storage, event stream, and content filters — so you can plug in your own behavior. Managed extras like proxy pools, tuned bot-evasion, authenticated sessions, scheduling, and billing live in xberg-enterprise. Features Feature Description ------- ----------- Structured extraction Text, metadata, links, images, assets, JSON-LD, Open Graph, hreflang, favicons, headings, response headers Markdown conversion Clean Markdown with citat","default_branch":null,"files":null,"tree":[],"storefront":"/r/xberg-io","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/xberg-io/crawlberg/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}