{"repo":"pragmar/mcp-server-webcrawl","free":true,"listed":false,"github":"https://github.com/pragmar/mcp-server-webcrawl","clone":"git clone https://github.com/pragmar/mcp-server-webcrawl.git","description":"MCP server tailored to connecting web crawler data and archives","language":"Python","stars":45,"topics":["mcp","mcp-server","mcp-servers","knowledgebase","warc","wget","archivebox","httrack","katana","siteone"],"license":null,"category":"mcp-servers","readme_excerpt":"Website GitHub Docs PyPi mcp-server-webcrawl Advanced search and retrieval for web crawler data. With mcp-server-webcrawl , your AI client filters and analyzes web content under your direction or autonomously. The server includes a fulltext search interface with boolean support, and resource filtering by type, HTTP status, and more. mcp-server-webcrawl provides the LLM a complete menu with which to search, and works with a variety of web crawlers: Crawler/Format Description Platforms Setup Guide --- --- --- --- [ ArchiveBox ][1] Web archiving tool macOS/Linux [Setup Guide][8] [ HTTrack ][2] GUI mirroring tool macOS/Windows/Linux [Setup Guide][9] [ InterroBot ][3] GUI crawler and analyzer macOS/Windows/Linux [Setup Guide][10] [ Katana ][4] CLI security-focused crawler macOS/Windows/Linux [Setup Guide][11] [ SiteOne ][5] GUI crawler and analyzer macOS/Windows/Linux [Setup Guide][12] [ WARC ][6] Standard web archive format varies by client [Setup Guide][13] [ wget ][7] CLI website mirroring tool macOS/Linux [Setup Guide][14] [1]: https://archivebox.io [2]: https://github.com/xroche/httrack [3]: https://interro.bot [4]: https://github.com/projectdiscovery/katana [5]: https://crawler.siteone.io [6]: https://en.wikipedia.org/wiki/WARC (file format) [7]: https://en.wikipedia.org/wiki/Wget [8]: https://pragmar.github.io/mcp-server-webcrawl/guides/archivebox.html [9]: https://pragmar.github.io/mcp-server-webcrawl/guides/httrack.html [10]: https://pragmar.github.io/mcp-server-webcrawl/","default_branch":null,"files":null,"tree":[],"storefront":"/r/pragmar","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/pragmar/mcp-server-webcrawl/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}