{"repo":"rumca-js/crawler-buddy","free":true,"listed":false,"github":"https://github.com/rumca-js/crawler-buddy","clone":"git clone https://github.com/rumca-js/crawler-buddy.git","description":"Crawling framework, RSS reader and parser","language":"HTML","stars":250,"topics":["feed-reader","rss","rss-client","rss-feed","rss-reader","webcrawling","webscraping"],"license":"GPL-3.0","category":"scrapers-browser-automation","readme_excerpt":"Crawling server The Crawling Server is an HTTP-based web crawler that delivers data in an easily accessible JSON format. - No need to rely on tools like yt-dlp or Beautiful Soup for extracting link metadata. - Metadata is standardized with consistent fields (e.g., title, description, date published, etc.). - Eliminates the need for custom HTTP wrappers to access RSS pages, even for sites with poorly configured bot protection. - Automatically discovers RSS feed URLs for websites and YouTube channels in many cases. - Simplifies data handling—no more parsing RSS files; just consume JSON! - Offers a unified interface for all metadata. - Running a containerized docker environment helps isolate problems from the host operating system - All your crawling / scraping / rss clients could use one source, or you can split it up by hosting multiple servers - Encoding? What encoding? All responses are in UTF Main Available Endpoints: - GET / - Provides index page - GET /info - Displays information about available crawlers. - GET /system - Information about system. - GET /debug - debug information Operation Endpoints: - GET /history - Displays the crawl history. - GET /queue - Displays information about the current queue. - GET /find - form for api/find Endpoint. Forms: - GET /get - form for api/get Endpoint. - GET /contents - form for contentsr Endpoint - GET /feeds - form for finding feeds for the specified URL - GET /social - form for social information - GET /link - form for link inform","default_branch":null,"files":null,"tree":[],"storefront":"/r/rumca-js","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/rumca-js/crawler-buddy/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}