{"repo":"SermetPekin/rss-discovery-engine","free":true,"listed":false,"github":"https://github.com/SermetPekin/rss-discovery-engine","clone":"git clone https://github.com/SermetPekin/rss-discovery-engine.git","description":"A self-expanding blog discovery system that automatically finds new blogs by exploring network relationships. Start with a few seed blogs and watch it recursively discover hundreds more through link analysis.","language":"HTML","stars":11,"topics":["blog","feed-reader","network","rss"],"license":"BSD-3-Clause","category":"scrapers-browser-automation","readme_excerpt":"RSS Discovery Engine A recursive crawler designed to map the independent web. It starts with a handful of seed blogs and follows the breadcrumbs—citations, blogrolls, and links—to discover the hidden network of writers and thinkers. 👉 View Live Demo Why this exists Most discovery today is algorithmic, driven by engagement metrics on centralized platforms. This engine takes a different approach: it trusts the writers you already read. By following who they link to, we can uncover a graph of high-quality, human-curated content that often flies under the radar of search engines and social feeds. How it works The engine uses a recursive strategy: 1. Ingest : Starts with a list of trusted \"seed\" blogs. 2. Crawl : Fetches the latest posts via RSS/Atom feeds. 3. Analyze : Scans content for outbound links to other domains. 4. Verify : Checks if those domains are valid blogs (active feeds, non-corporate, non-spam). 5. Expand : Adds verified blogs to the queue and repeats. The result is a directed graph of the blogosphere, visualized interactively to show communities and connections. Architecture Built with Python, this project emphasizes robustness and modularity: - Data Validation : Uses Pydantic for strict type checking and data modeling. - Configuration : Environment-aware settings management via pydantic-settings . - State Management : Resilient checkpointing system that saves progress automatically, allowing long-running crawls to be paused and resumed. - Graph Visualization : A","default_branch":null,"files":null,"tree":[],"storefront":"/r/SermetPekin","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/SermetPekin/rss-discovery-engine/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}