{"repo":"WillyEverGreen/acon","free":true,"listed":false,"github":"https://github.com/WillyEverGreen/acon","clone":"git clone https://github.com/WillyEverGreen/acon.git","description":"The intelligence layer for any web scraper. Pair with Scrapling, Playwright, or httpx to crawl smarter.","language":"Python","stars":41,"topics":["crawler","playwright","python","scraping","site-intelligence","spider","web-scraping"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Acon — The Intelligent Brain for Any Scraper Acon doesn't replace Scrapling or Firecrawl. It tells them where to look. --- Why Acon? Most crawlers are dumb. They follow links blindly, return raw HTML, and break the moment a site changes its structure. Before you can extract anything useful, you need to understand what you're dealing with. Acon is a site intelligence engine. It maps the structural \"skeleton\" of a website automatically — before any data extraction happens — so your scraper always knows where to look. --- 🏗️ The Core Thesis Most modern web scrapers suffer from \"URL Exhaustion\" — they spend 90% of their bandwidth fetching identical product or blog pages. Acon introduces a Topology Orchestrator that maps, classifies, and samples site structures, then stops the moment it has fully learned the site's DNA — no wasted requests. --- 📊 Real-World Benchmark Results (v0.1.2 - Final 10/10 Polish) The correct question : How many pages does each engine need to fully map a site's structure? Both crawlers given an uncapped budget . BFS runs until exhaustion. Acon stops the moment low information gain fires — meaning the site's structural DNA is fully mapped. Comparison Summary (4 Representative Sites) Site BFS Pages Acon Pages Request Reduction Time Saved Stopped By :--- :---: :---: :---: :---: :--- books.toscrape.com 200 6 97.0% 93.7% low information gain Hacker News 50 9 82.0% 89.0% low information gain Wikipedia 100 8 92.0% 93.7% low information gain PyPI 100 20 80.0% 93.","default_branch":null,"files":null,"tree":[],"storefront":"/r/WillyEverGreen","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/WillyEverGreen/acon/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}