{"repo":"pzaino/thecrowler","free":true,"listed":false,"github":"https://github.com/pzaino/thecrowler","clone":"git clone https://github.com/pzaino/thecrowler.git","description":"A Content Discovery and Development Platform. Empowering Cybersecurity, AI, Marketing, and Finance professionals and researchers to discover, analyze, and interact with the web in all its dimensions.","language":"Go","stars":61,"topics":["crawler","crawling","golang","indexer","indexing","search-engine","automation","content-detection","content-discovery","cyber-security"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"The CROWler Please consider supporting this project! The CROWler is empowering thousands of users (mostly professionals and enterprises) to build their own content discovery and intelligence based solutions, and it's growing fast. If you want to support the project, you can do it by: - Clicking the support button on the GitHub page - Buying some official merchandise from the CROWler Merch Store - Contacting the author for consultancies, paid support, or custom development. Important note : From release 2.0.0, the section selenium in the configuration file has been renamed to vdi . Please update your configuration file accordingly. The old selenium section is no longer supported. What is it? The CROWler is a self-hosted, event-driven Content Discovery and Intelligence development platform designed for advanced web, email, APIs (both REST and WebSockets), filesystems, network crawling, scraping, detection, and automation using real browsers, rulesets, plugins, and agents. Project status: Still under active development (WIP). Most components are usable. Beta testers welcome. Full daily progress stats. Additionally, the system is equipped with a powerful search API, providing a streamlined interface for data queries. This feature ensures easy integration and access to indexed data for various applications. The CROWler is designed to be micro-services based, so it can be easily deployed in a containerized environment. As a Content Discovery Development Platform, The CROWler integr","default_branch":null,"files":null,"tree":[],"storefront":"/r/pzaino","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/pzaino/thecrowler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}