{"repo":"malvads/mojo","free":true,"listed":false,"github":"https://github.com/malvads/mojo","clone":"git clone https://github.com/malvads/mojo.git","description":"Non sucking cross-platform extremely fast C++ crawler to convert entire websites into LLM readable data","language":"C++","stars":16,"topics":["crawler","llm","rag"],"license":"MIT","category":"ai-agents","readme_excerpt":"Mojo Extremely Fast Web Crawler for AI & LLM Data Ingestion Mojo is a high-performance, multithreaded web crawler tailored for creating high-quality datasets for Large Language Models (LLMs) and AI training. Written in modern C++20 with coroutines, it rapidly fetches entire websites and converts them into clean, structured Markdown, making it the ideal tool for building knowledge bases and RAG (Retrieval-Augmented Generation) pipelines. Installation You can download the latest pre-compiled binaries from the Releases page. Linux (Binary Packages) For maximum compatibility, we recommend using the official packages which automatically handle dependencies: Debian / Ubuntu / Kali / Mint: CentOS / RHEL / Fedora: macOS 1. Download mojo-macos-arm (M1/M2/M3) or mojo-macos-intel . 2. Move it to your bin folder and give it execution permissions: 3. Make sure to grant privileges to the binary via security settings, since it is not signed. Windows 1. Download mojo-windows-x64.exe . 2. Run it from your terminal (CMD/Powershell). Key Features - High Performance : Built with C++20 coroutines, Boost.Beast, and Boost.Asio, Mojo utilizes a thread-pool architecture with async I/O to maximize throughput, significantly outperforming Python-based crawlers in high-volume tasks due to C++ native performance. - RAG-Ready Data Ingestion : Automatically transforms noisy HTML into clean, token-efficient Markdown. Perfect for populating vector databases (Pinecone, Milvus, Weaviate) or providing context fo","default_branch":null,"files":null,"tree":[],"storefront":"/r/malvads","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/malvads/mojo/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}