{"repo":"wszqkzqk/qt-web-extractor","free":true,"listed":false,"github":"https://github.com/wszqkzqk/qt-web-extractor","clone":"git clone https://github.com/wszqkzqk/qt-web-extractor.git","description":"Multimodal web content extraction engine backed by Qt WebEngine.","language":"Python","stars":26,"topics":["chromium","content-extraction","headless-browser","open-webui","pdf-extraction","pyside6","qtwebengine","web-scraping","mcp","mcp-server"],"license":"GPL-3.0","category":"mcp-servers","readme_excerpt":"Qt Web Extractor A general-purpose multimodal web content extraction engine powered by Qt WebEngine (Chromium). Designed to extract fully-rendered content from modern web pages that rely on JavaScript, cookies, dynamic content loading, or client-side rendering — transforming complex, noisy web structures directly into clean Markdown text format, preserving links and readability for LLM processing or downstream pipelines. Also supports extracting text from PDF documents via Qt PDF and fetching images through the same browser engine, so vision-capable agents can see the web, not just read it. Key features: - Smart Markdown formatting — converts fully rendered web pages directly into clean Markdown, preserving links and structure without the noise of raw HTML. Ideal for feeding text into LLMs. - Full JavaScript rendering — handles SPAs, React/Vue/Angular apps, and any page that requires JS to display content. - Multimodal MCP support — serve images to vision-capable agents as native MCP image content, fetched through the same full browser engine. - PDF extraction — extract text from PDF documents or render pages as images for vision-capable agents, all through the same real browser engine. - Cookie & session support — access pages behind login walls or consent gates. - Multiple interfaces — use as a CLI tool, Python library, or HTTP service with a simple REST API. - Headless operation — runs in Qt offscreen mode, no display or GPU required. - Lightweight dependencies — no standa","default_branch":null,"files":null,"tree":[],"storefront":"/r/wszqkzqk","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/wszqkzqk/qt-web-extractor/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}