{"repo":"lightfeed/extractor","free":true,"listed":false,"github":"https://github.com/lightfeed/extractor","clone":"git clone https://github.com/lightfeed/extractor.git","description":"Use LLMs to robustly extract web data","language":"TypeScript","stars":320,"topics":["ai-agents","article-extractor","crawler","data-engineering","data-pipeline","etl","html-parser","html-to-markdown","llm","llm-extraction"],"license":"Apache-2.0","category":"data-pipelines","readme_excerpt":"Lightfeed Extractor Robust Web Data Extractor Using LLMs Overview Lightfeed Extractor is a Typescript library built for robust web data extraction using LLMs. Use natural language prompts to extract structured data from HTML, markdown, or plain text. Get complete, accurate results with great token efficiency — critical for production data pipelines. Features - 🧹 LLM-ready Markdown - Convert HTML to LLM-ready markdown, with options to extract only main content and clean URLs by removing tracking parameters. - ⚡️ LLM Extraction - Use LLMs in JSON mode to extract structured data according to input Zod schema. Token usage limit and tracking included. - 🛠️ JSON Recovery - Sanitize and recover failed JSON output. This makes complex schema extraction much more robust, especially with deeply nested objects and arrays. - 🔗 URL Validation - Handle relative URLs, remove invalid ones, and repair markdown-escaped links. - 🤖 Works with Playwright - Use Playwright to load pages, then extract structured data from the HTML content. - 🧭 AI Browser Navigation - Pair with @lightfeed/browser-agent to navigate pages using natural language commands before extracting structured data. [!TIP] Building retail competitor intelligence at scale? Go to lightfeed.ai - our full platform for tracking competitor pricing, sales, promotions, and SEO. Installation Install the extractor along with @langchain/core and your chosen LLM provider: Then add your LLM provider (we use LangChain for interoperability):","default_branch":null,"files":null,"tree":[],"storefront":"/r/lightfeed","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/lightfeed/extractor/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}