{"repo":"m92vyas/llm-reader","free":true,"listed":false,"github":"https://github.com/m92vyas/llm-reader","clone":"git clone https://github.com/m92vyas/llm-reader.git","description":"Turn Webpage to LLM friendly input text. Similar to Firecrawl and Jina Reader API. Makes RAG, AI web scraping, image & webpage links extraction easy.","language":"Python","stars":304,"topics":["extract-data","llm","llm-agent","scraper","scraping","scraping-websites","webscraping","ai-agent-tools","ai-agents","ai-web-scraper"],"license":"MIT","category":"ai-agents","readme_excerpt":"Webpage to LLM Ready Input Text Pre-processing webpage before giving it as input to the LLM improves extraction/scraping accuracy especially if you want to extract website and image links, tables required for most scraping operations like scraping an e-commerce website. Use this library to turn any webpage/url to LLM friendly text. Fully open source alternative to firecrawl and jina reader api. You can also refer to my other repo AI-web scraper for direct scraping tools that will do web search and scrapes multiple links with just a simple query . It supports multiple LLMs, Web Search and Extracts Data as per your written instructions. --- Update for Old Users: We have switched from Selenium to Playwright for concurrent web scraping support. Kindly install the required playwright dependencies as given below. --- Install: Import: Get processed LLM input text: --- Example Usage: suppose we want to scrape the product name, main product page link, image link and price from the url \"https://www.ikea.com/in/en/cat/corner-sofas-10671/\" using any openai model. --- Documentation: https://github.com/m92vyas/llm-reader/wiki/Documentation --- To Scrape and Crawl without getting Blocked: - You can replace the get page source function with any paid API or use your own proxies or scraping setup to get the page source. e.g. you can use a pay-as-you-go option like Scrappey to get page source without getting blocked and then pass the HTML to get processed text function to get LLM text for free.","default_branch":null,"files":null,"tree":[],"storefront":"/r/m92vyas","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/m92vyas/llm-reader/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}