{"repo":"scrapfly/python-scrapfly","free":true,"listed":false,"github":"https://github.com/scrapfly/python-scrapfly","clone":"git clone https://github.com/scrapfly/python-scrapfly.git","description":"Official Python SDK for the Scrapfly platform: web scraping, screenshots, AI extraction, crawling, and a remote anti-bot browser. Integrates with Scrapy, LlamaIndex, and LangChain.","language":"Python","stars":62,"topics":["anti-bot","anti-detect","antidetect-browser","bot-detection","browser-automation","data-extraction","extraction-api","headless-browser","proxy","python"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"Scrapfly SDK Installation pip install scrapfly-sdk You can also install extra dependencies pip install \"scrapfly-sdk[seepdup]\" for performance improvement pip install \"scrapfly-sdk[concurrency]\" for concurrency out of the box (asyncio / thread) pip install \"scrapfly-sdk[scrapy]\" for scrapy integration pip install \"scrapfly-sdk[webhook-server]\" for have a native webhook server using flask pip install \"scrapfly-sdk[all]\" Everything! For use of built-in HTML parser (via ScrapeApiResponse.selector property) additional requirement of either parsel or scrapy is required. For reference of usage or examples, please checkout the folder /examples in this repository. This SDK cover the following Scrapfly API endpoints: Web Scraping API Extraction API Screenshot API Integrations Scrapfly Python SDKs are integrated with LlamaIndex and LangChain. Both framework allows training Large Language Models (LLMs) using augmented context. This augmented context is approached by training LLMs on top of private or domain-specific data for common use cases: - Question-Answering Chatbots (commonly referred to as RAG systems, which stands for \"Retrieval-Augmented Generation\") - Document Understanding and Extraction - Autonomous Agents that can perform research and take actions In the context of web scraping, web page data can be extracted as Text or Markdown using Scrapfly's format feature to train LLMs with the scraped data. LlamaIndex Installation Install llama-index , llama-index-readers-web , and sc","default_branch":null,"files":null,"tree":[],"storefront":"/r/scrapfly","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/scrapfly/python-scrapfly/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}