{"repo":"D4Vinci/Scrapling","free":true,"listed":false,"github":"https://github.com/D4Vinci/Scrapling","clone":"git clone https://github.com/D4Vinci/Scrapling.git","description":"🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!","language":"Python","stars":74528,"topics":["ai","ai-scraping","automation","crawler","crawling","crawling-python","data","data-extraction","mcp","mcp-server","playwright","python","scraping","selectors","stealth","web-scraper","web-scraping","web-scraping-python","webscraping","xpath"],"license":"BSD-3-Clause","category":"data_extraction","readme_excerpt":"<!-- mcp-name: io.github.D4Vinci/Scrapling -->\n\n<h1 align=\"center\">\n    <a href=\"https://scrapling.readthedocs.io\">\n        <picture>\n          <source media=\"(prefers-color-scheme: dark)\" srcset=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/docs/assets/cover_dark.svg?sanitize=true\">\n          <img alt=\"Scrapling Poster\" src=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/docs/assets/cover_light.svg?sanitize=true\">\n        </picture>\n    </a>\n    <br>\n    <small>Effortless Web Scraping for the Modern Web</small>\n</h1>\n\n<p align=\"center\">\n    <a href=\"https://trendshift.io/repositories/14244\" target=\"_blank\"><img src=\"https://trendshift.io/api/badge/repositories/14244\" alt=\"D4Vinci%2FScrapling | Trendshift\" style=\"width: 250px; height: 55px;\" width=\"250\" height=\"55\"/></a>\n    <br/>\n    <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_AR.md\">العربيه</a> | <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_ES.md\">Español</a> | <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_PT_BR.md\">Português (Brasil)</a> | <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_FR.md\">Français</a> | <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_DE.md\">Deutsch</a> | <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_CN.md\">简体中文</a> | <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_JP.md\">日本語</a> |  <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_RU.md\">Русский</a> | <a href=\"https://github.com/D4Vinci/Scrapling/blob/main/docs/README_KR.md\">한국어</a>\n    <br/>\n    <a href=\"https://github.com/D4Vinci/Scrapling/actions/workflows/tests.yml\" alt=\"Tests\">\n        <img alt=\"Tests\" src=\"https://github.com/D4Vinci/Scrapling/actions/workflows/tests.yml/badge.svg\"></a>\n    <a href=\"https://badge.fury.io/py/Scrapling\" alt=\"PyPI version\">\n        <img alt=\"PyPI version\" src=\"https://badge.fury.io/py/Scrapling.svg\"></a>\n    <a href=\"https://hub.docker.com/r/pyd4vinci/scrapling\" target=\"_blank\">\n        <img alt=\"Docker Pulls\" src=\"https://img.shields.io/docker/pulls/pyd4vinci/scrapling?labelColor=%20%23FDB062&logo=Docker&labelColor=%20%23528bff\"></a>\n    <a href=\"https://clickpy.clickhouse.com/dashboard/scrapling\" rel=\"nofollow\"><img src=\"https://img.shields.io/pypi/dm/scrapling\" alt=\"PyPI package downloads\"></a>\n    <a href=\"https://github.com/D4Vinci/Scrapling/tree/main/agent-skill\" alt=\"AI Agent Skill directory\">\n        <img alt=\"Static Badge\" src=\"https://img.shields.io/badge/Skill-black?style=flat&label=Agent&link=https%3A%2F%2Fgithub.com%2FD4Vinci%2FScrapling%2Ftree%2Fmain%2Fagent-skill\"></a>\n    <a href=\"https://clawhub.ai/D4Vinci/scrapling-official\" alt=\"OpenClaw Skill\">\n        <img alt=\"OpenClaw Skill\" src=\"https://img.shields.io/badge/Clawhub-darkred?style=flat&label=OpenClaw&link=https%3A%2F%2Fclawhub.ai%2FD4Vinci%2Fscrapling-official\"></a>\n    <br/>\n    <a href=\"https://discord.gg/EMgGbDceNQ\" alt=\"Discord\" target=\"_blank\">\n      <img alt=\"Discord\" src=\"https://img.shields.io/discord/1360786381042880532?style=social&logo=discord&link=https%3A%2F%2Fdiscord.gg%2FEMgGbDceNQ\">\n    </a>\n    <a href=\"https://x.com/Scrapling_dev\" alt=\"X (formerly Twitter)\">\n      <img alt=\"X (formerly Twitter) Follow\" src=\"https://img.shields.io/twitter/follow/Scrapling_dev?style=social&logo=x&link=https%3A%2F%2Fx.com%2FScrapling_dev\">\n    </a>\n    <br/>\n    <a href=\"https://pypi.org/project/scrapling/\" alt=\"Supported Python versions\">\n        <img alt=\"Supported Python versions\" src=\"https://img.shields.io/pypi/pyversions/scrapling.svg\"></a>\n</p>\n\n<p align=\"center\">\n    <a href=\"https://scrapling.readthedocs.io/en/latest/parsing/selection.html\"><strong>Selection methods</strong></a>\n    &middot;\n    <a href=\"https://scrapling.readthedocs.io/en/latest/fetching/choosing.html\"><strong>Fetchers</strong></a>\n    &middot;\n    <a href=\"https://scrapling.readthedocs.io/en/latest/spiders/architecture.html\"><strong>Spiders</strong></a>\n    &middot;\n    <a href=\"https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html\"><strong>Proxy Rotation</strong></a>\n    &middot;\n    <a href=\"https://scrapling.readthedocs.io/en/latest/cli/overview.html\"><strong>CLI</strong></a>\n    &middot;\n    <a href=\"https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html\"><strong>MCP</strong></a>\n</p>\n\nScrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl.\n\nIts parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box. And its spider framework lets you scale up to concurrent, multi-session crawls with pause/resume and automatic proxy rotation - all in a few lines of Python. One library, zero compromises.\n\nBlazing fast crawls with real-time stats and streaming. Built by Web Scrapers for Web Scrapers and regular users, there's something for everyone.\n\n```python\nfrom scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher\nStealthyFetcher.adaptive = True\np = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)  # Fetch website under the radar!\nproducts = p.css('.product', auto_save=True)                                        # Scrape data that survives website design changes!\nproducts = p.css('.product', adaptive=True)                                         # Later, if the website structure changes, pass `adaptive=True` to find them!\n```\nOr scale up to full crawls\n```python\nfrom scrapling.spiders import Spider, Response\n\nclass MySpider(Spider):\n  name = \"demo\"\n  start_urls = [\"https://example.com/\"]\n\n  async def parse(self, response: Response):\n      for item in response.css('.product'):\n          yield {\"title\": item.css('h2::text').get()}\n\nMySpider().start()\n```\n\n<p align=\"center\">\n    <a href=\"https://dataimpulse.com/?utm_source=scrapling&utm_medium=banner&utm_campaign=scrapling\" target=\"_blank\" style=\"display:flex; justify-content:center; padding:4px 0;\">\n        <img src=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/DataImpulse.png\" alt=\"At DataImpulse, we specialize in developing custom proxy services for your business. Make requests from anywhere, collect data, and enjoy fast connections with our premium proxies.\" style=\"max-height:60px;\">\n    </a>\n</p>\n\n# Platinum Sponsors\n<table>\n  <tr>\n    <td width=\"200\">\n      <a href=\"https://go.nodemaven.com/scraplingjuly\" target=\"_blank\" title=\"Proxies with the Highest IP Scores\">\n        <img src=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/NodeMaven.jpg\" width=\"240\" height=\"100\">\n      </a>\n    </td>\n    <td>\n    <a href=\"https://go.nodemaven.com/scraplingjuly\" target=\"_blank\">NodeMaven</a> - reliable proxy provider with the highest quality IP on the market. Use promo code SCRAPLING35 for 35% discount on proxies.\n    </td>\n  </tr>\n  <tr>\n    <td width=\"200\">\n      <a href=\"https://proxidize.com/?utm_source=github&utm_medium=sponsorship&utm_campaign=scrapling&utm_content=d4vinci\" target=\"_blank\" title=\"Clean Proxies with No Nonsense.\">\n        <img src=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/proxidize.png\">\n      </a>\n    </td>\n    <td> <a href=\"https://proxidize.com/?utm_source=github&utm_medium=sponsorship&utm_campaign=scrapling&utm_content=d4vinci\" target=\"_blank\"><b>Proxidize</b></a> provides mobile and residential proxies for scraping, browser automation, SEO monitoring, AI agents, and data collection. <i>Use code <b>scrapling20</b> for 20% off</i>.\n    </td>\n  </tr>\n  <tr>\n    <td width=\"200\">\n      <a href=\"https://coldproxy.com/?utm_source=scrapling&utm_medium=github&utm_campaign=coldproxy&utm_content=platinum_sponsor\" target=\"_blank\" title=\"Residential, IPv6 & Datacenter Proxies for Web Scraping\">\n        <img src=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/coldproxy.png\">\n      </a>\n    </td>\n    <td> <a href=\"https://coldproxy.com/?utm_source=scrapling&utm_medium=github&utm_campaign=coldproxy&utm_content=platinum_sponsor\" target=\"_blank\"><b>ColdProxy</b></a> provides residential and datacenter proxies for stable web scraping, public data collection, and geo-targeted testing across 195+ countries.\n    </td>\n  </tr>\n  <tr>\n    <td width=\"200\">\n      <a href=\"https://hypersolutions.co/?utm_source=github&utm_medium=readme&utm_campaign=scrapling\" target=\"_blank\" title=\"Bot Protection Bypass API for Akamai, DataDome, Incapsula & Kasada\">\n        <img src=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/HyperSolutions.png\">\n      </a>\n    </td>\n    <td> Scrapling handles Cloudflare Turnstile. For enterprise-grade protection, <a href=\"https://hypersolutions.co?utm_source=github&utm_medium=readme&utm_campaign=scrapling\">\n        <b>Hyper Solutions</b>\n      </a> provides API endpoints that generate valid antibot tokens for <b>Akamai</b>, <b>DataDome</b>, <b>Kasada</b>, and <b>Incapsula</b>. Simple API calls, no browser automation required. </td>\n  </tr>\n  <tr>\n    <td width=\"200\">\n      <a href=\"https://birdproxies.com/t/scrapling\" target=\"_blank\" title=\"At Bird Proxies, we eliminate your pains such as banned IPs, geo restriction, and high costs so you can focus on your work.\">\n        <img src=\"https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/BirdProxies.jpg\">\n      </a>\n    </td>\n    <td>Hey, we built <a href=\"https://birdproxies.com/t/scrapling\">\n        <b>BirdProxies</b>\n      </a> because proxies shouldn't be complicated or overpriced. Fast residential and ISP proxies in 195+ locations, fair pricing, and real support. <br />\n      <b>Try our FlappyBird game on the landing page for free data!</b>\n    </td>\n  </tr>\n  <tr>\n    <td width=\"200\">\n      <a href=\"https://evomi.com?utm_source=github&utm_medium=banner&utm_campaign=d4vinci-scrapling\" target=\"_blank\" title=\"Evomi is your Swiss Quality Proxy Provider, start","default_branch":"main","files":252,"tree":[".bandit.yml",".dockerignore",".github/FUNDING.yml",".github/ISSUE_TEMPLATE/01-bug_report.yml",".github/ISSUE_TEMPLATE/02-feature_request.yml",".github/ISSUE_TEMPLATE/03-other.yml",".github/ISSUE_TEMPLATE/04-docs_issue.yml",".github/ISSUE_TEMPLATE/config.yml",".github/PULL_REQUEST_TEMPLATE.md",".github/workflows/code-quality.yml",".github/workflows/docker-build.yml",".github/workflows/release-and-publish.yml",".github/workflows/tests.yml",".gitignore",".pre-commit-config.yaml",".readthedocs.yaml","AI_POLICY.md","CODE_OF_CONDUCT.md","CONTRIBUTING.md","Dockerfile","LICENSE","MANIFEST.in","README.md","ROADMAP.md","agent-skill/README.md","agent-skill/Scrapling-Skill.zip","agent-skill/Scrapling-Skill/LICENSE.txt","agent-skill/Scrapling-Skill/SKILL.md","agent-skill/Scrapling-Skill/examples/01_fetcher_session.py","agent-skill/Scrapling-Skill/examples/02_dynamic_session.py","agent-skill/Scrapling-Skill/examples/03_stealthy_session.py","agent-skill/Scrapling-Skill/examples/04_spider.py","agent-skill/Scrapling-Skill/examples/README.md","agent-skill/Scrapling-Skill/references/fetching/choosing.md","agent-skill/Scrapling-Skill/references/fetching/dynamic.md","agent-skill/Scrapling-Skill/references/fetching/static.md","agent-skill/Scrapling-Skill/references/fetching/stealthy.md","agent-skill/Scrapling-Skill/references/integrations/scrapy.md","agent-skill/Scrapling-Skill/references/mcp-server.md","agent-skill/Scrapling-Skill/references/migrating_from_beautifulsoup.md","agent-skill/Scrapling-Skill/references/parsing/adaptive.md","agent-skill/Scrapling-Skill/references/parsing/main_classes.md","agent-skill/Scrapling-Skill/references/parsing/selection.md","agent-skill/Scrapling-Skill/references/spiders/advanced.md","agent-skill/Scrapling-Skill/references/spiders/architecture.md","agent-skill/Scrapling-Skill/references/spiders/generic-templates.md","agent-skill/Scrapling-Skill/references/spiders/getting-started.md","agent-skill/Scrapling-Skill/references/spiders/platform-templates.md","agent-skill/Scrapling-Skill/references/spiders/proxy-blocking.md","agent-skill/Scrapling-Skill/references/spiders/requests-responses.md","agent-skill/Scrapling-Skill/references/spiders/sessions.md","benchmarks.py","cleanup.py","docs/README_AR.md","docs/README_CN.md","docs/README_DE.md","docs/README_ES.md","docs/README_FR.md","docs/README_JP.md","docs/README_KR.md","docs/README_PT_BR.md","docs/README_RU.md","docs/ai/mcp-server.md","docs/api-reference/custom-types.md","docs/api-reference/fetchers.md","docs/api-reference/mcp-server.md","docs/api-reference/proxy-rotation.md","docs/api-reference/response.md","docs/api-reference/selector.md","docs/api-reference/spiders.md","docs/assets/cover_dark.png","docs/assets/cover_dark.svg","docs/assets/cover_light.png","docs/assets/cover_light.svg","docs/assets/favicon.ico","docs/assets/logo.png","docs/assets/main_cover.png","docs/assets/scrapling_shell_curl.png","docs/assets/spider_architecture.png","docs/assets/star-history.png","docs/benchmarks.md","docs/cli/extract-commands.md","docs/cli/interactive-shell.md","docs/cli/overview.md","docs/development/adaptive_storage_system.md","docs/development/scrapling_custom_types.md","docs/donate.md","docs/fetching/choosing.md","docs/fetching/dynamic.md","docs/fetching/static.md","docs/fetching/stealthy.md","docs/index.md","docs/integrations/scrapy.md","docs/overrides/main.html","docs/overview.md","docs/parsing/adaptive.md","docs/parsing/main_classes.md","docs/parsing/selection.md","docs/requirements.txt","docs/spiders/advanced.md","docs/spiders/architecture.md","docs/spiders/generic-templates.md","docs/spiders/getting-started.md","docs/spiders/platform-templates.md","docs/spiders/proxy-blocking.md","docs/spiders/requests-responses.md","docs/spiders/sessions.md","docs/stylesheets/extra.css","docs/tutorials/migrating_from_beautifulsoup.md","docs/tutorials/replacing_ai.md","images/BirdProxies.jpg","images/CoreClaw.jpg","images/DataImpulse.png","images/HyperSolutions.png","images/NodeMaven.jpg","images/ProxyEmpire.png","images/SerpApi.png","images/SwiftProxy.png","images/TWSC.png","images/TikHub.jpg","images/coldproxy.png","images/decodo.png","images/evomi.png","images/novada.jpg","images/petrosky.png","images/proxidize.png","images/proxiware.png","images/webshare.png","pyproject.toml","pytest.ini","ruff.toml","scrapling/__init__.py","scrapling/cli.py","scrapling/core/__init__.py","scrapling/core/_shell_signatures.py","scrapling/core/_types.py","scrapling/core/ai.py","scrapling/core/custom_types.py","scrapling/core/mixins.py","scrapling/core/shell.py","scrapling/core/storage.py","scrapling/core/translator.py","scrapling/core/utils/__init__.py","scrapling/core/utils/_shell.py","scrapling/core/utils/_utils.py","scrapling/engines/__init__.py","scrapling/engines/_browsers/__init__.py","scrapling/engines/_browsers/_base.py","scrapling/engines/_browsers/_config_tools.py","scrapling/engines/_browsers/_controllers.py","scrapling/engines/_browsers/_page.py","scrapling/engines/_browsers/_stealth.py","scrapling/engines/_browsers/_types.py","scrapling/engines/_browsers/_validators.py","scrapling/engines/constants.py","scrapling/engines/static.py","scrapling/engines/toolbelt/__init__.py","scrapling/engines/toolbelt/ad_domains.py","scrapling/engines/toolbelt/convertor.py","scrapling/engines/toolbelt/custom.py","scrapling/engines/toolbelt/fingerprints.py","scrapling/engines/toolbelt/navigation.py","scrapling/engines/toolbelt/proxy_rotation.py","scrapling/fetchers/__init__.py","scrapling/fetchers/chrome.py","scrapling/fetchers/requests.py","scrapling/fetchers/stealth_chrome.py","scrapling/integrations/__init__.py","scrapling/integrations/scrapy.py","scrapling/parser.py","scrapling/py.typed","scrapling/spiders/__init__.py","scrapling/spiders/cache.py","scrapling/spiders/checkpoint.py","scrapling/spiders/engine.py","scrapling/spiders/links.py","scrapling/spiders/request.py","scrapling/spiders/result.py","scrapling/spiders/robotstxt.py","scrapling/spiders/scheduler.py","scrapling/spiders/session.py","scrapling/spiders/spider.py","scrapling/spiders/templates/__init__.py","scrapling/spiders/templates/_utils.py","scrapling/spiders/templates/crawler.py","scrapling/spiders/templates/feed.py","scrapling/spiders/templates/shopify.py","scrapling/spiders/templates/sitemap.py","scrapling/spiders/throttle.py","server.json","setup.cfg","tests/__init__.py","tests/ai/__init__.py","tests/ai/test_ai_mcp.py","tests/cli/__init__.py","tests/cli/test_cli.py","tests/cli/test_shell_functionality.py","tests/core/__init__.py","tests/core/test_shell_core.py","tests/core/test_storage_core.py","tests/fetchers/__init__.py","tests/fetchers/async/__init__.py","tests/fetchers/async/test_dynamic.py","tests/fetchers/async/test_dynamic_session.py","tests/fetchers/async/test_requests.py","tests/fetchers/async/test_requests_session.py","tests/fetchers/async/test_stealth.py","tests/fetchers/async/test_stealth_session.py","tests/fetchers/sync/__init__.py","tests/fetchers/sync/test_dynamic.py","tests/fetchers/sync/test_requests.py","tests/fetchers/sync/test_requests_session.py","tests/fetchers/sync/test_stealth_session.py","tests/fetchers/test_base.py","tests/fetchers/test_constants.py","tests/fetchers/test_impersonate_list.py","tests/fetchers/test_merge_request_args.py","tests/fetchers/test_pages.py","tests/fetchers/test_proxy_rotation.py","tests/fetchers/test_response_handling.py","tests/fetchers/test_utils.py","tests/fetchers/test_validator.py","tests/integrations/__init__.py","tests/integrations/test_scrapy.py","tests/parser/__init__.py","tests/parser/test_adaptive.py","tests/parser/test_ancestor_navigation.py","tests/parser/test_attributes_handler.py","tests/parser/test_find_similar_advanced.py","tests/parser/test_general.py","tests/parser/test_parser_advanced.py","tests/parser/test_selectors_filter.py","tests/requirements.txt","tests/spiders/__init__.py","tests/spiders/test_cache.py","tests/spiders/test_checkpoint.py","tests/spiders/test_engine.py","tests/spiders/test_feed.py","tests/spiders/test_force_stop_checkpoint.py","tests/spiders/test_links.py","tests/spiders/test_request.py","tests/spiders/test_result.py","tests/spiders/test_robotstxt.py","tests/spiders/test_scheduler.py","tests/spiders/test_session.py","tests/spiders/test_shopify.py","tests/spiders/test_sitemap.py","tests/spiders/test_spider.py","tests/spiders/test_templates.py","tests/spiders/test_throttle.py","tox.ini","zensical.toml"],"storefront":"/r/D4Vinci","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/D4Vinci/Scrapling/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}