{"repo":"FaustoS88/PinescriptV6-docs-crawler","free":true,"listed":false,"github":"https://github.com/FaustoS88/PinescriptV6-docs-crawler","clone":"git clone https://github.com/FaustoS88/PinescriptV6-docs-crawler.git","description":"A Python tool for crawling and processing TradingView's PineScript V6 documentation. Built with Crawl4Ai framework, it extracts, cleans, and organizes documentation into searchable markdown files, making it easier to reference and analyze PineScript features and syntax.","language":"Python","stars":27,"topics":["crawl4ai","documentation","pinescript","python","tradingview","web-scraping"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Pine Script v6 Documentation Crawler & RAG Processor An automated crawler and heading-aware chunking pipeline for the official TradingView Pine Script v6 documentation. Built with Crawl4AI and designed to feed high-quality chunks directly into RAG ingestion pipelines (e.g. pgvector, Pinecone, Qdrant). Features - Incremental Crawling : SHA-256 content hashing via manifest.json — skips unchanged pages automatically on subsequent runs. - Clean Overwrites : Overwrites page name.md instead of accumulating timestamped files. - Heading-Aware Chunking : chunker.py splits documents at H1/H2 boundaries into self-contained sections ( <3000 chars, 200 char overlap). Never cuts inside Pine Script code blocks. - Breadcrumb Context : Each chunk includes section breadcrumbs (e.g. [\"Language\", \"Types\", \"Arrays\"] ) for rich context during vector retrieval. - JSONL Combined Output : Exports chunks all docs.jsonl for one-pass vector DB population. Quick Start Repository Structure Regenerating Docs (CI) A GitHub Actions workflow ( .github/workflows/refresh-docs.yml ) lets you re-crawl and reprocess the docs on demand from the Actions tab ( workflow dispatch ). Pine Script v6 is stable, so there is no timer by default — run it when you actually have a reason: a new language feature landed, you're cloning fresh, or you're doing a pre-release check. A commented-out quarterly cron lives in the file if you ever want a passive drift check instead. Thanks to incremental crawling, a run that finds no cha","default_branch":null,"files":null,"tree":[],"storefront":"/r/FaustoS88","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/FaustoS88/PinescriptV6-docs-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}