{"repo":"SagarBiswas-MultiHAT/WebSource-Harvester","free":true,"listed":false,"github":"https://github.com/SagarBiswas-MultiHAT/WebSource-Harvester","clone":"git clone https://github.com/SagarBiswas-MultiHAT/WebSource-Harvester.git","description":"WebSource Harvester is an educational web-source harvester that crawls a site (BFS, depth-controlled), downloads browser-visible assets (HTML, CSS, JS, images, fonts, PDFs), and rewrites paths so pages work offline, including nested routes. It enforces same-origin limits and is designed for learning, offline analysis, and safe portfolio demos.","language":"Python","stars":26,"topics":["python","authorized-testing","bfs","crawler","css","educational","fonts","github-actions","html","images"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Web Source Code Downloader & Crawler &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Author: @SagarBiswas-MultiHAT \\ Category: Educational Web Crawling & Client-Side Security Analysis \\ Status: Learning-grade, interview-safe, portfolio-ready “This project performs depth-controlled crawling and client-side source reconstruction, capturing everything a browser can observe from a given URL, while intentionally respecting server-side trust boundaries.” --- --- Tested example: python \".\\PasourceDownloader.pyssword-Strength-Checker\" https://sagarbiswas-multihat.github.io/ --depth 2 --- Overview This project is an educational website source code downloader and crawler that extracts and reconstructs everything a browser can observe from a given URL. It crawls a site with depth-controlled BFS , downloads client-visible resources (HTML, CSS, JS, images, fonts, PDFs, etc.), and rewrites links so the pages work offline , even on nested paths like /blog/ . The tool respects server trust boundaries and does not attempt to fetch backend code, databases, or private data. --- What this project provides (accurate scope) This tool captures everything a browser can retrieve from a URL: - HTML pages (multiple pages via crawling) - Linked CSS files - JavaScript files - Images (including srcset ) - Fonts and media files - PDFs and other static assets - XML files (e.g., sitemap.xml ) - Correct offline reconstruction via path rewriting Ideal for: - Learning how real websites are structured - Offline inspection an","default_branch":null,"files":null,"tree":[],"storefront":"/r/SagarBiswas-MultiHAT","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/SagarBiswas-MultiHAT/WebSource-Harvester/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}