{"repo":"juvalen/mb-checker","free":true,"listed":false,"github":"https://github.com/juvalen/mb-checker","clone":"git clone https://github.com/juvalen/mb-checker.git","description":"Python scripts, first traverses chrome Bookmark file and second removes stale entries. Includes Jenkinsfile to generate docker images.","language":"Python","stars":12,"topics":["bookmarks","crawling","parallel-programming","python","tree","docker"],"license":"MIT","category":"deployment-docker-iac","readme_excerpt":"Bookmark cleansing R4.1 This is a simple python script utility to weed your good old bookmark file. After gathering and classifying bookmarks for more than 20 years one may hit dead URLs just when expecting them work. In order to keep the bookmark list current I created this script. Feed this python scripts with a Chrome bookmark file and a list of http return codes to be pruned and it will crawl through them and try to reach each entry. All successfull bookmarks will be copied to a cleaner json file, and failing URLs will be copied to additional files named as the specified return code. Empty bookmark folders can be optionally removed. Due to the large number of agents involved in Internet traffic, results achieved have not as reliable as to think about complete automation. Results may be inexact due to different redirect strategies, moved to https... So far, the suggestion is to keep the original bookmark file for some time, load the clean one in your browser, and review the excluded entries for yet valuable ones. Tasks are divided between two scripts. There is one script that runs on-line and crawls all entries included in the bookmarks and queues requests to workers that grab URLs in parallel performing these four steps: - workers are created an listen to queue - main loop pushes URLs to queue - workers read URLs from queue and try to reach them - workers write returned code to file A second script may be run off-line that reads the previous script output plus and a list ","default_branch":null,"files":null,"tree":[],"storefront":"/r/juvalen","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/juvalen/mb-checker/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}