{"repo":"novitae/njsparser","free":true,"listed":false,"github":"https://github.com/novitae/njsparser","clone":"git clone https://github.com/novitae/njsparser.git","description":"🦩 A NextJS data parser, to scrape peacefully","language":"HTML","stars":47,"topics":["javascript","next","nextjs","parser","scraping","scraper"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"NJSParser A powerful parser and explorer for any website built with NextJS. - Parses flight data (from the self. next f.push scripts). - Parses next data from NEXT DATA script. - Parses build manifests . - Searches for build id . - Many other things ... It uses only lxml , orjson , pydantic to garantee a fast and efficient data parsing and processing. Installation: Use CLI You can use the cli from 3 different commands: - njsp - njsparser - python3 -m njsparser.cli It has only one functionality of displaying informations about the website, like this: For more informations, use the --help argument with the command. Parsing next f . The data you find in next f is called flight data, and contains data under react format. You can parse it easily with njsparser the way it follows. We will build a parser for the flight data example 1. In the website you want to parse, make sure you see the self. next f.push in the begining of script contained the data you search for. Here I am searching for the description \"I should really have a better hobby, but this is it...\" (in blue) in my page, and I can also see the self. next f.push (in green). 2. Then I will do this simple script, to parse, then dump the flight data of my website, and see what objects I am searching for: 3. In my dumped flight data, I will search for the same string: 4. Then I will do to the closed \"value\" root to my found string, and look at the value of \"cls\" . Here it is \"Data\" : 5. Now that I know the \"cls\" (class) of o","default_branch":null,"files":null,"tree":[],"storefront":"/r/novitae","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/novitae/njsparser/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}