{"repo":"dadoonet/fscrawler","free":true,"listed":false,"github":"https://github.com/dadoonet/fscrawler","clone":"git clone https://github.com/dadoonet/fscrawler.git","description":"Elasticsearch File System Crawler (FS Crawler)","language":"Java","stars":1450,"topics":["java","elasticsearch","crawler","tika"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"File System Crawler for Elasticsearch Welcome to FSCrawler for Elasticsearch This crawler helps to index binary documents such as PDF, Open Office, MS Office. Main features : Local file system (or a mounted drive) crawling and index new files, update existing ones and removes old ones. Remote file system over SSH/FTP crawling. REST interface to let you \"upload\" your binary documents to elasticsearch. Latest versions Current \"most stable\" versions are: Elasticsearch FSCrawler Released Docs --------------- -------------- ------------ ----------------------------------------------------------------------- 7.x, 8.x, 9.x 3.0-SNAPSHOT 3.0-SNAPSHOT Quick start Run Elasticsearch with start-local : Run FSCrawler with Docker: Then open Kibana and watch for your documents coming to the fscrawler alias: Or search for some text: Or count by file.content type : Note: /resumes contains the documents you want to index Job settings will be stored in /.fscrawler/fscrawler/ settings.yaml Read the documentation for more details and specifically the tutorial page. Project information Stats Version in preparation Latest release Build & quality License Read more about the Apache2 License. Thanks Thanks to JetBrains for the IntelliJ IDEA License! The best IDE out there! Thanks to SonarCloud for the free analysis! You guys rock!","default_branch":null,"files":null,"tree":[],"storefront":"/r/dadoonet","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/dadoonet/fscrawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}