{"repo":"StabRise/spark-pdf","free":true,"listed":false,"github":"https://github.com/StabRise/spark-pdf","clone":"git clone https://github.com/StabRise/spark-pdf.git","description":"PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it","language":"Scala","stars":81,"topics":["ocr","ocr-recognition","pdf","pdf-document","pdf-document-processor","spark","spark-datasource","big-data","data-engineering","data-extraction"],"license":"AGPL-3.0","category":"media-processing","readme_excerpt":"--- ⭐ Star us on GitHub — it motivates us a lot! Source Code : https://github.com/StabRise/spark-pdf Quick Start Jupyter Notebook Spark 3.5.x on Databricks : PdfDataSourceDatabricks.ipynb Quick Start Jupyter Notebook Spark 3.x.x : PdfDataSource.ipynb Quick Start Jupyter Notebook Spark 4.0.x : PdfDataSourceSpark4.ipynb With Spark Connect : PdfDataSourceSparkConnect.ipynb --- Welcome to the Spark PDF The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame. If you found useful this project, please give a star to the repository. 👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog). Key features: - Read PDF documents to the Spark DataFrame - Support efficient read PDF files lazy per page - Support big files, up to 10k pages - Support scanned PDF files (call OCR for text recognition from the images) - No need to install Tesseract OCR, it's included in the package - 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark. - Works with Spark Connect Requirements - Java 8, 11, 17 - Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0 - Ghostscript 9.50 or later (only for the GhostScript reader) Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13). Installation Binary package is available in the Maven Central Repository. - Spark 3.5. : com.stabrise:spark-pdf-spark35 2.12:0.1.17 - Spark 3.4. : com","default_branch":null,"files":null,"tree":[],"storefront":"/r/StabRise","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/StabRise/spark-pdf/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}