{"repo":"datalab-to/lift","free":true,"listed":false,"github":"https://github.com/datalab-to/lift","clone":"git clone https://github.com/datalab-to/lift.git","description":"Extract structured data from documents quickly and accurately.","language":"Python","stars":888,"topics":["ai","extract","ocr","pdf","python"],"license":"Apache-2.0","category":"media-processing","readme_excerpt":"Datalab State of the Art models for Document Intelligence lift lift extracts structured JSON from PDFs and images by passing a schema. It's a 9B vision model that returns a JSON object matching your schema, with schema-constrained decoding guaranteeing valid output. Try lift on Datalab Our managed platform runs improved extraction with higher accuracy than the open weights, plus per-field verification, citations, and confidence scores. If you have high volume workloads, we offer a batch processing service that has processed 1B+ pages per week. Get started with $20 in free credits per month — sign up - takes under 30 seconds - or try lift in our public playground. Commercial self-hosting requires a license — see Commercial usage. For on-prem licensing, contact us. Features - Extract structured data from documents - Pass any JSON schema - Handles multi-page documents in a single pass, including values that span pages - Two inference modes: local (HuggingFace) and remote (vLLM server) - CLI for single files, inline schemas, or whole directories - Schema Studio: a Streamlit app to build, save, and test schemas against your documents Quickstart The easiest way to start is with the CLI tools: Benchmarks Evaluated on a 225-document extraction benchmark (6–64 pages per document, 11,000 scored fields) with adversarial cases planted throughout: cross-page values, exhaustive lists, fields that must be left null, near-miss distractors, multi-source aggregation. Scoring is deterministic e","default_branch":null,"files":null,"tree":[],"storefront":"/r/datalab-to","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/datalab-to/lift/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}