FrontierThe story, in brief

Datalab Releases lift: A 9B Open-Weights Vision Model That Extracts Structured JSON From PDFs Using Schemas

90.2% field accuracy. Datalab's 9B vision model turns PDFs into valid JSON without hallucination—and it's open-weights.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A practical, production-ready vision model for document automation that solves a real enterprise pain point (PDF-to-structured-data) with schema constraints and trained abstention. Open-weights release makes it immediately deployable.

The key facts

8 to know
  1. Model size: 9B parameters

  2. Open-weights release

  3. 90.2% field accuracy on 225-document benchmark

  4. Schema-constrained decoding for output validity

  5. Trained abstention to avoid hallucinations on missing fields

  6. Converts PDFs and images to schema-matching JSON

  7. Released by Datalab

  8. Published June 23, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Datalab released lift, a 9B open-weights vision model that turns PDFs and images into schema-matching JSON. It uses schema-constrained decoding for valid structure and trained abstention to return null instead of hallucinating absent fields, scoring 90.2% field accuracy on a 225-document benchmark.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier