FrontierThe story, in brief

Datalab Lift vs the Field: How a 9B Schema-First Extractor Compares with NuExtract3, LlamaExtract, Marker, and Docling

9B beats the field. Datalab's Lift outperforms NuExtract3, LlamaExtract, and Docling on schema-first extraction—no Markdown middleman required.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Document extraction is becoming a critical enterprise AI workload. A smaller, purpose-built model achieving better accuracy than larger alternatives signals a shift toward specialized extractors over general-purpose LLMs for structured data tasks.

The key facts

5 to know
  1. Datalab Lift: 9B parameters

  2. Schema-first extraction approach (JSON Schema input → JSON output)

  3. Comparison benchmarks: NuExtract3, LlamaExtract, Marker, Docling

  4. Direct PDF/image-to-schema extraction without Markdown intermediary

  5. Competitive performance on document extraction benchmarks

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Datalab’s Lift is a focused document extraction tool with a specific promise: give it a PDF or image plus a JSON Schema, and it returns schema-shaped JSON directly. Instead of converting a document to Markdown first and then asking another model to extract fields, Lift reads rendered page images…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier