Categories:
Tools
OCR open-source document AI RAG

How LightOnOCR-3 Turns Messy PDFs Into Structured Data in One Pass

Feature image for How LightOnOCR-3 Turns Messy PDFs Into Structured Data in One Pass

Every document pipeline hits the same wall sooner or later. The OCR engine reads the words. A separate layout detector figures out where the tables live. A third component tries to salvage something from the charts. Three models, three failure points, and a pile of glue code nobody wants to maintain.

LightOnOCR-3, which LightOn released on October 8, does all three jobs in a single pass. It comes in three sizes (0.8B, 1B, and 4B parameters), all under an Apache 2.0 license, with the 0.8B and 4B variants built on Qwen3.5’s vision-language architecture. The smallest runs on a modest GPU. None of them require a sales call.

If you build document ingestion for a living, this is the kind of release that deletes a week of integration work. Here is what shipped and how to put it to work.

What shipped

LightOn published the models with a training writeup covering the supervised fine-tuning mixture, the iterative bounding-box data pipeline, and a reinforcement learning stage trained on manually verified examples. That last part matters more than it sounds. RL on OCR outputs is where a lot of open models quietly fall apart on edge cases, and LightOn at least documented the process instead of hand-waving it.

The demo space is live on Hugging Face and the GitHub repo is public, so you can test before committing to anything.

The real differentiator is the grounding mode. That deserves its own section.

Grounding mode: one prompt, full structure

Send the word “grounding” as your prompt and the model returns more than transcribed text. You get labeled bounding boxes for every document region, short descriptions for images, and chart data extracted as HTML tables.

The bounding-box coordinates are normalized to a 0-1000 range and emitted as compact inline markers rather than wrapped in JSON or HTML. That choice keeps token overhead low, which matters when you are processing a few hundred pages and paying per token. Or waiting on them.

In practice, the output for a scanned invoice looks like this: the text block marked as a table, the table marked with its coordinates, and the chart beside it rendered as rows and columns you can drop straight into a database. One request. One model.

Why one model beats three

I have stitched together enough pipelines to have opinions here.

The multi-component approach fails in a specific way: each model is individually fine and jointly terrible. The OCR gets most characters right, the layout detector finds most regions, and the chart parser handles the easy cases. Then the errors compound at the seams. A table row loses its header because the layout box was off by twenty pixels. A caption gets assigned to the wrong figure. Nobody notices until a customer does.

A single model trained end to end on all three tasks makes one set of mistakes instead of three interacting sets. When it fails, the failure is at least legible. You can look at the grounding output and see exactly which region went wrong.

There is also the bill. Three components mean three inference calls or three deployments to babysit. LightOnOCR-3 is one call, and the 0.8B model is small enough that self-hosting often costs less than the API you would otherwise pay for.

The RAG angle: structure going in means better retrieval coming out

Retrieval quality is decided at ingestion time, not query time. When a PDF gets flattened into an undifferentiated wall of text, your chunker cuts mid-table, splits a caption from its figure, and buries the header row three chunks away from the data it describes. Every one of those cuts is a retrieval miss waiting to happen.

Grounding output fixes this at the source. Because you get regions with labels and coordinates, you can chunk by semantic unit instead of by character count: one chunk per table with its header intact, one per section, one per figure and its caption together. The coordinates even let you render a thumbnail of each region for hybrid search.

It is the least glamorous lever in RAG and the one that pays the most.

What to do now

  1. Try the demo space on Hugging Face with three of your own documents: an invoice, a report with charts, and something scanned badly on purpose. Do not test on marketing PDFs. They are too easy.
  2. Pull the GitHub repo and run the 0.8B model locally on the same documents. Compare the grounding output against what your current pipeline produces.
  3. If you build RAG ingestion, prototype chunking by region label instead of by token count. Measure recall on a small eval set before and after.
  4. If the results hold, route by document complexity: keep the 0.8B model for clean scans and swap in the 4B for the hard ones.

Limits worth knowing

The models are new and independent benchmarks are thin so far. The training writeup is detailed, but it is LightOn evaluating LightOn, so test on your own documents before trusting any vendor’s numbers, these included. Handwriting-heavy documents and low-resource languages deserve extra verification; the writeup says little about either.

On sizes: the 1B model exists, but on most hardware the real choice is between 0.8B for speed and 4B for accuracy.

One open model will not solve every document problem you have. It replaces the three-part pipeline most teams still hand-maintain, with weights you can inspect, fine-tune, and self-host. For document AI, that is a good week.

Related Articles