On 8 October 2026, LightOn published LightOnOCR-3: three OCR models that transcribe a page and, on a second prompt, mark where each block sits. The Hugging Face licence field is apache-2.0 on LightOnOCR-3-0.8B, LightOnOCR-3-1B, and LightOnOCR-3-4B. The blog says Apache 2.0 for research and commercial use.
LightOn says the 1B keeps the previous architecture, while the 0.8B and 4B use the Qwen3.5 vision-language architecture. The cards agree: the 0.8B lists base model Qwen/Qwen3.5-0.8B, the 4B lists Qwen/Qwen3.5-4B, and the 1B lists no base model.
Empty prompt, or the word grounding
An empty text prompt transcribes the page, the same interface as earlier LightOnOCR releases. The single word grounding adds labelled bounding boxes, short image descriptions, and chart data as HTML tables. Coordinates are normalized from 0 to 1000. Text blocks hold paragraphs and titles. Image blocks pair a box with a short description. Chart blocks hold an HTML table of the figure’s data. Labels and coordinates are compact inline markers. Chart tables stay HTML.
That layout is what a document pipeline can pass downstream. Chunking for retrieval (RAG: splitting a document so a later search returns the right passage) needs boundaries. A box tagged as a paragraph, a table, or a figure is a boundary the pipeline can keep, instead of cutting a caption in half. LightOn offers the format for that use. The scores below are LightOn’s, not a measurement made for this article.
On the same 512 olmOCR-bench pages, LightOn says the 0.8B and 4B models in grounding mode generate 9 to 14% fewer output tokens on average than Chandra-OCR-2 and Infinity-Parser2-Pro. Those outputs include boxes, image descriptions, and chart tables. LightOn notes that Infinity-Parser2-Pro, for example, does not generate image descriptions, and that the comparison is of complete outputs.
Tables LightOn printed
Higher is better, and LightOn says the tables keep the top models from a wider set it ran. On olmOCR-bench, LightOnOCR-3-4B scores 86.3 overall. Infinity Parser Pro, listed at 35.1 billion parameters, scores 87.6. Chandra 2, at the same listed 4 billion, scores 85.8. The 0.8B scores 85.5 and the 1B 84.5. LightOn says the 4B leads its ArXiv column (91.1), the 1B leads multi-column pages (85.9), and the 0.8B leads long tiny text (94.1). A footnote says the older LightOnOCR-2-1B overall of 83.2 drops headers and footers, so it is not comparable.
On ParseBench’s five-category overall, LightOn puts the 4B at 75.1 and the 0.8B at 74.6, ahead of Infinity Parser Pro at 74.3. The 1B scores 71.4. The 4B leads charts at 66.1. Infinity Parser Pro leads visual grounding at 74.9. On the French set fr-bench-pdf2md, LightOn says the 4B leads overall at 74.1.
LightOn says it did not train its native format to the markup each benchmark prefers, and that it normalizes text before scoring. A rewrite can move an edit-distance score without changing the words read off the page. The functions, it says, are in the repo. A one-point gap on these tables is that normalized score.
The GitHub repo is a client, a command-line tool, and a viewer for vLLM, plus the benchmark scripts. Its README says grounding works on LightOnOCR-3 only. A demo space is linked from the post. Empty prompt for a transcript. The word grounding for boxes on a 0 to 1000 grid. Check apache-2.0 on the card you download: the 1B is the previous architecture, and the 0.8B and 4B are Qwen3.5.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.