WZ-IT Logo

Local Vision-Language Models for OCR: Reading Invoices and Scans

Timo Wevelsiep
Timo Wevelsiep
•
#OCR #VisionLanguageModel #DocumentProcessing #LocalAI #Qwen #olmOCR

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

Local Vision-Language Models for OCR: Reading Invoices and Scans

Read documents locally instead of sending them to an external AI service? WZ-IT builds the extraction pipeline with open models and validation rules on your infrastructure, see AI document processing in the AI hub. Book a meeting

Since 2025 a new group of open models has appeared that read document pages as images and output text, Markdown or JSON: olmOCR, PaddleOCR-VL, GLM-OCR, DeepSeek-OCR, LightOnOCR and general vision-language models such as Qwen3.8. Almost all of them run on a single GPU, and many are licensed under Apache 2.0 or MIT. For invoices, delivery notes and scanned letters this is a real alternative to cloud services, with no document leaving your own network.

This article compares the models from an operator's perspective: what they do, which licence they are under, how to read the benchmark scores on model cards, what hardware they need and how to limit the risk of invented numbers. The overall process with e-invoicing, validation rules, approval and German bookkeeping requirements (GoBD) is covered in the article Extract and validate documents with AI. All information is as of September 2026.

Table of Contents

  1. Three tool classes: OCR, OCR VLM, general VLM
  2. Open models at a glance, as of September 2026
  3. Licences: not every model can be used commercially
  4. How to read benchmarks
  5. The number problem: plausible is not read
  6. Pipeline patterns for invoices and scans
  7. Hardware: 24 GB, 96 GB or AI Cube
  8. German documents: language, number format, original quality
  9. Our approach at WZ-IT
  10. Further guides

Three tool classes: OCR, OCR VLM, general VLM

Class Examples Input Output Typical error
Classic OCR Tesseract, PaddleOCR (PP-OCR), OCRmyPDF image, scan characters with positions, text layer in the PDF visible character errors, broken tables
OCR VLM, specialised for documents olmOCR 2, PaddleOCR-VL, GLM-OCR, LightOnOCR-2, DeepSeek-OCR-2, granite-docling page image Markdown, HTML, tables, partly JSON omitted or added passages, fluent text instead of a gap
General VLM Qwen3.8-27B, Qwen3-VL page image plus instruction free answer or JSON following a schema plausible but unread values

Classic OCR reads characters but does not understand structure. A field such as "invoice number" then requires rules or templates. OCR VLMs are small models, usually between 0.3 and 8 billion parameters, trained to transcribe a page completely and in reading order, including tables. General VLMs are large multimodal language models that process an image and an instruction. They can extract specific fields but are slower and more prone to filling gaps with something plausible.

Docling belongs to none of the three classes but connects them: the MIT-licensed project of the LF AI & Data Foundation detects layout, reading order and table structure, integrates classic OCR for scans and can optionally use a VLM such as granite-docling (docling --pipeline vlm). The current version is 2.131.0 from 29 September 2026 (Docling releases).

Open models at a glance, as of September 2026

Model Vendor Parameters Licence Released Base Output
Qwen3.8-27B Alibaba Qwen 27 billion Apache 2.0 August 2026 own architecture, natively multimodal free answer, JSON
olmOCR-2-7B-1025 Allen Institute for AI 7 billion (base) Apache 2.0 October 2025 Qwen2.5-VL-7B-Instruct Markdown
PaddleOCR-VL-1.6 Baidu PaddlePaddle 0.9 billion Apache 2.0 28 May 2026 ERNIE 4.5 0.3B Markdown, JSON with layout
GLM-OCR Z.ai 0.9 billion MIT January 2026 GLM-V, GLM-0.5B decoder Markdown, JSON following a schema
LightOnOCR-2-1B LightOn 1 billion Apache 2.0 January 2026 own architecture text in reading order
DeepSeek-OCR-2 DeepSeek about 3 billion Apache 2.0 January 2026 own architecture Markdown
granite-docling-258M IBM Research 258 million Apache 2.0 17 September 2025 Idefics3, Granite 165M DocTags for Docling

How the candidates compare:

  • Qwen3.8-27B is not a specialised OCR model but a general language model with image understanding. The model card explicitly lists documents as a use case, a context length of 262,144 tokens and a thinking mode that can be switched off per request. Switching it off makes sense for field extraction because long reasoning chains increase response time. Qwen provides an FP8 variant itself (Qwen3.8-27B-FP8).
  • olmOCR 2 is designed to convert whole PDF pages into text and runs through the olmOCR toolkit with vLLM, which renders, rotates and retries pages as needed. Allen AI recommends the FP8 variant for practical use (model card). Training rewards passing unit tests on documents (olmOCR 2, arXiv 2510.19817).
  • PaddleOCR-VL-1.6 has the same architecture as version 1.5 and can be swapped in without changes. The first version listed 109 supported languages (PaddleOCR-VL), and version 1.5 added seal and stamp recognition (PaddleOCR-VL-1.5).
  • GLM-OCR combines layout analysis (PP-DocLayout-V3) with parallel recognition and, besides transcription, supports information extraction in which the prompt specifies a JSON schema. It runs on vLLM, SGLang and Ollama.
  • LightOnOCR-2-1B explicitly lists German among its supported languages and is served directly with vllm serve.
  • granite-docling-258M is the smallest model in the list and is built for the Docling pipeline. The model card lists English as its language.

Licences: not every model can be used commercially

Hugging Face hosts further OCR models with good benchmark scores whose licence rules out or leaves open use in a company. Two examples:

Model Licence according to model card Consequence for companies
Chandra OCR 2 modified OpenRAIL-M use excluded for more than 2 million US dollars in prior-year revenue or more than 2 million US dollars in funding, except personal or research use; commercial licence from the vendor (LICENSE, Attachment A)
Nanonets-OCR2-3B not stated base model Qwen2.5-VL-3B-Instruct is under the Qwen Research License, which only permits non-commercial use (licence text)

Three rules for the check:

  1. Separate model licence and code licence. For Chandra the code is Apache 2.0, the weights are not. The weights are what matters for operation.
  2. Check the base model. A fine-tune inherits the terms of its base model. olmOCR 2 builds on Qwen2.5-VL-7B, which is under Apache 2.0. For the 3B variant of the same base model that would not be the case.
  3. Check pipeline components individually. Docling itself is MIT, while the integrated models are subject to their own licences, as the project explicitly states.

A detailed overview of licence models for open language models is provided in the article Open LLM licences for commercial use. This is a technical classification, not legal advice.

How to read benchmarks

Every model card reports top scores. The figures come from the vendors themselves and refer to different benchmarks and versions.

Model Benchmark Score according to vendor
PaddleOCR-VL-1.6 OmniDocBench v1.6 96.33
GLM-OCR OmniDocBench v1.5 94.62
PaddleOCR-VL-1.5 OmniDocBench v1.5 94.5
Qwen3.8-27B OmniDocBench 1.5 91.1
Chandra OCR 2 olmOCR-Bench 85.8
olmOCR-2-7B-1025-FP8 olmOCR-Bench, toolkit v0.4.0 82.4 ± 1.1

Sources: model cards of PaddleOCR-VL-1.6, GLM-OCR, PaddleOCR-VL-1.5, Qwen3.8-27B, Chandra OCR 2 and olmOCR 2.

What the table does not show:

  • Versions are not comparable. OmniDocBench was updated to v1.6 in April 2026, with 296 additional, harder pages and a new matching method, and to v1.7 shortly afterwards. A v1.5 score and a v1.6 score do not measure the same thing.
  • The overall score includes formulas. The OmniDocBench overall score is the mean of text accuracy, table structure (TEDS) and formula recognition (CDM). Formulas are irrelevant for invoices and central for scientific papers.
  • The documents are not yours. In its base version OmniDocBench comprises 1,651 pages from ten document types such as academic papers, financial reports, newspapers and textbooks, with English, Chinese and mixed pages. olmOCR-Bench tests arXiv papers, old scans, tables and headers and footers, among others. German delivery notes with stamps, carbon copies and ballpoint annotations are not specifically represented in either.
  • They measure transcription, not extraction. Both benchmarks assess whether a page is converted into text correctly. Whether a model then assigns the right invoice number and gross amount from that text is a separate question.

Benchmarks help with the shortlist. The decision is made on a sample of your own documents, evaluated field by field.

The number problem: plausible is not read

Classic OCR makes visible errors: an "8" becomes a "B", a table row falls apart. A VLM makes invisible errors. It generates text like a language model, and where the original is illegible, it produces a value that fits the context.

Two studies document this:

Study Finding
KIE-HVQA, arXiv 2506.20168 With degraded originals (ID cards and invoices with simulated scanning defects), multimodal models rely on linguistic priors and produce hallucinations, especially when a precise answer is not possible.
Visual Merit or Linguistic Crutch, arXiv 2601.03714 Without linguistic context, DeepSeek-OCR's performance dropped from about 90 to 20 percent. Classic pipeline OCR was considerably more robust against such perturbations than end-to-end models. Fewer visual tokens correlated with a stronger reliance on language priors.

For invoices this is the critical point. An IBAN, an invoice number or an amount has no linguistic context from which the correct digit could be inferred. That is exactly where a plausible value is most dangerous.

Countermeasures at a glance:

Measure Effect
Structured output following a JSON schema The model can only return defined fields and types; vLLM (structured outputs) and Ollama (structured outputs) enforce the schema during generation.
Allow empty fields The schema permits null, and the instruction asks for null instead of a guess when something is illegible.
Arithmetic rules Net plus tax equals gross, sum of line items equals total. A contradiction indicates a reading error.
Check digits IBAN (MOD 97), VAT ID via VIES, invoice number format per supplier.
Second reading path Also read the amount from the PDF text layer or with Tesseract; review if they differ.
Master data matching Bank details and supplier against the supplier master, order number against the ERP.

The validation rules in detail, including VIES and IBAN checks, are described in the article Extract and validate documents with AI. The principle here: the model delivers candidates, the rules decide.

Pipeline patterns for invoices and scans

Before any model there is a routing step: what actually arrives as an image?

Input Route Model needed
XRechnung, ZUGFeRD parse and validate the XML no
Digital PDF with text layer read text and positions directly, e.g. with Docling possibly for field assignment
Scan, photo, PDF without text layer OCR or VLM yes

For the third case, three patterns have become established:

Pattern A: transcription, then extraction. A small OCR VLM (PaddleOCR-VL, GLM-OCR, olmOCR) converts the page into Markdown with tables. A language model extracts the fields from it as JSON. Advantage: the transcription is traceable and can be archived, and both steps can be tested separately. Disadvantage: two models in operation.

Pattern B: direct extraction. A general VLM such as Qwen3.8-27B receives the page image and a JSON schema and returns the fields directly. Advantage: one model, with access to the visual arrangement, for example which amount sits in the total line. Disadvantage: a larger model and no intermediate stage to which an error can be attributed.

Pattern C: Docling pipeline with extraction. Docling handles layout, tables and OCR, and a language model or VLM extracts from the structured document. This fits when contracts, manuals or reports are processed alongside receipts, for example for a RAG application.

In all three patterns the same validation rules and the same approval step follow. Which pattern fits depends on volume, original quality and available hardware, and is decided in a test with your own documents, not on a benchmark.

Hardware: 24 GB, 96 GB or AI Cube

For GPU memory the model weights come first: parameter count times bytes per parameter, i.e. 2 bytes in BF16, 1 byte in FP8 and about 0.5 bytes with 4-bit quantisation. KV cache, image tokens and runtime overhead come on top. The following values are this calculation for the weights, not a measurement.

Model Weights BF16 Weights FP8 24 GB (RTX PRO 4000 Blackwell) 96 GB (RTX PRO 6000 Blackwell Max-Q) AI Cube, 128 GB unified memory
granite-docling-258M about 0.5 GB - yes yes yes
PaddleOCR-VL-1.6, GLM-OCR, LightOnOCR-2-1B about 2 GB - yes yes yes
DeepSeek-OCR-2 about 7 GB - yes yes yes
olmOCR-2-7B-1025 about 17 GB about 9 GB yes, FP8; project states at least 12 GB yes yes
Qwen3.8-27B about 55 GB about 28 GB only 4-bit quantised yes, FP8 or BF16 yes

This leads to two typical set-ups. An OCR VLM around 1 billion parameters runs on a 24 GB GPU with plenty of headroom for parallel pages, and a medium-sized language model for extraction can run alongside it. Anyone who wants to use Qwen3.8-27B for direct extraction without heavy quantisation needs 96 GB or the AI Cube. How quantisation affects quality and memory is explained in the article on LLM quantization, and the calculation including KV cache is in VRAM sizing.

WZ-IT offers two ways to run this. The managed GPU server runs in a German data centre, either as GPU Server 24 with an NVIDIA RTX PRO 4000 Blackwell (24 GB GDDR7 ECC) from €699 net per month or as GPU Server 96 with an RTX PRO 6000 Blackwell Max-Q (96 GB GDDR7 ECC) for €1,799 net per month, each plus a setup fee. The AI Cube is an appliance in your own network, costs €6,490 net plus AI Cube Care at €349.90 net per month, and suits cases where documents should not leave the building. For the choice of inference software, see the comparison vLLM, Ollama or llama.cpp.

German documents: language, number format, original quality

The language information on the model cards differs considerably:

Model Language information according to model card
GLM-OCR including German, English, French, Spanish
LightOnOCR-2-1B including German, English, French, Italian, Dutch
PaddleOCR-VL 109 languages (first version), later "multilingual"
DeepSeek-OCR-2 "multilingual"
olmOCR-2-7B-1025 English
granite-docling-258M English

A missing language entry does not mean a model cannot read German; Latin script with umlauts is recognised by many models. It means the vendor does not promise it and it has to be verified in the test.

Three characteristics of German documents belong in every test sample:

  • Number format. 1.234,56 instead of 1,234.56. A model that has mostly seen English documents may swap thousands and decimal separators. Normalisation belongs in the code after the model, and the arithmetic rule catches remaining errors.
  • Umlauts and ß in company names and addresses, which must match exactly for supplier master matching.
  • Original quality. Faxes, carbon copies, delivery notes photographed at an angle, stamps over the amount. This is exactly where the risk of plausible rather than read values increases. PaddleOCR-VL-1.5 was explicitly developed further for skew, warping, screen photos and poor lighting, but that does not replace your own test.

Handwriting remains the hardest case. Some model cards advertise it, but none of the benchmarks named provides reliable figures for German handwriting on business documents.

Our approach at WZ-IT

  1. Define the sample. One document type, 50 to 100 representative documents including poor scans, with the correct values per document as reference.
  2. Build the routing. Handle e-invoices and digital PDFs with a text layer separately; only images and scans go to a model.
  3. Test two to three candidates. For example an OCR VLM with downstream extraction against a general VLM with direct extraction, licence checked in advance, on the target hardware.
  4. Evaluate field by field. Share of correct values per field, plus the share of documents without any correction. Invented values are counted separately.
  5. Validation rules and approval. Structured output, arithmetic rules, check digits, second reading path, master data matching, and approval for everything that fails a rule.
  6. Operation. Pin the model version, run a regression test with the sample before every model change, with support, consulting and implementation by WZ-IT.

The document processing pilot covers one document class with up to 20 target fields and starts at €17,900 net. If it is still open which process should be automated first, the AI process assessment evaluates document types and volumes. WZ-IT also runs Ollama and vLLM individually, see Ollama and vLLM.

Further guides

Which model reads your documents most reliably? We test candidates with a sample of your documents on your hardware and evaluate them field by field before a pipeline is built. Book a meeting

Sources

Enquiry

Check the model choice for your documents

We test OCR and vision-language models against a sample of your documents on your infrastructure and measure which fields are correct without manual correction.

What is your situation?

How should we get back to you?

Frequently Asked Questions

Answers to important questions about this topic

Classic OCR such as Tesseract recognises characters and returns text with positions. A vision-language model (VLM) processes the page as an image and generates text, Markdown or JSON like a language model. It recognises tables, reading order and fields without templates, but it can also produce values that do not appear on the page.

Specialised OCR models such as olmOCR-2-7B-1025, PaddleOCR-VL-1.6, GLM-OCR, LightOnOCR-2-1B and DeepSeek-OCR-2 convert pages into text or Markdown. General VLMs such as Qwen3.8-27B can also extract specific fields and follow instructions. Docling combines layout detection, OCR and VLMs in one pipeline. All models named here are licensed under Apache 2.0 or MIT.

No. Chandra OCR 2 is released under a modified OpenRAIL-M licence that excludes use by companies with more than 2 million US dollars in annual revenue, except for personal or research use. Nanonets-OCR2-3B states no licence on its model card and is built on Qwen2.5-VL-3B, which is under the non-commercial Qwen Research License. The licence has to be checked per model and per base model.

Only to a limited extent. OmniDocBench contains English, Chinese and mixed pages from papers, reports, newspapers and textbooks, and its overall score averages text, tables and formulas. olmOCR-Bench is also predominantly English. German or other European invoices with decimal commas, stamps and poor scans are not specifically represented in either benchmark. The scores are vendor-reported and do not replace a test with your own documents.

The risk is documented. With blurred or damaged originals, models fall back on linguistic probabilities and output plausible rather than read values, as the study behind the KIE-HVQA benchmark with invoices and ID cards shows. A study of DeepSeek-OCR found that recognition performance fell from about 90 to 20 percent without linguistic context. Amounts, IBANs and invoice numbers therefore need validation rules outside the model.

No. In the DeepSeek-OCR study, classic pipeline OCR proved more robust against semantically corrupted text than end-to-end models. Tesseract and the text layer of a digital PDF remain useful as a second, independent reading path: if the amount returned by the VLM differs from the OCR string, the document goes to review.

Models around 1 billion parameters such as PaddleOCR-VL or GLM-OCR need about 2 GB for their weights in BF16 and run on a 24 GB GPU with plenty of headroom. The olmOCR project states a minimum of 12 GB of GPU memory. Qwen3.8-27B needs about 55 GB in BF16 and about 28 GB in FP8 for its weights, so it fits a 96 GB GPU or the AI Cube with 128 GB unified memory, and a 24 GB GPU only with 4-bit quantisation.

No. XRechnung and the XML part of a ZUGFeRD invoice are parsed and validated against the EN 16931 rules, not read. For ZUGFeRD, the structured part takes precedence. OCR and VLMs concern paper, scans, photos and PDF files without a structured part. The full process with e-invoicing, validation rules and German bookkeeping requirements is covered in the article on document processing.

No. The model cards of the OCR VLMs named here do not describe a calibrated per-field confidence score, and a certainty stated by the model itself is not a measured accuracy. Reliable signals are arithmetic rules, format checks such as the IBAN check digit, matching against master data and comparison with a second reading path.

Timo Wevelsiep

Written by

Timo Wevelsiep

Co-Founder & CEO

Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.

LinkedIn

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.