Local Vision-Language Models for OCR: Reading Invoices and Scans

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

Read documents locally instead of sending them to an external AI service? WZ-IT builds the extraction pipeline with open models and validation rules on your infrastructure, see AI document processing in the AI hub. Book a meeting
Since 2025 a new group of open models has appeared that read document pages as images and output text, Markdown or JSON: olmOCR, PaddleOCR-VL, GLM-OCR, DeepSeek-OCR, LightOnOCR and general vision-language models such as Qwen3.8. Almost all of them run on a single GPU, and many are licensed under Apache 2.0 or MIT. For invoices, delivery notes and scanned letters this is a real alternative to cloud services, with no document leaving your own network.
This article compares the models from an operator's perspective: what they do, which licence they are under, how to read the benchmark scores on model cards, what hardware they need and how to limit the risk of invented numbers. The overall process with e-invoicing, validation rules, approval and German bookkeeping requirements (GoBD) is covered in the article Extract and validate documents with AI. All information is as of September 2026.
Table of Contents
- Three tool classes: OCR, OCR VLM, general VLM
- Open models at a glance, as of September 2026
- Licences: not every model can be used commercially
- How to read benchmarks
- The number problem: plausible is not read
- Pipeline patterns for invoices and scans
- Hardware: 24 GB, 96 GB or AI Cube
- German documents: language, number format, original quality
- Our approach at WZ-IT
- Further guides
Three tool classes: OCR, OCR VLM, general VLM
| Class | Examples | Input | Output | Typical error |
|---|---|---|---|---|
| Classic OCR | Tesseract, PaddleOCR (PP-OCR), OCRmyPDF | image, scan | characters with positions, text layer in the PDF | visible character errors, broken tables |
| OCR VLM, specialised for documents | olmOCR 2, PaddleOCR-VL, GLM-OCR, LightOnOCR-2, DeepSeek-OCR-2, granite-docling | page image | Markdown, HTML, tables, partly JSON | omitted or added passages, fluent text instead of a gap |
| General VLM | Qwen3.8-27B, Qwen3-VL | page image plus instruction | free answer or JSON following a schema | plausible but unread values |
Classic OCR reads characters but does not understand structure. A field such as "invoice number" then requires rules or templates. OCR VLMs are small models, usually between 0.3 and 8 billion parameters, trained to transcribe a page completely and in reading order, including tables. General VLMs are large multimodal language models that process an image and an instruction. They can extract specific fields but are slower and more prone to filling gaps with something plausible.
Docling belongs to none of the three classes but connects them: the MIT-licensed project of the LF AI & Data Foundation detects layout, reading order and table structure, integrates classic OCR for scans and can optionally use a VLM such as granite-docling (docling --pipeline vlm). The current version is 2.131.0 from 29 September 2026 (Docling releases).
Open models at a glance, as of September 2026
| Model | Vendor | Parameters | Licence | Released | Base | Output |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | Alibaba Qwen | 27 billion | Apache 2.0 | August 2026 | own architecture, natively multimodal | free answer, JSON |
| olmOCR-2-7B-1025 | Allen Institute for AI | 7 billion (base) | Apache 2.0 | October 2025 | Qwen2.5-VL-7B-Instruct | Markdown |
| PaddleOCR-VL-1.6 | Baidu PaddlePaddle | 0.9 billion | Apache 2.0 | 28 May 2026 | ERNIE 4.5 0.3B | Markdown, JSON with layout |
| GLM-OCR | Z.ai | 0.9 billion | MIT | January 2026 | GLM-V, GLM-0.5B decoder | Markdown, JSON following a schema |
| LightOnOCR-2-1B | LightOn | 1 billion | Apache 2.0 | January 2026 | own architecture | text in reading order |
| DeepSeek-OCR-2 | DeepSeek | about 3 billion | Apache 2.0 | January 2026 | own architecture | Markdown |
| granite-docling-258M | IBM Research | 258 million | Apache 2.0 | 17 September 2025 | Idefics3, Granite 165M | DocTags for Docling |
How the candidates compare:
- Qwen3.8-27B is not a specialised OCR model but a general language model with image understanding. The model card explicitly lists documents as a use case, a context length of 262,144 tokens and a thinking mode that can be switched off per request. Switching it off makes sense for field extraction because long reasoning chains increase response time. Qwen provides an FP8 variant itself (Qwen3.8-27B-FP8).
- olmOCR 2 is designed to convert whole PDF pages into text and runs through the olmOCR toolkit with vLLM, which renders, rotates and retries pages as needed. Allen AI recommends the FP8 variant for practical use (model card). Training rewards passing unit tests on documents (olmOCR 2, arXiv 2510.19817).
- PaddleOCR-VL-1.6 has the same architecture as version 1.5 and can be swapped in without changes. The first version listed 109 supported languages (PaddleOCR-VL), and version 1.5 added seal and stamp recognition (PaddleOCR-VL-1.5).
- GLM-OCR combines layout analysis (PP-DocLayout-V3) with parallel recognition and, besides transcription, supports information extraction in which the prompt specifies a JSON schema. It runs on vLLM, SGLang and Ollama.
- LightOnOCR-2-1B explicitly lists German among its supported languages and is served directly with
vllm serve. - granite-docling-258M is the smallest model in the list and is built for the Docling pipeline. The model card lists English as its language.
Licences: not every model can be used commercially
Hugging Face hosts further OCR models with good benchmark scores whose licence rules out or leaves open use in a company. Two examples:
| Model | Licence according to model card | Consequence for companies |
|---|---|---|
| Chandra OCR 2 | modified OpenRAIL-M | use excluded for more than 2 million US dollars in prior-year revenue or more than 2 million US dollars in funding, except personal or research use; commercial licence from the vendor (LICENSE, Attachment A) |
| Nanonets-OCR2-3B | not stated | base model Qwen2.5-VL-3B-Instruct is under the Qwen Research License, which only permits non-commercial use (licence text) |
Three rules for the check:
- Separate model licence and code licence. For Chandra the code is Apache 2.0, the weights are not. The weights are what matters for operation.
- Check the base model. A fine-tune inherits the terms of its base model. olmOCR 2 builds on Qwen2.5-VL-7B, which is under Apache 2.0. For the 3B variant of the same base model that would not be the case.
- Check pipeline components individually. Docling itself is MIT, while the integrated models are subject to their own licences, as the project explicitly states.
A detailed overview of licence models for open language models is provided in the article Open LLM licences for commercial use. This is a technical classification, not legal advice.
How to read benchmarks
Every model card reports top scores. The figures come from the vendors themselves and refer to different benchmarks and versions.
| Model | Benchmark | Score according to vendor |
|---|---|---|
| PaddleOCR-VL-1.6 | OmniDocBench v1.6 | 96.33 |
| GLM-OCR | OmniDocBench v1.5 | 94.62 |
| PaddleOCR-VL-1.5 | OmniDocBench v1.5 | 94.5 |
| Qwen3.8-27B | OmniDocBench 1.5 | 91.1 |
| Chandra OCR 2 | olmOCR-Bench | 85.8 |
| olmOCR-2-7B-1025-FP8 | olmOCR-Bench, toolkit v0.4.0 | 82.4 ± 1.1 |
Sources: model cards of PaddleOCR-VL-1.6, GLM-OCR, PaddleOCR-VL-1.5, Qwen3.8-27B, Chandra OCR 2 and olmOCR 2.
What the table does not show:
- Versions are not comparable. OmniDocBench was updated to v1.6 in April 2026, with 296 additional, harder pages and a new matching method, and to v1.7 shortly afterwards. A v1.5 score and a v1.6 score do not measure the same thing.
- The overall score includes formulas. The OmniDocBench overall score is the mean of text accuracy, table structure (TEDS) and formula recognition (CDM). Formulas are irrelevant for invoices and central for scientific papers.
- The documents are not yours. In its base version OmniDocBench comprises 1,651 pages from ten document types such as academic papers, financial reports, newspapers and textbooks, with English, Chinese and mixed pages. olmOCR-Bench tests arXiv papers, old scans, tables and headers and footers, among others. German delivery notes with stamps, carbon copies and ballpoint annotations are not specifically represented in either.
- They measure transcription, not extraction. Both benchmarks assess whether a page is converted into text correctly. Whether a model then assigns the right invoice number and gross amount from that text is a separate question.
Benchmarks help with the shortlist. The decision is made on a sample of your own documents, evaluated field by field.
The number problem: plausible is not read
Classic OCR makes visible errors: an "8" becomes a "B", a table row falls apart. A VLM makes invisible errors. It generates text like a language model, and where the original is illegible, it produces a value that fits the context.
Two studies document this:
| Study | Finding |
|---|---|
| KIE-HVQA, arXiv 2506.20168 | With degraded originals (ID cards and invoices with simulated scanning defects), multimodal models rely on linguistic priors and produce hallucinations, especially when a precise answer is not possible. |
| Visual Merit or Linguistic Crutch, arXiv 2601.03714 | Without linguistic context, DeepSeek-OCR's performance dropped from about 90 to 20 percent. Classic pipeline OCR was considerably more robust against such perturbations than end-to-end models. Fewer visual tokens correlated with a stronger reliance on language priors. |
For invoices this is the critical point. An IBAN, an invoice number or an amount has no linguistic context from which the correct digit could be inferred. That is exactly where a plausible value is most dangerous.
Countermeasures at a glance:
| Measure | Effect |
|---|---|
| Structured output following a JSON schema | The model can only return defined fields and types; vLLM (structured outputs) and Ollama (structured outputs) enforce the schema during generation. |
| Allow empty fields | The schema permits null, and the instruction asks for null instead of a guess when something is illegible. |
| Arithmetic rules | Net plus tax equals gross, sum of line items equals total. A contradiction indicates a reading error. |
| Check digits | IBAN (MOD 97), VAT ID via VIES, invoice number format per supplier. |
| Second reading path | Also read the amount from the PDF text layer or with Tesseract; review if they differ. |
| Master data matching | Bank details and supplier against the supplier master, order number against the ERP. |
The validation rules in detail, including VIES and IBAN checks, are described in the article Extract and validate documents with AI. The principle here: the model delivers candidates, the rules decide.
Pipeline patterns for invoices and scans
Before any model there is a routing step: what actually arrives as an image?
| Input | Route | Model needed |
|---|---|---|
| XRechnung, ZUGFeRD | parse and validate the XML | no |
| Digital PDF with text layer | read text and positions directly, e.g. with Docling | possibly for field assignment |
| Scan, photo, PDF without text layer | OCR or VLM | yes |
For the third case, three patterns have become established:
Pattern A: transcription, then extraction. A small OCR VLM (PaddleOCR-VL, GLM-OCR, olmOCR) converts the page into Markdown with tables. A language model extracts the fields from it as JSON. Advantage: the transcription is traceable and can be archived, and both steps can be tested separately. Disadvantage: two models in operation.
Pattern B: direct extraction. A general VLM such as Qwen3.8-27B receives the page image and a JSON schema and returns the fields directly. Advantage: one model, with access to the visual arrangement, for example which amount sits in the total line. Disadvantage: a larger model and no intermediate stage to which an error can be attributed.
Pattern C: Docling pipeline with extraction. Docling handles layout, tables and OCR, and a language model or VLM extracts from the structured document. This fits when contracts, manuals or reports are processed alongside receipts, for example for a RAG application.
In all three patterns the same validation rules and the same approval step follow. Which pattern fits depends on volume, original quality and available hardware, and is decided in a test with your own documents, not on a benchmark.
Hardware: 24 GB, 96 GB or AI Cube
For GPU memory the model weights come first: parameter count times bytes per parameter, i.e. 2 bytes in BF16, 1 byte in FP8 and about 0.5 bytes with 4-bit quantisation. KV cache, image tokens and runtime overhead come on top. The following values are this calculation for the weights, not a measurement.
| Model | Weights BF16 | Weights FP8 | 24 GB (RTX PRO 4000 Blackwell) | 96 GB (RTX PRO 6000 Blackwell Max-Q) | AI Cube, 128 GB unified memory |
|---|---|---|---|---|---|
| granite-docling-258M | about 0.5 GB | - | yes | yes | yes |
| PaddleOCR-VL-1.6, GLM-OCR, LightOnOCR-2-1B | about 2 GB | - | yes | yes | yes |
| DeepSeek-OCR-2 | about 7 GB | - | yes | yes | yes |
| olmOCR-2-7B-1025 | about 17 GB | about 9 GB | yes, FP8; project states at least 12 GB | yes | yes |
| Qwen3.8-27B | about 55 GB | about 28 GB | only 4-bit quantised | yes, FP8 or BF16 | yes |
This leads to two typical set-ups. An OCR VLM around 1 billion parameters runs on a 24 GB GPU with plenty of headroom for parallel pages, and a medium-sized language model for extraction can run alongside it. Anyone who wants to use Qwen3.8-27B for direct extraction without heavy quantisation needs 96 GB or the AI Cube. How quantisation affects quality and memory is explained in the article on LLM quantization, and the calculation including KV cache is in VRAM sizing.
WZ-IT offers two ways to run this. The managed GPU server runs in a German data centre, either as GPU Server 24 with an NVIDIA RTX PRO 4000 Blackwell (24 GB GDDR7 ECC) from €699 net per month or as GPU Server 96 with an RTX PRO 6000 Blackwell Max-Q (96 GB GDDR7 ECC) for €1,799 net per month, each plus a setup fee. The AI Cube is an appliance in your own network, costs €6,490 net plus AI Cube Care at €349.90 net per month, and suits cases where documents should not leave the building. For the choice of inference software, see the comparison vLLM, Ollama or llama.cpp.
German documents: language, number format, original quality
The language information on the model cards differs considerably:
| Model | Language information according to model card |
|---|---|
| GLM-OCR | including German, English, French, Spanish |
| LightOnOCR-2-1B | including German, English, French, Italian, Dutch |
| PaddleOCR-VL | 109 languages (first version), later "multilingual" |
| DeepSeek-OCR-2 | "multilingual" |
| olmOCR-2-7B-1025 | English |
| granite-docling-258M | English |
A missing language entry does not mean a model cannot read German; Latin script with umlauts is recognised by many models. It means the vendor does not promise it and it has to be verified in the test.
Three characteristics of German documents belong in every test sample:
- Number format. 1.234,56 instead of 1,234.56. A model that has mostly seen English documents may swap thousands and decimal separators. Normalisation belongs in the code after the model, and the arithmetic rule catches remaining errors.
- Umlauts and ß in company names and addresses, which must match exactly for supplier master matching.
- Original quality. Faxes, carbon copies, delivery notes photographed at an angle, stamps over the amount. This is exactly where the risk of plausible rather than read values increases. PaddleOCR-VL-1.5 was explicitly developed further for skew, warping, screen photos and poor lighting, but that does not replace your own test.
Handwriting remains the hardest case. Some model cards advertise it, but none of the benchmarks named provides reliable figures for German handwriting on business documents.
Our approach at WZ-IT
- Define the sample. One document type, 50 to 100 representative documents including poor scans, with the correct values per document as reference.
- Build the routing. Handle e-invoices and digital PDFs with a text layer separately; only images and scans go to a model.
- Test two to three candidates. For example an OCR VLM with downstream extraction against a general VLM with direct extraction, licence checked in advance, on the target hardware.
- Evaluate field by field. Share of correct values per field, plus the share of documents without any correction. Invented values are counted separately.
- Validation rules and approval. Structured output, arithmetic rules, check digits, second reading path, master data matching, and approval for everything that fails a rule.
- Operation. Pin the model version, run a regression test with the sample before every model change, with support, consulting and implementation by WZ-IT.
The document processing pilot covers one document class with up to 20 target fields and starts at €17,900 net. If it is still open which process should be automated first, the AI process assessment evaluates document types and volumes. WZ-IT also runs Ollama and vLLM individually, see Ollama and vLLM.
Further guides
- Extract and validate documents with AI, the overall process with e-invoicing, validation rules, approval and GoBD.
- Process documents with AI: from OCR to IDP, the basics of intelligent document processing.
- Classify e-mails and tickets locally with AI, the step before: sorting incoming items before documents are extracted.
- Local speech-to-text under GDPR, the same question for audio instead of images.
- Which LLM to self-host, choosing open language models for extraction.
- Install Paperless-ngx on Ubuntu, the archive for extracted documents.
- AI solutions from WZ-IT, the hub with all offerings for local AI.
Which model reads your documents most reliably? We test candidates with a sample of your documents on your hardware and evaluate them field by field before a pipeline is built. Book a meeting
Sources
- Qwen3.8-27B, model card
- Qwen3.8-27B-FP8, model card
- olmOCR-2-7B-1025, model card
- olmOCR-2-7B-1025-FP8, model card
- olmOCR toolkit on GitHub
- olmOCR, Allen AI project page
- olmOCR 2: Unit Test Rewards for Document OCR, arXiv 2510.19817
- olmOCR-Bench, dataset
- PaddleOCR-VL-1.6, model card
- PaddleOCR-VL-1.5, model card
- PaddleOCR-VL, model card of the first version
- PaddleOCR on GitHub
- GLM-OCR, model card
- GLM-OCR in the Ollama library
- LightOnOCR-2-1B, model card
- DeepSeek-OCR-2, model card
- granite-docling-258M, model card
- Docling on GitHub
- Docling releases
- Chandra OCR 2, model card
- Chandra OCR 2, licence text
- Nanonets-OCR2-3B, model card
- Qwen2.5-VL-3B-Instruct, Qwen Research License
- OmniDocBench on GitHub
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models, arXiv 2506.20168
- Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR, arXiv 2601.03714
- vLLM, structured outputs
- Ollama, structured outputs
- Tesseract, trained models tessdata_best
- OCRmyPDF on GitHub
Check the model choice for your documents
We test OCR and vision-language models against a sample of your documents on your infrastructure and measure which fields are correct without manual correction.
Frequently Asked Questions
Answers to important questions about this topic
Classic OCR such as Tesseract recognises characters and returns text with positions. A vision-language model (VLM) processes the page as an image and generates text, Markdown or JSON like a language model. It recognises tables, reading order and fields without templates, but it can also produce values that do not appear on the page.
Specialised OCR models such as olmOCR-2-7B-1025, PaddleOCR-VL-1.6, GLM-OCR, LightOnOCR-2-1B and DeepSeek-OCR-2 convert pages into text or Markdown. General VLMs such as Qwen3.8-27B can also extract specific fields and follow instructions. Docling combines layout detection, OCR and VLMs in one pipeline. All models named here are licensed under Apache 2.0 or MIT.
No. Chandra OCR 2 is released under a modified OpenRAIL-M licence that excludes use by companies with more than 2 million US dollars in annual revenue, except for personal or research use. Nanonets-OCR2-3B states no licence on its model card and is built on Qwen2.5-VL-3B, which is under the non-commercial Qwen Research License. The licence has to be checked per model and per base model.
Only to a limited extent. OmniDocBench contains English, Chinese and mixed pages from papers, reports, newspapers and textbooks, and its overall score averages text, tables and formulas. olmOCR-Bench is also predominantly English. German or other European invoices with decimal commas, stamps and poor scans are not specifically represented in either benchmark. The scores are vendor-reported and do not replace a test with your own documents.
The risk is documented. With blurred or damaged originals, models fall back on linguistic probabilities and output plausible rather than read values, as the study behind the KIE-HVQA benchmark with invoices and ID cards shows. A study of DeepSeek-OCR found that recognition performance fell from about 90 to 20 percent without linguistic context. Amounts, IBANs and invoice numbers therefore need validation rules outside the model.
No. In the DeepSeek-OCR study, classic pipeline OCR proved more robust against semantically corrupted text than end-to-end models. Tesseract and the text layer of a digital PDF remain useful as a second, independent reading path: if the amount returned by the VLM differs from the OCR string, the document goes to review.
Models around 1 billion parameters such as PaddleOCR-VL or GLM-OCR need about 2 GB for their weights in BF16 and run on a 24 GB GPU with plenty of headroom. The olmOCR project states a minimum of 12 GB of GPU memory. Qwen3.8-27B needs about 55 GB in BF16 and about 28 GB in FP8 for its weights, so it fits a 96 GB GPU or the AI Cube with 128 GB unified memory, and a 24 GB GPU only with 4-bit quantisation.
No. XRechnung and the XML part of a ZUGFeRD invoice are parsed and validated against the EN 16931 rules, not read. For ZUGFeRD, the structured part takes precedence. OCR and VLMs concern paper, scans, photos and PDF files without a structured part. The full process with e-invoicing, validation rules and German bookkeeping requirements is covered in the article on document processing.
No. The model cards of the OCR VLMs named here do not describe a calibrated per-field confidence score, and a certainty stated by the model itself is not a measured accuracy. Reliable signals are arithmetic rules, format checks such as the IBAN check digit, matching against master data and comparison with a second reading path.

Written by
Timo Wevelsiep
Co-Founder & CEO
Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.
LinkedInLet's Talk About Your Idea
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.





