Paperless-ngx AI Setup: Native AI from Version 3 with Local Ollama

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

Using Paperless-ngx with AI? WZ-IT sets up the native AI in Paperless-ngx, connects local models and operates the instance long term if required, see Paperless-ngx at WZ-IT and Paperless-AI and AI integration. Book a meeting
Until recently, using Paperless-ngx with AI meant adding a separate project such as Paperless-AI. Since version 3.0.0, released on 22 July 2026, Paperless-ngx ships AI itself: suggestions for title, tags, correspondent and document type, a chat about a single document or the whole archive, and a vector index for similar documents. Version 3.1.0 added a workflow action that applies these suggestions automatically.
At the same time, the Paperless-AI repository has carried the notice "This repo is currently not maintained" since March 2026 (clusterzx/paperless-ai). For new installations the native integration is the obvious route; for existing Paperless-AI setups the question is how to switch.
This post covers the setup with a local Ollama server via Docker Compose, every AI variable with its default, model choice for non-English archives, automatic tagging via workflows, data flows and the migration from Paperless-AI. All details refer to Paperless-ngx 3.2.1 (as of October 2026).
Table of Contents
- Key facts at a glance
- What the native AI in Paperless-ngx does
- What has changed since version 3
- Requirements and hardware
- Step by step: Paperless-ngx with Ollama via Docker Compose
- All AI variables at a glance
- Choosing generation and embedding models
- Automatic tagging with the workflow action
- Privacy: which data leaves the system
- Native AI vs. Paperless-AI compared
- Moving off Paperless-AI
- Limitations as of October 2026
- Our approach at WZ-IT
- Further guides
Key facts at a glance
| Question | Answer (Paperless-ngx 3.2.1) |
|---|---|
| From which version? | 3.0.0 (22 Jul 2026); workflow action from 3.1.0 (27 Aug 2026) |
| Enabled by default? | no, PAPERLESS_AI_ENABLED is false |
| Backends for the language model | ollama, openai-like |
| Backends for embeddings | ollama, huggingface, openai-like |
| Features | AI suggestions, document chat, similar documents, "Apply AI Suggestions" workflow action |
| Classic classifier | stays active, AI suggestions are added |
| Configuration | PAPERLESS_AI_* environment variables or the UI; the database value takes precedence |
| Vector index | local in the data directory (data/llm_index), SQLite with sqlite-vec |
| Data flow | document content goes to the configured model; it stays local only with a local backend |
Sources: Configuration, AI section, Advanced Usage, AI features, releases, source code vector_store.py.
What the native AI in Paperless-ngx does
The documentation describes three features. All are optional and none replaces the existing non-LLM matching (Advanced Usage):
| Feature | What it does | Requirement |
|---|---|---|
| AI suggestions | Suggests title, tags, correspondent, document type, storage path and dates; appears in the menu of the suggest button next to the classic suggestions | AI enabled, LLM backend set |
| Similar documents and RAG | Vector index over text and metadata; suggestions are grounded in similar existing documents | embedding backend set as well |
| Document chat | Questions about a single document or all visible documents, with links to the source documents | LLM index active |
| "Apply AI Suggestions" workflow action | applies suggestions automatically in the background | from 3.1.0, AI enabled |
Two details matter in operation:
- Suggestions on open. For documents with an inbox tag, suggestions are requested automatically as soon as the document is opened. Since 3.2.0 this can be switched off under Settings > Documents (Usage, Document Suggestions). With a local model, every open is a request to the GPU.
- Permissions. The chat only searches documents the signed-in user may see. The source code filters the vector search to the permitted document IDs (chat.py). Since 3.0.0, suggestions require change permission on the document (release 3.0.0).
What has changed since version 3
| Version | Date | Relevant AI changes |
|---|---|---|
| 3.0.0 | 22 Jul 2026 | Paperless AI (suggestions, chat, LLM index), Ollama embeddings, configurable timeout, chunk and context size, output language, Remote OCR via Azure AI |
| 3.0.5 | 1 Aug 2026 | suggestion cache keyed by model and endpoint, better handling of empty fields |
| 3.1.0 | 27 Aug 2026 | "Apply AI Suggestions" workflow action; suggestions prefer existing tags, types, correspondents and storage paths; Remote OCR can be limited to selected documents |
| 3.1.3 | 4 Sep 2026 | workflow action reliably runs after the document is created; Remote OCR endpoint is validated |
| 3.2.0 | 19 Sep 2026 | automatic suggestions for inbox documents can be disabled; documents without text are skipped by the workflow action; better LLM error messages |
| 3.2.1 | 20 Sep 2026 | bug fixes unrelated to AI (mail fetch, OCR, search index) |
If you are coming from 2.x, read the breaking changes of 3.0.0 before upgrading. They include, among others, the removal of document encryption, of API version 1 and of Python 3.10 support (release 3.0.0). For Docker installations the usual rule applies: back up first, then update the image.
Requirements and hardware
- Paperless-ngx 3.1.0 or later, preferably the current 3.2.1 if you want the workflow action. The guide Paperless-ngx on Ubuntu with Caddy shows a Docker installation with Caddy.
- Ollama as a container on the same host or as a separate server on the internal network (Ollama, Docker). On a separate server, the Ollama port should only be reachable from the Paperless network. The post on Bleeding Llama explains why an openly reachable Ollama is a risk.
- A GPU for usable response times. Ollama also runs on CPU, but suggestions then take considerably longer. NVIDIA GPUs in containers require the NVIDIA Container Toolkit.
As a rough guide, the model plus its context has to fit into GPU memory. The download sizes from the Ollama library are the lower bound:
| Model (Ollama) | Purpose | Download | Context window |
|---|---|---|---|
| gemma3:4b | generation, small | 3.3 GB | 128K |
| gemma3:12b | generation, medium | 8.1 GB | 128K |
| qwen3:8b | generation, medium | 5.2 GB | 40K |
| llama3.1:8b | Paperless-ngx default for Ollama | 4.9 GB | 128K |
| bge-m3 | embeddings, multilingual | 1.2 GB | 8K |
| embeddinggemma | Paperless-ngx default embedding for Ollama | 622 MB | 2K |
By default, Paperless-ngx sends a context of 8,192 tokens to Ollama (PAPERLESS_AI_LLM_CONTEXT_SIZE, sent as num_ctx for Ollama). Memory usage is therefore above the download size. The article Sizing GPU VRAM for LLMs explains how to calculate GPU memory for models and context.
Step by step: Paperless-ngx with Ollama via Docker Compose
The starting point is the official docker-compose.postgres.yml. It is extended with an Ollama service and the AI variables in the webserver service.
1. Extend the Compose file
services:
broker:
image: docker.io/valkey/valkey:9-alpine
restart: unless-stopped
volumes:
- redisdata:/data
db:
image: docker.io/library/postgres:18
restart: unless-stopped
volumes:
- pgdata:/var/lib/postgresql
environment:
POSTGRES_DB: paperless
POSTGRES_USER: paperless
POSTGRES_PASSWORD: paperless # change in production
ollama:
image: docker.io/ollama/ollama:latest
restart: unless-stopped
volumes:
- ollama:/root/.ollama
# no "ports:" needed, Paperless reaches Ollama via the Compose network
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
webserver:
image: ghcr.io/paperless-ngx/paperless-ngx:latest # pin a version in production
restart: unless-stopped
depends_on:
- db
- broker
- ollama
ports:
- "8000:8000"
volumes:
- data:/usr/src/paperless/data
- media:/usr/src/paperless/media
- ./export:/usr/src/paperless/export
- ./consume:/usr/src/paperless/consume
env_file: docker-compose.env
environment:
PAPERLESS_REDIS: redis://broker:6379
PAPERLESS_DBHOST: db
PAPERLESS_DBENGINE: postgresql
# AI
PAPERLESS_AI_ENABLED: "true"
PAPERLESS_AI_LLM_BACKEND: ollama
PAPERLESS_AI_LLM_ENDPOINT: http://ollama:11434
PAPERLESS_AI_LLM_MODEL: gemma3:12b
PAPERLESS_AI_LLM_EMBEDDING_BACKEND: ollama
PAPERLESS_AI_LLM_EMBEDDING_MODEL: bge-m3
PAPERLESS_AI_LLM_REQUEST_TIMEOUT: "300"
volumes:
data:
media:
pgdata:
redisdata:
ollama:
The deploy.resources.reservations.devices block reserves the NVIDIA GPU for the Ollama container (Docker, GPU support in Compose). Remove it if there is no GPU. If Ollama runs on another server, drop the ollama service and point PAPERLESS_AI_LLM_ENDPOINT to its internal address.
2. Start and pull the models
docker compose pull
docker compose up -d
docker compose exec ollama ollama pull gemma3:12b
docker compose exec ollama ollama pull bge-m3
3. Build the LLM index
Rebuild the index from scratch when enabling it for the first time and after every change of the embedding model (Administration, LLM index):
docker compose exec webserver document_llmindex rebuild
After that, a scheduled task keeps the index up to date, by default daily at 02:10 (PAPERLESS_LLM_INDEX_TASK_CRON, default 10 2 * * *). For large archives the first build takes correspondingly long; embeddings run on the same GPU as the suggestions.
4. Check in the UI
- Open a document and expand the menu next to the suggest button. AI suggestions appear there in addition to the classic suggestions.
- Open the chat in the top toolbar. On the detail page it refers to the open document, in lists to all visible documents.
The AI settings can also be set under Settings > Application Configuration. Database values take precedence over environment variables (Advanced Usage). If you manage the configuration in Compose, leave the fields in the UI empty, otherwise a change in the Compose file has no effect.
5. On updates
Since version 3, migrating the LLM index is part of the update procedure. In the Docker image, document_llmindex migrate runs automatically when the container starts (init script init-llmindex-migrate). On bare-metal installations, the command is a separate step after the database migration (Administration, Updating).
The index lives in the data volume under llm_index and is therefore part of the existing backup.
All AI variables at a glance
| Variable | Default | Meaning |
|---|---|---|
PAPERLESS_AI_ENABLED |
false |
master switch for all AI features |
PAPERLESS_AI_LLM_BACKEND |
none | ollama or openai-like; no AI without this value |
PAPERLESS_AI_LLM_MODEL |
none; internally llama3.1 (Ollama) or gpt-3.5-turbo (openai-like) |
model for suggestions and chat |
PAPERLESS_AI_LLM_ENDPOINT |
none | backend URL; required for Ollama |
PAPERLESS_AI_LLM_API_KEY |
none | API key, usually only for openai-like |
PAPERLESS_AI_LLM_EMBEDDING_BACKEND |
none | ollama, huggingface or openai-like; enables the LLM index |
PAPERLESS_AI_LLM_EMBEDDING_MODEL |
none; internally embeddinggemma (Ollama), sentence-transformers/all-MiniLM-L6-v2 (Hugging Face), text-embedding-3-small (openai-like) |
embedding model |
PAPERLESS_AI_LLM_EMBEDDING_ENDPOINT |
value of PAPERLESS_AI_LLM_ENDPOINT |
separate endpoint for embeddings |
PAPERLESS_AI_LLM_EMBEDDING_CHUNK_SIZE |
1024 |
size of text chunks for embeddings |
PAPERLESS_AI_LLM_CONTEXT_SIZE |
8192 |
context size for prompts and RAG; sent as num_ctx for Ollama |
PAPERLESS_AI_LLM_REQUEST_TIMEOUT |
120 |
timeout in seconds per request |
PAPERLESS_AI_LLM_OUTPUT_LANGUAGE |
none (UI language) | language of the suggestions |
PAPERLESS_AI_LLM_ALLOW_INTERNAL_ENDPOINTS |
true |
false blocks endpoints with internal addresses |
PAPERLESS_LLM_INDEX_TASK_CRON |
10 2 * * * |
schedule for updating the index |
Source: Paperless-ngx configuration, AI section, version 3.2.1.
On PAPERLESS_AI_LLM_ALLOW_INTERNAL_ENDPOINTS: for a local Ollama the value has to stay true, otherwise Paperless-ngx blocks addresses such as http://ollama:11434. Set it to false only if you exclusively use an external provider and want to rule out requests into the internal network.
Choosing generation and embedding models
Paperless-ngx uses two separate models. The project wiki gives guidance, explicitly without benchmarks (AI Model Recommendations):
- Generation model (
PAPERLESS_AI_LLM_MODEL): an instruction or chat model that reliably produces structured output. Start with a small model and only move up if quality is insufficient. - Embedding model (
PAPERLESS_AI_LLM_EMBEDDING_MODEL): the Hugging Face defaultall-MiniLM-L6-v2is meant for primarily English archives. For multilingual archives the wiki listsintfloat/multilingual-e5-small,intfloat/multilingual-e5-baseandBAAI/bge-m3.
For a German or mixed-language archive this means:
| Decision | Recommendation |
|---|---|
| Embeddings via Ollama | bge-m3 (multilingual, 8K context) instead of the default embeddinggemma with 2K context |
| Embeddings without Ollama | huggingface with intfloat/multilingual-e5-small; runs inside the Paperless container and downloads the model on first start |
| Generation, starting point | a multilingual mid-size model such as gemma3:12b or qwen3:8b, tested on your own documents |
| Timeouts | according to the wiki, try a smaller model first, then raise PAPERLESS_AI_LLM_REQUEST_TIMEOUT |
| Changing the embedding model | rebuild the index with document_llmindex rebuild |
Test with a sample of typical documents: invoices, contracts, letters from authorities, poor-quality scans. The article Which LLM to self-host? helps to classify models for self-hosting. If text recognition itself is the problem, see the comparison of local VLM OCR models.
Automatic tagging with the workflow action
Suggestions alone do not change a document. For automatic assignment, the "Apply AI Suggestions" workflow action has existed since 3.1.0 (Usage, Workflows). It requests the same suggestions as the document detail page and applies them.
| Option | Behaviour |
|---|---|
| Fields | title, tags, correspondent, document type, storage path and created date, selectable individually; suggestions for unselected fields are discarded |
| Create missing items | off by default: only existing tags, correspondents and document types are assigned; storage paths are never created |
| Overwrite existing values | off by default: only empty fields are filled; title and date are almost always set, so overwriting has to be enabled for them |
| Tags | always added, never replaced |
| Triggers | all except "Consumption Started", because the text only exists after processing |
| Execution | queued as a separate background task; documents without text are skipped |
A proven setup to start with:
- Trigger "Document Added", filtered on the inbox tag or an intake folder.
- Action "Apply AI Suggestions" with correspondent, document type and tags, "create missing items" off.
- Second action (assignment): a tag such as
ai-suggested, so that AI-processed documents can be reviewed specifically. - After a few weeks, review samples and only then add title and date with overwriting.
The documentation points out that every matching document triggers a request to the model. A workflow across the whole archive can occupy the task queue for a long time and delay processing of new documents. For existing documents, work in small batches.
Privacy: which data leaves the system
For suggestions, chat and embeddings, Paperless-ngx sends document content and metadata to the configured backend (Usage, AI Features). Where the data goes depends solely on the configuration:
| Configuration | Data flow | Assessment |
|---|---|---|
ollama on the same host or internal network |
stays within your network | suitable for confidential documents |
huggingface for embeddings |
local inside the Paperless container; only the first model download needs internet | local |
openai-like with a self-hosted server (vLLM, LiteLLM in front of own inference) |
stays within your network | local, as long as the gateway does not route to cloud models |
openai-like with a cloud provider |
document content goes to the provider | check data processing agreement, third-country transfer and cost |
Remote OCR (PAPERLESS_REMOTE_OCR_ENGINE=azureai) |
the document files go to Microsoft Azure AI Document Intelligence | needs its own assessment, independent of the LLM choice |
Remote OCR is a separate feature from version 3.0.0 and off by default (Usage, Remote OCR). When enabled in always mode, every supported document is sent to Azure. Since 3.1.0, PAPERLESS_REMOTE_OCR_MODE=workflow_only limits it to individual documents via workflows. If you choose local AI for privacy reasons, do not enable Remote OCR in passing.
The documentation treats document content as untrusted data towards the model. A document may contain text phrased like an instruction to the model. For automatic workflows this is an argument for using "create missing items" and "overwrite" sparingly. Background in the article Prompt injection protection.
Native AI vs. Paperless-AI compared
| Aspect | Native AI in Paperless-ngx 3.2.1 | Paperless-AI 3.0.9 |
|---|---|---|
| Maintenance | part of the main project, regular releases | not maintained according to the repository (notice since 31 Mar 2026), last release 4 Nov 2025 |
| Installation | environment variables in the existing container | separate container with API token and its own web UI |
| Backends | ollama, openai-like |
Ollama, OpenAI and OpenAI-compatible providers |
| Suggestions in the Paperless UI | yes, in the suggest menu | no, separate UI (/manual) |
| Automatic processing | workflow action with Paperless triggers and filters | own polling on a schedule (default every 30 minutes) with tag filter |
| Fields | title, tags, correspondent, document type, storage path, date | title, tags, correspondent, document type, date, custom fields |
| Custom prompt | no documented setting | freely editable system prompt |
| Chat | inside the Paperless UI, with links to source documents | separate chat UI with RAG |
| Permissions | chat and suggestions follow the signed-in user's permissions | all documents visible to the API token's user |
| Vector index | in the Paperless data directory, part of the normal backup | separate data store in the Paperless-AI container |
Sources: Paperless-ngx Advanced Usage, Paperless-AI README, Paperless-AI sample configuration, Paperless-AI releases.
Paperless-AI only keeps an advantage where a custom system prompt or filling custom fields is needed. That has to be weighed against a project without security updates that holds an API token with broad access to the archive.
Moving off Paperless-AI
The switch takes six steps and loses no data. Paperless-AI writes its results directly into Paperless-ngx, so tags, correspondents and titles remain in place.
- Back up and upgrade. Back up Paperless-ngx, review the breaking changes of 3.0.0 and upgrade to 3.2.1.
- Record the configuration. From the Paperless-AI configuration, note the model and endpoint (
OLLAMA_API_URL,OLLAMA_MODELorCUSTOM_BASE_URL), the filter tag (TAGS), the processed tag (AI_PROCESSED_TAG_NAME) and any changes toSYSTEM_PROMPT. - Stop Paperless-AI. Stop the container but do not delete it yet, so that two systems do not process the same documents in parallel.
- Enable native AI. Enter the existing Ollama endpoint as
PAPERLESS_AI_LLM_ENDPOINT, set the embedding backend and rundocument_llmindex rebuild. - Rebuild the workflow. Use the former filter tag as the workflow filter, create "Apply AI Suggestions" with the fields used so far and set the processed tag with an assignment action.
- Verify and retire. Compare a sample of new documents. Then remove the Paperless-AI container and its data and revoke its API token in Paperless-ngx.
What does not carry over: a customised system prompt and filling custom fields. If you depend on these, clarify this with the business teams before switching. The existing guides for Paperless-AI and the Paperless-AI installer remain available for existing installations.
Limitations as of October 2026
Some limitations from the 3.0 days still apply in 3.2.1. They can be traced in the source code:
| Point | Status in 3.2.1 | Evidence |
|---|---|---|
| References in chat | at most three referenced documents per answer; retrieval uses five text chunks | MAX_CHAT_REFERENCES = 3, CHAT_RETRIEVER_TOP_K = 5 in chat.py |
| Conversation history | every question is sent on its own; history only lives in the browser and is lost on reload | request contains only document ID and question, chat.service.ts |
| Formatting | answers are shown as plain text without Markdown rendering | chat.component.html |
| Backends | only ollama and openai-like; other providers via an OpenAI-compatible endpoint |
configuration |
| Prompts | no documented setting for custom prompts | configuration |
| Custom fields | not part of AI suggestions | Usage, Workflows |
In practice: suggestions and the workflow action are fit for production use with a review step. The chat works as a search aid for questions like "When does the lease end?", but does not replace research across many documents with a complete list of sources. If you need that, build a dedicated RAG application as described in Processing documents with AI.
Our approach at WZ-IT
WZ-IT operates Paperless-ngx for businesses and sets up the AI integration. When introducing the native AI we work in five steps:
- Take stock. Record version, document volume, languages, existing rules and any running Paperless-AI installation.
- Decide where the model runs. Depending on confidentiality: Ollama on existing hardware, a GPU server or an AI Cube on premises; cloud APIs only after an explicit decision.
- Set up and test. Upgrade to 3.2.1, set the AI variables, build the index and assess suggestions on a sample together with the business teams.
- Build workflows. Automatic assignment with a review tag, existing documents in batches, controlled retirement of Paperless-AI.
- Operate. Updates including index migration, backups of documents and index, monitoring of the task queue.
Managed open source such as Paperless-ngx costs from EUR 129.90 excl. VAT per workload and month at WZ-IT (Managed Open Source). For local models with sensitive documents, the AI Cube is available at EUR 6,490 excl. VAT one-off plus AI Cube Care from EUR 349.90 excl. VAT per month. Support, consulting and implementation by WZ-IT.
Further guides
- Install Paperless-ngx on Ubuntu with Caddy, the base installation with Docker and automatic certificates.
- Install Paperless-AI, the previous extension for existing installations.
- Paperless-AI one-liner installer, installation by script.
- Local VLM OCR models, when text recognition itself needs to improve.
- Self-hosted AI document processing, capturing invoices and delivery notes in structured form.
- Processing documents with AI, from OCR to intelligent document processing.
- Which LLM to self-host?, classifying models for self-hosting.
- vLLM vs. Ollama, which inference server fits which job.
- AI solutions by WZ-IT, the hub for all local AI offerings.
Introducing Paperless-ngx with local AI? We set up the native AI, translate an existing Paperless-AI configuration into workflows and operate the instance long term if required. Book a meeting
Sources
- Paperless-ngx release 3.0.0
- Paperless-ngx release 3.0.5
- Paperless-ngx release 3.1.0
- Paperless-ngx release 3.1.3
- Paperless-ngx release 3.2.0
- Paperless-ngx release 3.2.1
- Paperless-ngx releases
- Paperless-ngx docs, Configuration: AI
- Paperless-ngx docs, Advanced Usage: AI features
- Paperless-ngx docs, Usage: AI Features
- Paperless-ngx docs, Usage: Apply AI Suggestions
- Paperless-ngx docs, Usage: Remote OCR
- Paperless-ngx docs, Administration: LLM index
- Paperless-ngx, init script init-llmindex-migrate (v3.2.1)
- Paperless-ngx wiki, AI Model Recommendations
- Paperless-ngx, docker-compose.postgres.yml (v3.2.1)
- Paperless-ngx source, chat.py (v3.2.1)
- Paperless-ngx source, chat.service.ts (v3.2.1)
- Paperless-ngx source, chat.component.html (v3.2.1)
- Paperless-ngx source, vector_store.py (v3.2.1)
- Paperless-AI repository with maintenance notice
- Paperless-AI releases
- Paperless-AI sample configuration .env.example
- Ollama docs, Docker
- Docker docs, GPU support in Compose
- Ollama library: gemma3
- Ollama library: qwen3
- Ollama library: llama3.1
- Ollama library: bge-m3
- Ollama library: embeddinggemma
Set up AI in Paperless-ngx or have it operated
We enable the native AI in your Paperless-ngx instance, connect a local model, build workflows for automatic tagging and, if required, take over ongoing operation.
Frequently Asked Questions
Answers to important questions about this topic
Since version 3.0.0, released on 22 July 2026. The release includes AI-assisted suggestions, a document chat and a vector index for similar documents and RAG. Version 3.1.0 from 27 August 2026 added the 'Apply AI Suggestions' workflow action. The current version is 3.2.1 from 20 September 2026 (as of October 2026).
No. PAPERLESS_AI_ENABLED defaults to false, and no AI feature works without PAPERLESS_AI_LLM_BACKEND being set. If you configure nothing, Paperless-ngx keeps working without AI after the upgrade to version 3.
No, not on its own. AI suggestions appear in the suggest menu on the document detail page and only apply after a click. Documents are changed automatically only if you create a workflow with the 'Apply AI Suggestions' action. Even then, by default the action only fills empty fields, does not create new tags or correspondents, and only adds tags instead of replacing existing ones.
No. The classic classifier, a locally trained machine learning model without an LLM, stays active. AI suggestions appear in addition. The matching rules for tags, correspondents and document types also keep working unchanged.
Yes. PAPERLESS_AI_LLM_BACKEND accepts 'ollama' and 'openai-like'. For Ollama, PAPERLESS_AI_LLM_ENDPOINT is required, for example http://ollama:11434 on the same Docker network. With 'openai-like' you can also connect vLLM, LiteLLM or other OpenAI-compatible servers.
Only with a local backend. Paperless-ngx sends document content and metadata to the configured model. With Ollama or a self-hosted OpenAI-compatible server this stays within your network. With a cloud API the content leaves the server. The same applies to Remote OCR via Azure AI Document Intelligence, which sends documents to Microsoft.
For most use cases, no. The native AI covers suggestions, automatic assignment via workflows and document chat. According to its repository, Paperless-AI has not been maintained since March 2026, and its last release is v3.0.9 from 4 November 2025. The remaining reasons for Paperless-AI are a freely written system prompt or filling custom fields, which the native integration does not offer.
The Paperless-ngx wiki lists multilingual embedding models such as intfloat/multilingual-e5-small, intfloat/multilingual-e5-base and BAAI/bge-m3. For generation it recommends a small instruction model that reliably produces structured output, tested on your own documents. Multilingual models such as Gemma 3 or Qwen 3 from the Ollama library are a sensible starting point.
Yes. The chat relies on the LLM index, which is only built when AI is enabled and PAPERLESS_AI_LLM_EMBEDDING_BACKEND is set. Suggestions work without the index but are grounded in similar existing documents when it is available. After enabling it for the first time or changing the embedding model, build the index with document_llmindex rebuild.
Managed open source such as Paperless-ngx costs from EUR 129.90 excl. VAT per workload and month at WZ-IT. For local models with sensitive documents there is the AI Cube at EUR 6,490 excl. VAT one-off plus AI Cube Care from EUR 349.90 excl. VAT per month. We clarify the exact scope in a short call.

Written by
Timo Wevelsiep
Co-Founder & CEO
Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.
LinkedInLet's Talk About Your Idea
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.





