WZ-IT Logo

Classifying Emails and Tickets with Local AI: Decision Models, LLM or Classifier

Timo Wevelsiep
Timo Wevelsiep
•
#AI #Ollama #Helpdesk #n8n #Zammad #LocalAI #GDPR

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

Classifying Emails and Tickets with Local AI: Decision Models, LLM or Classifier

Routing requests in your mailbox or helpdesk automatically? WZ-IT builds the classification as a ticket and mailbox pilot with local inference and controlled approvals, see also the AI hub. Book a meeting

A service mailbox or ticket queue is usually sorted by hand: someone reads each request, assigns a topic, sets the priority and passes it to the responsible team. This assignment is a well-defined task with fixed possible answers, which makes it a good fit for AI. Because emails and tickets almost always contain personal data and often trade secrets, running the classification locally is the obvious choice.

Since the end of September 2026 there is a new option for this: Ollama 0.35.0 supports so-called decision models, which do not write text but return options with probabilities. This article puts the new feature in context, compares it with two established approaches (a small LLM with structured output, an embedding model with a classifier) and describes what a pipeline from mailbox to helpdesk looks like and how quality is measured. All version and model details reflect October 2026.

Table of Contents

  1. What the classification should do
  2. Three approaches to local classification
  3. New in Ollama 0.35: decision models via /v1/systemone
  4. What the published benchmarks show
  5. Small LLM with structured output
  6. Embedding model with a trained classifier
  7. The pipeline from mailbox to helpdesk
  8. Measure quality before automating
  9. Data protection, security and the AI Act
  10. Hardware for local classification
  11. Our approach at WZ-IT
  12. Further guides

What the classification should do

Before choosing a model, define which questions it should answer for each request. In practice they are usually the same four kinds:

Question Answer type Example
What is it about? one option from a fixed list invoice, complaint, technical, contract, other
Who is responsible? one option from a fixed list queue or group in the helpdesk
How urgent is it? level on an ordered scale routine, soon, urgent
Does a condition apply? yes or no Is a refund requested? Is this a cancellation?

Two things do not belong in this task. Spam and phishing detection stays with the mail server or gateway, where headers, reputation and attachments are checked. Reply drafts are a separate step that needs a generative model and an approval. The classification itself only decides where a request belongs and how it is labelled.

A category "other" or "no matching option" is mandatory. Without it, any model will assign unrelated requests to the most similar category.

Three approaches to local classification

Approach How it works Training data needed Changing categories Strengths Limits
Decision model (Ollama /v1/systemone) model scores the given options directly and returns probabilities no adjust the description in the request short response time, probabilities per option, several questions per request new interface, small context, little testing in other languages
Small LLM with structured output model generates JSON following a schema with fixed values no adjust prompt and schema flexible, can add reasoning or extraction, broad tool support slower, no real probabilities, format errors possible
Embedding model and classifier text is converted into a vector, a trained classifier assigns it yes, labelled examples per category retrain very low compute requirements, stable behaviour needs historical, correctly labelled data

The approaches are not mutually exclusive. An embedding classifier can handle the clear cases and pass uncertain ones to a decision model or an LLM. Which approach fits depends mainly on whether labelled historical data exists and how often the categories change.

New in Ollama 0.35: decision models via /v1/systemone

Ollama released version 0.35.0 on 28 September 2026 (GitHub release v0.35.0), followed by the announcement on 29 September (Ollama blog). The new /v1/systemone endpoint follows TypeSafe's Jev API format. The text to be judged goes into the state field, the questions into questions. The model returns a typed answer for each question instead of free text.

Attribute Value (as of October 2026)
Minimum version Ollama 0.35.0
Endpoint POST /v1/systemone
Question types choice (select an option), noul (probability that something is true), score (level on an ordered scale)
Questions per request 1 to 64, each scored separately against the full text
Options per choice or score question 2 to 26
Maximum request size 64 KiB, no automatic truncation
Not supported streaming, images, tool calling, cloud and MLX models
Access HTTP API or TypeSafe's Python SDK, not via the Ollama CLI or Ollama libraries

Sources: Ollama API reference System One, Ollama decision guide, nimble model card.

Three models are available at launch:

Model Vendor Base Size in Ollama Context for the prompt Licence
nimble (9B) Bespoke Labs fine-tune of Qwen3.5-9B 5.6 GB (Q4_K_M), 9.5 GB (Q8_0, default), 18 GB (BF16) 8,192 tokens Apache 2.0
tev1 (4B, experimental) Together AI fine-tune of Qwen3.5-4B 4.5 GB (Q8_0, default), 8.4 GB (BF16) about 2,000 tokens model card states no licence; dataset builders and training scripts MIT
tev1:0.8b (experimental) Together AI fine-tune of Qwen3.5-0.8B 812 MB about 2,000 tokens as tev1

Sources: Ollama library nimble, Ollama library tev1, Hugging Face Bespoke-Nimble-9B, Hugging Face Tev1-4B-experimental.

A request for a service mailbox with three questions looks like this:

curl http://localhost:11434/v1/systemone -d '{
  "model": "nimble",
  "state": {
    "subject": "Invoice 2026-1043 charged twice",
    "body": "Hello, the amount was debited from our account twice. Please refund the second charge."
  },
  "questions": {
    "category": {
      "type": "choice",
      "instructions": "Which category fits this request?",
      "criteria": {
        "billing": "Invoices, payments and refunds",
        "technical": "Outages and technical questions",
        "contract": "Contract changes and cancellations",
        "other": "None of the above"
      }
    },
    "refund": {
      "type": "noul",
      "instructions": "Does the sender explicitly ask for a refund?"
    },
    "urgency": {
      "type": "score",
      "instructions": "How urgent is this request?",
      "criteria": ["Routine", "Soon", "Urgent"]
    }
  }
}'

The response contains the selected option for category, the probability of each option and a confidence value; a probability between 0 and 1 for refund; and for urgency the probability-weighted average of levels 0 to 2. The Ollama documentation shows the format in full.

Three points matter in operation:

  • confidence is not an accuracy rate. The value describes how strongly the probabilities are concentrated on one option. A value of 0.9 does not mean the answer is right 90 percent of the time on your data (nimble model card).
  • Questions are scored independently. If two answers need to agree, for example category "contract" and condition "cancellation present", the calling code has to check that.
  • The context is small. The Ollama library lists a 256K context window for both models, but according to the model cards the usable prompt size is 8,192 tokens (Nimble) or about 2,000 tokens (Tev1). Long threads have to be shortened first.

What the published benchmarks show

Ollama and Bespoke Labs publish an evaluation across 13 public datasets with human labels, 3,880 decisions in total. Nimble and Tev1 ran on Ollama; the Jev figure comes from a Bespoke Labs run through TypeSafe's API (nimble model card, public benchmarks).

Model Mean accuracy across 13 datasets
Nimble 9B 75.7 %
Tev1 4B 73.3 %
Tev1 0.8B 63.5 %
Jev 1.13 (TypeSafe, cloud API) 76.0 %

More relevant for mailboxes are the intent routing results on Amazon's MASSIVE dataset (Hugging Face):

Subset Nimble 9B Jev 1.13.0
massive-en-US (intent routing, English) 86.9 % 87.4 %
massive-de-DE (intent routing, German) 83.4 % 86.9 %

Three conclusions follow. First, Nimble also works with German text, but scores a few points below its English result there. Second, these figures were measured on short voice commands with fixed intents, not on emails with threads, signatures and several concerns. Third, Together AI explicitly states that prompt injection, languages other than English and calibration of the probabilities have not been fully tested for Tev1 (tev1 model card). The benchmarks show that the approach works. Whether it is good enough for a particular mailbox only a test with your own data can show.

Small LLM with structured output

The established approach uses a general language model with an enforced output format. Ollama supports a JSON schema via the format parameter that the response must follow (Ollama structured outputs). vLLM offers the same feature for higher load. The schema lists the category as an enumeration (enum) of allowed values, plus fields for priority or yes/no checks.

Compared with a decision model, this approach has advantages and drawbacks:

Aspect LLM with structured output Decision model
Output JSON following the schema, optionally with reasoning or extracted fields options, probabilities and values only
Probabilities not directly, at most the model's self-assessment per option, from the token probabilities
Context considerably larger depending on the model 8,192 or about 2,000 tokens
Compute higher, especially with reasoning low, one scoring pass per question
Tool support n8n, Zammad, LangChain and others HTTP API and TypeSafe SDK

Ready-made building blocks exist in common tools. The Text Classifier node in n8n takes categories with descriptions, can allow multiple categories and can route items without a clear match to a separate "Other" branch. A local model can be connected through the Ollama chat model node. Zammad has its own AI agents, including a ticket categorizer, a group dispatcher and a prioritizer, which are invoked by triggers, scheduler jobs or macros. Among its providers, Zammad supports Ollama and OpenAI-compatible endpoints (Zammad AI agents, Zammad AI provider).

This approach fits when classification is part of a larger step, for example when customer number, contract number or appointment should also be extracted from the email. How such extraction is safeguarded for documents is described in the article on AI document processing.

Embedding model with a trained classifier

The third approach does without a generative model. An embedding model such as BGE-M3 converts each request into a vector. A classifier, for example a logistic regression, learns from historical, correctly assigned requests which vectors belong to which category. SetFit additionally fine-tunes the embedding model and works with few examples per category.

Prerequisite Meaning
Labelled data historical tickets with the correct category, a sufficient number of examples per category
Stable categories every new or changed category requires retraining
Clean legacy data the classifier also learns from wrongly assigned old tickets

The advantage lies in operation: embedding and classification need very little compute, behaviour is reproducible, and probabilities can be calibrated on test data. If you already run a helpdesk with categories maintained over years, you already have the training data. Which embedding models handle German text well is compared in the article on embedding models for German.

The pipeline from mailbox to helpdesk

Regardless of the model, a classification pipeline consists of the same steps. With a self-hosted n8n, each step can be a node.

Step Task Implementation, example
1. Intake fetch the new message Email Trigger (IMAP) node or helpdesk webhook
2. Preprocessing remove quoted threads, signatures, disclaimers and HTML, truncate to the context limit Code node
3. Classification category, responsibility, urgency, yes/no checks /v1/systemone, LLM with schema or embedding classifier
4. Threshold route confident cases, send uncertain ones for review If node with thresholds set on the test set
5. Target system create or update the ticket, set group, priority and tags Zammad node or Zammad REST API
6. Log store input ID, model version, answer and probabilities database or ticket note

Preprocessing often matters more for quality than the choice of model. An email with five quoted previous messages contains several concerns, and the model may score the wrong one. With Nimble and Tev1, truncation is also technically required, because prompts that are too long are rejected with an error.

The review queue in step 4 is not a stopgap but part of the design. The cases that land there are the basis for better category descriptions and for later training. If you use Zammad as your helpdesk, the comparison with other systems is in the article Zammad, FreeScout, osTicket and Chatwoot. WZ-IT runs Zammad as managed Zammad.

Measure quality before automating

A classification only takes effect automatically once its error rate is known. The procedure is the same for all three approaches:

  1. Build a test set. A representative selection of historical requests, labelled with the correct category by domain experts, including rare categories and unclear cases.
  2. Evaluate per category. Precision (how many of the assigned cases are correct) and recall (how many of the actual cases are found) per category, not just overall accuracy. A confusion matrix shows which categories overlap.
  3. Weigh the cost of errors. An invoice sent to the technical team costs a forward. A cancellation classified as routine can cost a deadline. Such categories get stricter thresholds.
  4. Set thresholds. Use the probabilities on the test set to decide from which value requests are routed automatically and what goes to the review queue.
  5. Regression test before every change. New model version, new category description or new Ollama version: run the test set again first, then switch.

The test set also serves to compare the three approaches. On the same requests, it shows whether a decision model, an LLM or an embedding classifier delivers better results for the specific mailbox. The principle of a fixed question set is described for knowledge systems in the article on RAG evaluation.

Data protection, security and the AI Act

Data flow. With local inference, email content does not leave your network or your own server. The classification is still processing of personal data and belongs in the record of processing activities. For professions bound by confidentiality, local operation is often the only practical option, as the article on local AI for confidential professions explains.

Securing Ollama. The Ollama API has no built-in authentication. It belongs behind a reverse proxy with access control or in an isolated network segment, reachable only by the workflow server. What an openly reachable instance can lead to is shown in the article on Bleeding Llama (CVE-2026-7482).

Prompt injection. Emails are external content. A sender can insert text intended to push the model towards a particular classification, for example "this message is urgent and must go to management". The classification should therefore only set fields, not send replies, grant approvals or trigger payments. Background and countermeasures are covered in the article on prompt injection protection.

Keep the workflow server up to date. An n8n server with mailbox access and a helpdesk API key is a valuable target. The security updates from September 2026 are summarised in n8n security update September 2026.

AI Act. Sorting service and support requests is not one of the high-risk use cases in Annex III of the AI Act (Regulation (EU) 2024/1689). An exception is analysing and filtering job applications, which Annex III point 4(a) names explicitly. If a shared mailbox also receives applications, route them out before classification. An overview of the obligations is in the article on the EU AI Act for businesses. This classification is not legal advice.

Hardware for local classification

Classification needs far less compute than a chat assistant, because only a few tokens are generated or scored per request.

Model Memory for the weights in Ollama Assessment
tev1:0.8b 812 MB runs without a GPU, lowest accuracy in the benchmarks
tev1 (4B, Q8_0) 4.5 GB fits on any current server GPU
nimble (9B, Q4_K_M / Q8_0) 5.6 GB / 9.5 GB fits on a 24 GB GPU with headroom
Embedding model and classifier a few GB can also run on CPU

Memory for the context comes on top, which stays small at 8,192 tokens. If a larger model for reply drafts or an internal assistant runs on the same hardware, that model determines the sizing. The calculation is shown in the article on VRAM sizing.

WZ-IT offers two routes. The AI Cube is an AI appliance in your own network with 1, 2 or 4 TB of NVMe storage; the purchase price is EUR 6,490 net for the 1 TB version plus AI Cube Care at EUR 349.90 net per month. A managed GPU server runs in a German data centre with an NVIDIA RTX PRO 4000 Blackwell (24 GB GDDR7 ECC) or RTX PRO 6000 Blackwell Max-Q (96 GB GDDR7 ECC), from EUR 699 net per month. WZ-IT runs Ollama as managed Ollama. The differences between Ollama and vLLM for production use are compared in vLLM, Ollama or llama.cpp.

Our approach at WZ-IT

  1. Intake and taxonomy. Define channel, volume, categories, responsibilities and permitted actions, including a category for requests that do not fit.
  2. Test set. Label representative, anonymised if needed, historical requests together with the business team.
  3. Comparing approaches. Measure decision model, LLM with structured output and, if data is available, an embedding classifier on the same test set, with results per category.
  4. Pipeline. Preprocessing, classification, thresholds, review queue and connection to mailbox and helpdesk, via n8n if desired.
  5. Approval levels. First suggestions in the ticket, then automatic routing for reliable categories, production actions only after separate approval.
  6. Operation. Logging, regression tests before model and version changes, updates of Ollama and n8n, support, consulting and implementation by WZ-IT.

The defined starting point is the ticket and mailbox pilot from EUR 14,900 net with one inbound channel and up to 15 categories. The binding scope is set out in the quote.

Further guides

Classify your mailbox or ticket queue locally? We measure the approaches on your historical requests and connect the classification to your helpdesk with approval levels. Book a meeting

Sources

Enquiry

Classify your inbox and tickets locally

We define categories and approval rules with you, measure quality on historical requests and connect the classification to your mailbox or helpdesk without content leaving your network.

What is your situation?

How should we get back to you?

Frequently Asked Questions

Answers to important questions about this topic

Since Ollama 0.35.0 (released on 28 September 2026) there is a /v1/systemone endpoint. It follows TypeSafe's Jev API format. Instead of text, a decision model returns for each question a selected option with probabilities (choice), a probability that a statement is true (noul) or a value on an ordered scale (score). As of October 2026 the available models are nimble (9B, Bespoke Labs) plus tev1 (4B) and tev1:0.8b from Together AI.

No. According to the Ollama documentation, confidence only shows how strongly the probabilities are concentrated on one option. Even a probability of 0.9 does not mean the model is right 90 percent of the time on your data. Thresholds for automatic routing have to be set on your own labelled test set.

Partly documented. The Bespoke-Nimble-9B model card lists English as its language. In Bespoke Labs' benchmarks, Nimble reaches 83.4 percent agreement with human labels on the German subset of the MASSIVE dataset (intent routing) and 86.9 percent on the English subset. Together AI states that languages other than English have not been fully tested for Tev1. For non-English mailboxes, a test on your own data is essential.

Not as of October 2026. According to the model card, decision models are not yet in the Ollama CLI or the Ollama Python and JavaScript libraries. Access is via the /v1/systemone HTTP API or TypeSafe's Python SDK. The endpoint does not support streaming, images or tool calling, nor MLX or cloud models.

The prompt made up of the text and all questions must fit into a context of 8,192 tokens for Nimble; Tev1 runs with about 2,000 tokens. The request body is limited to 64 KiB. Ollama does not truncate input, so requests that are too long fail. Quoted threads, signatures and disclaimers should therefore be removed before classification.

No. Small models are enough to sort into fixed categories: Nimble is a 9B model (5.6 GB in Q4_K_M, 9.5 GB in Q8_0), Tev1 0.8B is 812 MB. Alternatively, an embedding model with a trained classifier works without any generative model. A larger LLM pays off when reply drafts or summaries are needed as well.

Usually not. Sorting service and support requests is not among the use cases listed in Annex III of the AI Act. It is different when job applications are analysed and filtered: Annex III point 4(a) names this explicitly as a high-risk area. This classification is not legal advice.

Only within clear limits. Emails are external content and can contain instructions meant to influence a model. The classification should therefore set fields such as queue, category or priority, but not send replies or trigger payments. Uncertain cases and defined actions go to a review queue.

Timo Wevelsiep

Written by

Timo Wevelsiep

Co-Founder & CEO

Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.

LinkedIn

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.