WZ-IT Logo

Local, GDPR-Compliant Speech Recognition: Whisper, Voxtral and Parakeet Compared

Timo Wevelsiep
Timo Wevelsiep
•
#AI #SpeechToText #Whisper #Voxtral #Parakeet #GDPR #Transcription

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

Local, GDPR-Compliant Speech Recognition: Whisper, Voxtral and Parakeet Compared

Transcribing dictation, meetings or calls in your own network? WZ-IT runs speech-to-text models such as Whisper or Voxtral on the AI Cube in your company network or on a managed GPU server, see the AI hub. Book a meeting

Speech recognition is one of the AI applications with the clearest benefit: dictation becomes text, meetings become minutes, phone calls become searchable notes. At the same time, audio is particularly sensitive. A recording contains the voice itself in addition to the content, often names, health details or client matters, and it is frequently created in situations where the people involved do not expect it to be passed on.

Cloud transcription services process this data at the provider. Open models such as OpenAI Whisper, Mistral Voxtral and NVIDIA Parakeet, by contrast, run entirely on your own hardware. This article compares the models based on their model cards (license, size, German support), shows the common runtimes, speaker separation with pyannote and the legal framework of the GDPR, German criminal law and the EU AI Act. All model details reflect September 2026.

Table of Contents

  1. What local speech recognition means
  2. The models at a glance: Whisper, Voxtral, Parakeet
  3. German: what the vendor figures say
  4. Runtimes and interfaces
  5. Speaker separation with pyannote
  6. Hallucinations and quality assurance
  7. Hardware: what fits on 24 GB, 96 GB and the AI Cube
  8. Legal framework: GDPR, German Criminal Code and AI Act
  9. Use cases: meetings, dictation, phone
  10. Our approach at WZ-IT
  11. Further guides

What local speech recognition means

Local means: the model and the audio files sit on a server the company controls itself, in its own network or on a dedicated server in a data centre. Neither audio nor transcript goes to a model provider. A typical chain consists of four steps:

Step Task Examples
Capture and preprocessing Convert audio to a uniform format (usually 16 kHz mono), detect silence ffmpeg, voice activity detection (Silero VAD)
Recognition (ASR) Convert speech to text, with timestamps Whisper, Voxtral, Parakeet
Speaker separation Assign segments to speakers pyannote
Further processing Summary, minutes, hand-over to business software local language model, n8n, interface

Recognition itself is the most compute-intensive part, but not the only one that touches data. Summarisation by a language model should also run locally, otherwise the transcript leaves the building in the last step after all.

The models at a glance: Whisper, Voxtral, Parakeet

Model Vendor Parameters License Languages Released
Whisper large-v3 OpenAI 1.55 bn MIT (repository), model card states Apache 2.0 99 6 Nov 2023
Whisper large-v3-turbo OpenAI 809 m MIT 99 (transcription only, no translation) 1 Oct 2024
Voxtral Mini 3B (2507) Mistral AI 3 bn according to Mistral Apache 2.0 8, including German 15 Jul 2025
Voxtral Small 24B (2507) Mistral AI 24 bn Apache 2.0 8, including German 15 Jul 2025
Voxtral Mini 4B Realtime (2602) Mistral AI 4 bn Apache 2.0 13, including German February 2026
Parakeet TDT 0.6B v3 NVIDIA 600 m CC BY 4.0 25 European, including German 14 Aug 2025
Canary-1B-v2 NVIDIA 978 m CC BY 4.0 25 European, including German 14 Aug 2025

Sources: model cards of Whisper large-v3, Whisper large-v3-turbo, Voxtral Mini 3B, Voxtral Small 24B, Voxtral Mini 4B Realtime, Parakeet TDT 0.6B v3 and Canary-1B-v2; Whisper release dates from the announcements of large-v3 and turbo.

Whisper is the benchmark other models are measured against. The Whisper repository releases code and weights under the MIT license. Whisper processes audio in 30-second windows; for longer recordings the runtime stitches the segments together. large-v3-turbo reduces the decoder layers from 32 to 4, which makes it considerably faster and, according to OpenAI, comes with a minor quality degradation. turbo is not trained for translation into English.

Voxtral is a language model with audio input. Besides plain transcription it answers questions about the spoken content and summarises it without a separate language model. According to Mistral, Voxtral handles up to 30 minutes of audio for transcription and up to 40 minutes for understanding tasks with a 32,000-token context. Voxtral Mini 4B Realtime is designed for live transcription, with a configurable delay between 80 ms and 2.4 s.

Parakeet is a pure recognition model from the NVIDIA NeMo framework. It delivers punctuation, capitalisation and word- and segment-level timestamps. According to the model card, it processes up to 24 minutes of audio in one pass with full attention and up to three hours with local attention. Canary-1B-v2 from the same line additionally translates between English and the 24 other languages.

On licensing: MIT and Apache 2.0 permit commercial use with no conditions beyond the license notice. CC BY 4.0 also permits it but requires attribution. For an internal tool this is unproblematic; in a product for third parties the notice belongs in the documentation.

Not all Voxtral models are open. Voxtral Mini Transcribe V2 with built-in speaker diarization, announced in February 2026, is offered by Mistral only through its own API. From this generation, open weights exist only for Voxtral Realtime.

German: what the vendor figures say

For decision-makers, the word error rate (WER) in German is the obvious metric. The available figures, however, come from the vendors themselves:

Model Dataset WER German Source
Parakeet TDT 0.6B v3 FLEURS 5.04% Model card
Parakeet TDT 0.6B v3 CoVoST 4.84% Model card
Voxtral Mini 4B Realtime FLEURS, 480 ms delay 6.19% Model card
Whisper large-v3 no separate German figure in the model card - Model card

These figures are not directly comparable. Normalisation of punctuation and numbers, decoding parameters and test setup differ, and a realtime model with 480 ms delay works under different conditions than a model that sees the entire file. FLEURS and CoVoST consist of read sentences in good recording quality. Telephone audio at 8 kHz, room reverberation in meetings, dialect and specialist vocabulary (drug names, case numbers, product names) degrade the results of every model.

Mistral states that Voxtral outperforms Whisper large-v3 in its own tests (Mistral). That, too, is a vendor statement. What counts for the selection is a test with 30 to 60 minutes of your own recordings that have been transcribed correctly by hand. The WER can be calculated from them with standard tools such as jiwer, and specialist terms belong in a separate checklist.

Runtimes and interfaces

A model on its own is not yet a service. Between model and application sits a runtime that accepts audio, runs the model and returns text with timestamps.

Tool License Purpose
faster-whisper MIT Whisper based on CTranslate2, according to the project up to four times faster than the reference implementation at the same accuracy, int8 on GPU and CPU
WhisperX BSD-2-Clause Whisper with voice activity detection, word timestamps via wav2vec2 alignment and speaker separation via pyannote
Speaches MIT OpenAI-compatible server for transcription (faster-whisper) and speech generation, with streaming
vLLM Apache 2.0 Transcriptions API at /v1/audio/transcriptions for ASR models; Mistral recommends vLLM for Voxtral
NVIDIA NeMo Apache 2.0 Framework for Parakeet and Canary

For integration, the interface matters more than the model. Speaches and vLLM offer an endpoint in the format of the OpenAI audio API. Applications already built for a cloud transcription service can be switched to your own server by changing the base URL and model name. WZ-IT runs Speaches as a managed service and vLLM for language and audio models. How vLLM differs from Ollama and llama.cpp is covered in vLLM, Ollama or llama.cpp.

If you use Nextcloud, transcription can also be connected through the Nextcloud Assistant; the changes in version 35 are described in the article on Nextcloud 35.

Speaker separation with pyannote

Meeting minutes need to show who said what. Whisper, Voxtral 2507 and Parakeet do not provide this information. It is produced in a second step, diarization.

The common open solution is pyannote. The current open model speaker-diarization-community-1 is licensed under CC BY 4.0, runs with pyannote.audio 4.x and can be operated fully offline once downloaded. The download requires registering on Hugging Face and sharing contact details with the developers. The stronger precision-2 model, by contrast, runs only through the pyannoteAI API and is therefore not suitable for local operation.

Two points belong in the planning:

  • Placeholders instead of names. Diarization assigns labels such as SPEAKER_00. Mapping them to people happens manually or by matching against stored voice profiles.
  • Voice profiles are sensitive. Data that allows the unique identification of a person based on behavioural or physiological characteristics is biometric data under Art. 4(14) GDPR. If it is processed for identification, Art. 9 GDPR applies. For most minutes, manual assignment is sufficient.

Overlapping speech remains a weakness of all methods. Mistral notes for its own API model that with overlapping speech the model typically transcribes one speaker (Mistral).

Hallucinations and quality assurance

Speech recognition with generative models can produce text that was never spoken. The Whisper large-v3 model card states this limitation explicitly. A study by Koenecke et al. (Careless Whisper, ACM FAccT 2024) found entire hallucinated phrases or sentences in roughly 1% of transcriptions; 38% of these hallucinations contained explicit harms such as inaccurate associations.

Measures that reduce the risk:

Measure Effect
Voice activity detection before recognition Silent segments are not passed to the model; WhisperX lists this as a means against hallucinations
Set the language explicitly prevents misdetection of the language in short segments
Keep timestamps in the transcript every passage can be checked against the recording
Review before filing mandatory for medical letters, legal briefs and minutes with legal effect
Test specialist vocabulary your own checklist with names, drugs, case numbers

A transcript is a draft. For documents that decisions are based on, a human review step belongs in the workflow, and the audio should remain available until the review is complete.

Hardware: what fits on 24 GB, 96 GB and the AI Cube

Speech recognition needs far less GPU memory than large language models. Vendor figures:

Model Memory requirement according to vendor
Whisper large-v3 about 10 GB VRAM (OpenAI)
Whisper large-v3-turbo about 6 GB VRAM (OpenAI)
Voxtral Mini 3B about 9.5 GB in bf16 or fp16 (model card)
Voxtral Mini 4B Realtime GPU with at least 16 GB (model card)
Voxtral Small 24B about 55 GB in bf16 or fp16 (model card)
Parakeet TDT 0.6B v3 at least 2 GB RAM to load (model card)

For WZ-IT's offerings this means:

Platform Memory Suitable for
Managed GPU Server 24 (NVIDIA RTX PRO 4000 Blackwell) 24 GB GDDR7 ECC Whisper, Parakeet, Voxtral Mini and Realtime; a smaller language model for summaries in parallel, depending on load
Managed GPU Server 96 (NVIDIA RTX PRO 6000 Blackwell Max-Q) 96 GB GDDR7 ECC additionally Voxtral Small 24B in bf16, or speech recognition alongside a larger language model
AI Cube (NVIDIA GB10) 128 GB unified memory speech recognition and language model on one device in your own network

The Managed GPU Server 24 costs €699 net per month plus €499 one-time setup, the Managed GPU Server 96 €1,799 net per month plus €999 setup. The AI Cube costs €6,490 net one-time, AI Cube Care for ongoing operations €349.90 net per month. How many concurrent users and tasks a system can handle is described in the article on sizing a local AI server. The calculation for language models is shown in GPU VRAM sizing.

Throughput depends on model, runtime, batch size and audio quality. Reliable figures come only from measuring on the target hardware with realistic recordings. For the AI Cube with its Arm processor, it is also necessary to check whether the chosen runtime is available for the platform.

Local operation solves the question of the data path, not every data protection obligation. The following overview is a classification, not legal advice.

Rule Content Relevance for transcription
Art. 6 GDPR legal basis for any processing applies to recording, transcript and summary
Art. 9 GDPR special categories such as health data and biometric data for identification medical dictation, therapy sessions, voice profiles
Art. 28 GDPR processing on behalf required for cloud services, also for operation by an IT service provider
Art. 35 GDPR data protection impact assessment to be checked for extensive processing of sensitive data or systematic monitoring
Section 203 StGB protection of professional secrets disclosure to contributing service providers only where necessary and with a confidentiality obligation
Section 201 StGB confidentiality of the spoken word recording privately spoken words without authorisation is a criminal offence
Art. 5(1)(f) AI Act ban on emotion recognition in the workplace mood analysis of call centre employees is prohibited

For professionals bound by secrecy, Section 203(3) StGB permits disclosure to other persons contributing to their work insofar as this is necessary for that work; subsection (4) makes the professional responsible for ensuring those persons have been bound to confidentiality. With local processing on your own hardware, there is no disclosure to a model provider. Maintenance by an IT service provider still falls under these rules and needs a corresponding agreement.

For recordings in an employment context, the works council's co-determination right under Section 87(1) no. 6 of the Works Constitution Act applies if the recording is suitable for monitoring performance or behaviour. How an AI assistant can be agreed with the works council is described in AI assistant and the works council. The other obligations under the AI Act are summarised in EU AI Act for companies.

Three decisions have proven useful in practice: a retention period for the audio (for example until the transcript is approved), separate access rights for audio and text, and logging of who transcribed and accessed which recording.

Use cases: meetings, dictation, phone

Application Requirement Suitable setup
Meeting minutes multiple speakers, summary, action items Whisper or Parakeet, pyannote, local language model for the minutes
Dictation in a law firm or practice single speaker, specialist vocabulary, professional secrecy Whisper large-v3 or Voxtral Mini, review step, hand-over to business software
Live captions and notes low latency Voxtral Mini 4B Realtime or streaming via Speaches
Call centre and phone 8 kHz audio, many hours per day, co-determination batch processing, test with real telephone quality, no emotion recognition for employees
Knowledge capture from interviews long recordings, reuse in a knowledge base transcription with timestamps, then preparation for RAG

For law firms and medical practices, the articles Local AI for law firms and Local AI for medical practices describe the wider framework. Preparing conversations with departing employees is part of the knowledge transfer offering, which includes transcription where needed.

Our approach at WZ-IT

  1. Assessment. Audio sources, volume per day, languages and specialist vocabulary, protection needs and further processing. This determines whether batch transcription, live operation or both are needed.
  2. Model test with your own recordings. A selection of Whisper, Voxtral and Parakeet against hand-checked reference transcripts, with word error rate and a checklist for specialist terms.
  3. Platform. Operation on the AI Cube in your own network or on a managed GPU server in Germany, with an OpenAI-compatible interface via Speaches or vLLM.
  4. Integration. Speaker separation, summarisation by a local language model, hand-over to business software or filing, retention periods and access rights according to your data protection concept.
  5. Operations. Updates, monitoring and regression testing before model changes. Support, consulting and implementation by WZ-IT.

Audio processing is a separate integration and is implemented after a professional and technical review. The framework for ongoing operations is described in Managed AI.

Further guides

Transcription without a cloud service? We test Whisper, Voxtral and Parakeet with your recordings and run speech recognition on the AI Cube or a managed GPU server. Book a meeting

Sources

Enquiry

Run speech recognition in your own network

We assess your audio sources and protection needs, choose model and runtime, test with your own recordings and run transcription on the AI Cube or a managed GPU server.

What is your situation?

How should we get back to you?

Frequently Asked Questions

Answers to important questions about this topic

As of September 2026, three families are widely used: OpenAI Whisper (large-v3 and large-v3-turbo, MIT license according to the repository), Mistral Voxtral (Mini 3B, Small 24B and Mini 4B Realtime, all Apache 2.0) and NVIDIA Parakeet TDT 0.6B v3 plus Canary-1B-v2 (CC BY 4.0, attribution required). All of them support German and allow commercial use.

No. Local operation avoids transferring data to third parties and thus a data processing agreement with a cloud provider. A legal basis, informing data subjects, retention periods for audio and transcript, access rights and, where required, a data protection impact assessment remain mandatory. Health data is additionally subject to Art. 9 GDPR.

Yes, it happens. The Whisper large-v3 model card states that transcripts may include text that was not actually spoken in the audio. A study (Koenecke et al., ACM FAccT 2024) found entire hallucinated phrases or sentences in roughly 1% of the transcriptions examined. Voice activity detection before recognition reduces the risk but does not replace human review of medical letters or legal briefs.

No. Diarization with pyannote separates speakers and assigns placeholders such as SPEAKER_00 and SPEAKER_01. Who is behind a placeholder is assigned by a person or by an additional step. If that step matches against a stored voice profile, the data may be biometric data under Art. 4(14) GDPR, to which Art. 9 GDPR applies.

There is no general ranking. Vendors publish their own figures, for example a word error rate of 5.04% on FLEURS German for Parakeet TDT 0.6B v3 and 6.19% for Voxtral Mini 4B Realtime at 480 ms delay. Test conditions differ, and technical terms, dialects or telephone quality change results considerably. Only a test with your own recordings is reliable.

No. According to OpenAI, Whisper large-v3-turbo needs about 6 GB of GPU memory, Parakeet TDT 0.6B v3 is even smaller at 600 million parameters, and faster-whisper also runs on CPU with int8. A 24 GB GPU is enough for all models except Voxtral Small 24B, which according to Mistral needs about 55 GB of GPU memory in bf16.

No. In Germany, recording another person's privately spoken words without authorisation is a criminal offence under Section 201 of the German Criminal Code (StGB). Recording usually requires the consent of those involved, plus a legal basis under the GDPR. In an employment context, the works council must also be involved if recordings are suitable for monitoring performance or behaviour.

No, not for employees. Since 2 February 2025, Art. 5(1)(f) of the EU AI Act prohibits AI systems that infer emotions of natural persons in the workplace, except for medical or safety reasons. Plain transcription and topic analysis are not covered by this prohibition.

No. Voxtral Mini 3B, Voxtral Small 24B and Voxtral Mini 4B Realtime are available under Apache 2.0 on Hugging Face. Voxtral Mini Transcribe V2 with built-in speaker diarization, announced in February 2026, is offered by Mistral only through its own API. It is therefore not an option for local operation.

Timo Wevelsiep

Written by

Timo Wevelsiep

Co-Founder & CEO

Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.

LinkedIn

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.