Managed GPU Server 20
A compact starting point for a dedicated, fully managed model and inference environment.
€499 excl. VAT one-time setup
- NVIDIA RTX 4000 SFF Ada
- 20 GB GDDR6 ECC GPU memory
- 6,144 CUDA cores
- 70 W maximum GPU power consumption
Dedicated NVIDIA GPU servers in a German data centre including an agreed model, managed vLLM inference, Open WebUI, proactive 24/7 monitoring and updates.
Companies worldwide trust WZ-IT
The following are trademarks of their respective owners: NVIDIA (NVIDIA Corporation). WZ-IT is an independent service provider and has no business, partnership, or contractual relationship with these companies. We offer independent migration, installation, hosting, and operations services.
Both configurations include the dedicated GPU server, one agreed model, vLLM, Open WebUI and managed operations by WZ-IT.
A compact starting point for a dedicated, fully managed model and inference environment.
€499 excl. VAT one-time setup
A managed model and inference environment for larger models, longer contexts and higher concurrency.
€999 excl. VAT one-time setup
WZ-IT does not only provide compute capacity. We keep the agreed technical foundation traceable and supportable throughout ongoing operations.
Linux, administrative access, baseline firewall protection, NVIDIA drivers, vLLM, Open WebUI and one agreed model are configured.
Host, GPU health, capacity and agreed services are monitored proactively. Relevant events enter the operating process.
The operating system, drivers, vLLM and Open WebUI are checked against security advisories and updated in a controlled manner. Critical application vulnerabilities are prioritised.
Incidents are accepted, assessed and handled with documentation according to the agreed service level.
Configuration, responsibilities and relevant changes remain traceable. Your point of contact knows the agreed environment.
The server is located in Germany. External APIs, telemetry, integrations and backup targets are treated and agreed as separate data paths.
The one-time setup creates a documented and monitored model environment. Custom integrations are scoped separately.
Check model, quantisation, context, concurrency and additional compute loads against 20 or 96 GB of GPU memory.
Configure the host system, Linux, access, network, NVIDIA drivers, vLLM, Open WebUI and the agreed model.
Set up monitoring, alert paths, maintenance windows, service level and operations documentation.
Jointly verify the GPU, model endpoint, Open WebUI, reachability and monitoring, then hand over the agreed state.
Additional models, RAG, SSO, custom integrations and data pipelines, extensive backup, multi-GPU and high availability can be added to match the project.
Parameter count alone is not enough. Quantisation, context length, concurrent users, batch size, KV cache and runtime determine memory and performance requirements.
Suitable for compact inference models, embeddings, image and video processing, development environments and clearly bounded production APIs. We verify the specific model fit before provisioning.
More GPU memory provides headroom for larger models, longer contexts, higher concurrency, multimodal pipelines and selected fine-tuning methods.
Multi-GPU, separate inference and training systems, high availability and clusters are designed technically rather than forced into a standard price.
Both configurations are provisioned as usable, managed model environments. Model selection remains dependent on GPU memory and the actual workload.

One agreed model is delivered through a managed vLLM inference layer. API, runtime, updates and monitoring are included.
A managed chat and user interface for the deployed model. Baseline configuration, operations, updates and monitoring are included.
A fixed managed GPU server is not the most economical form for every workload. Utilisation and required operational responsibility determine the right model.
Hourly GPU cloud
Managed dedicated GPU server
Each page has a distinct role, keeping hardware, model serving and ongoing platform operations out of one ambiguous offer.
When you need an operated model endpoint, an OpenAI-compatible API or a complete inference layer.
View serviceWhen an existing or individually designed AI platform with models, runtimes and applications should be taken into operations.
View serviceWhen the AI hardware should remain at your site and operate as a local appliance.
View serviceNon-binding enquiry
Tell us the model, quantisation, context, concurrent usage and required application layer. We assess whether 20 GB, 96 GB or a custom architecture fits.
Pricing, scope, GPU selection, access and extensions
Managed GPU Server 20 costs €599.90 net per month plus a one-time €499 net setup fee. Managed GPU Server 96 costs €1,599.90 net per month plus a one-time €999 net setup fee. Host system, RAM, storage, network, term and provisioning are confirmed in the proposal.
Yes. The standard configurations are based on a dedicated server with one dedicated NVIDIA GPU. Multi-GPU and cluster configurations are designed separately.
The baseline includes the agreed server and GPU infrastructure, Linux base system, NVIDIA drivers, one agreed model, vLLM, Open WebUI, secure administration, proactive 24/7 monitoring, updates and CVE assessment of managed components, incident handling according to the service level, operations documentation and a personal point of contact.
That does not depend on parameter count alone. Quantisation, context length, concurrent requests, batch size, KV cache, runtime and additional workloads affect GPU memory. We therefore confirm model fit before provisioning.
Yes. Your own models, containers and applications can run within the agreed technical framework. Administrative rights and change paths are aligned so managed responsibility remains traceable.
Yes. One model matched to the configuration, the managed vLLM inference layer and Open WebUI are included in both packages. Additional models, RAG, SSO, data migration and custom integrations are assessed separately.
Yes. Multi-GPU, separate systems for inference and training, redundant model endpoints and clusters are possible. These architectures require individual sizing and a separate proposal.
The offered server is located in a German data centre. Whether the entire data path remains in Germany also depends on external APIs, telemetry, model sources, integrations and backup targets. These paths are reviewed and documented within the agreed scope.
07.09.2026
Since 1 January 2025 every German company has to be able to receive e-invoices. Part of the incoming documents now arrive as structured XML and...
24.05.2026
Anyone bringing AI into production business processes quickly faces an uncomfortable question: what is actually happening in there? Which prompt went to which model, why...
10.05.2026
Three unauthenticated API calls. No login, no exploit framework, no privilege escalation. Three POST requests to a default port, and the machine's memory is on...
05.05.2026
The EU AI Act's high-risk obligations do not start on 2 August 2026. Stand-alone high-risk systems under Annex III now have until 2 December 2027,...
24.11.2025
OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a...
From local AI integration to architecture, data sovereignty and ongoing operations.
“WZ-IT moved our studio infrastructure from decentralised individual devices to a central platform: every site is securely connected via VPN, new devices are onboarded automatically and an entire site is provisioned from a template, without manual steps on location. What impressed me most is the breadth and depth of their knowledge: Timo and Robin are not a typical IT provider who sets up a server and leaves. The two of them think their way into highly complex infrastructure and software topics, work through every requirement we put in front of them, and build networking, provisioning and operations so that everything fits together in the end. WZ-IT is an excellent partner for complex software, network and architecture projects.”

Steve Kirchner
Managing Director, nextGYM GmbH

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.