Inference runtime
Deploy SGLang, drivers, kernels and parameters reproducibly.
WZ-IT designs SGLang as an inference layer for local AI, benchmarks models and sizes GPUs, concurrency, context and availability for the real workload.
Companies worldwide trust WZ-IT
The following are trademarks of their respective owners: SGLang (the SGLang project contributors). WZ-IT is an independent service provider and has no business, partnership, or contractual relationship with these companies. We offer independent migration, installation, hosting, and operations services.
Response time and throughput depend on model, quantisation, context length, batch size, concurrency and GPU architecture. A single token figure does not describe business operations.
We test target models with representative inputs and design routing, model cache, metrics, API protection and scaling as part of the platform.
User interface, RAG, identities, governance and business evaluation are supplied by other components. SGLang primarily exposes models through APIs.
GPU type, instance count and target metrics are determined during assessment, so there is no flat standard price.
Model, quantisation, runtime, GPU, context and scheduling form one system that must be benchmarked under realistic load.
Deploy SGLang, drivers, kernels and parameters reproducibly.
Version weights, quantisation, context and APIs.
Optimise KV cache, batching, concurrency and memory limits.
Monitor latency, throughput, failures and GPU use.
WZ-IT operates the runtime; model quality, licensing and permitted use are shared or customer-owned.
| Area | Responsibility | Scope and boundaries |
|---|---|---|
| SGLang runtime | WZ-IT | Deployment, drivers, containers, monitoring and updates. |
| Model artefacts | WZ-IT | Controlled download, storage and launch configuration. |
| Load profile | Shared | We measure while the customer supplies representative prompts. |
| API integration | Shared | Authentication, limits and clients are tested together. |
| Model approval | Customer | Suitability, licences and outputs remain customer-owned. |
| Routing and scaling | Optional WZ-IT service | Multiple models, GPUs and gateways are scoped separately. |
Serve supported language and multimodal models with optimised scheduling and caching.
Benchmark latency, time to first token, token throughput and concurrency with realistic load profiles.
Expose OpenAI-compatible endpoints to approved applications and internal platforms.
Document versions, quantisations, launch parameters and rollout paths transparently.
Protect API authentication, network paths, tenant boundaries and administrative access.
Connect Open WebUI, LiteLLM, agents, RAG services and custom applications to the inference layer.
Assess existing data and configuration, migrate them in a test run and move to managed operations through a controlled cutover.
Secure SSO, roles, administrative paths and external access for the application and existing infrastructure.
Back up all stateful components consistently and document the recovery path for the agreed scope.
Monitor and update the application and its technical dependencies and operate them under the agreed service level.
A clearly defined operating scope instead of an opaque hosting flat fee.
We set up a test instance for you, usually on the next business day. No payment details required. After seven days it is deleted unless you continue.
We combine the right compute size with ongoing operations, backups, monitoring and a service level appropriate for the criticality of SGLang. High availability and recovery targets are designed separately where needed.
We also design custom hosting architectures, integrations and migrations around SGLang. Contact us for a technical assessment.
One managed standard SGLang application is included in the Starter workload. Business and higher levels add a flexible operations allowance for planned work during regular service hours. Select compute, additional applications, storage and the appropriate service level.
A workload is one compute instance with the applications agreed for it.
One standard app per workload is already included. Additional dedicated servers count as separate workloads.
€7.99 excl. VAT per additional 100 GB per month, including daily encrypted offsite backup with 7-day retention. Billing is based on provisioned capacity, not actual usage.
We also implement custom backup schedules and retention periods. For example: daily backups retained for 30 days and one monthly backup retained for 12 months. Scope, storage requirements and additional costs are agreed separately.
+€299 excluding VAT per workload and month
Standard hosting runs on a single instance without high availability. Availability for custom architectures is agreed separately.
Zero downtime on request
Distributed operation with several instances, a database cluster and rolling updates, so that a node failure causes no interruption. Only for applications suited to it; design and price after technical assessment.
This page covers installation and the agreed platform operations for SGLang. Custom code, the access layer and special confidentiality requirements remain clearly separated responsibilities that can be added when needed.
Maintain the application
Keep custom extensions, interfaces and dependencies controlled, updated and supportable.
from €699.90 excl. VAT / month
View serviceSecure access
Add identity-based access and private network paths as a separate operational layer.
from €349.90 excl. VAT / month
View serviceOperate confidentially
For professional secrecy holders, define contractual, access and infrastructure requirements in a dedicated operating model.
from €499.90 excl. VAT / month
View serviceThree relevant adjacent paths instead of a long list of further products.
All prices are net and exclude statutory VAT. The offers are addressed to businesses.
View the overall systemEnquiry
Briefly describe the current state and objective for SGLang. We assess infrastructure, integration, and ongoing operations.
Weights, KV cache, context, batch size, GPU type and TTFT targets must be considered together.
| Usage scenario | Technical starting point | Key factors |
|---|---|---|
| One model and few chats | Model and GPU benchmark | Quantisation, context and interactivity are measured. |
| Several teams or API clients | Concurrency test | Throughput, queueing and response time are measured. |
| Long context or agents | Memory and prefill assessment | KV cache and input load determine capacity. |
| Multiple GPUs or models | Cluster and routing design | Parallelism and failover are planned. |
Pricing follows model, hardware, context, concurrency and target metrics.
The runtime runs on dedicated GPU hardware close to calling applications.
European inference service after benchmarking.
Deployment in your account.
Operation on suitable local GPU servers.
Combine local and external inference endpoints.
Clients use an authenticated, measured API path rather than direct GPU access.
OpenAI-compatible or project-specific clients.
TLS, authentication, quotas and request limits.
Separate keys, models and usage boundaries.
Scheduling, batching, cache and model serving.
Accelerators and kernels for the selected model.
Versioned weights and quantisations.
TTFT, token rate, queue, memory and failures.
Benchmark results only apply to the tested model, runtime, hardware, context and load profile.
Answers about models, GPUs, benchmarks and operations.
It is relevant for performance-focused serving, especially where caching and concurrency matter.
Only after measuring context, input, output and concurrency against target latency.
No. GPU type and quantity are quoted after assessment.
Yes. We deploy it on suitable local GPU hardware.
Yes. We benchmark both with identical models, prompts and concurrency.
As an alternative to managed hosting in the data centre, WZ-IT provides the hardware, configures SGLang, and handles hardening, monitoring, updates, backup and technical support. Access can be limited to the internal network or enabled through VPN and existing identities.
from EUR 349 excl. VAT / month · plus one-time provisioning and initial setup

01.06.2026
Your own ChatGPT, running entirely in-house: a familiar chat interface, a large language model on your own hardware behind it, without a single prompt going...
24.11.2025
OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a...
09.11.2025
Local AI inference means running a language or multimodal model on owned hardware rather than through a public API. Organisations gain control over data paths,...
These solutions are often used together with SGLang
These solutions offer similar functionalities and can be evaluated together
These solutions are direct alternatives with similar use cases
No risk: worst case, you leave with a clearer understanding of your project than before.


“WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.”
“WZ-IT moved our studio infrastructure from decentralised individual devices to a central platform: every site is securely connected via VPN, new devices are onboarded automatically and an entire site is provisioned from a template, without manual steps on location. What impressed me most is the breadth and depth of their knowledge: Timo and Robin are not a typical IT provider who sets up a server and leaves. The two of them think their way into highly complex infrastructure and software topics, work through every requirement we put in front of them, and build networking, provisioning and operations so that everything fits together in the end. WZ-IT is an excellent partner for complex software, network and architecture projects.”

Steve Kirchner
Managing Director, nextGYM GmbH

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.