Inference runtime
Deploy SGLang, drivers, kernels and parameters reproducibly.
WZ-IT designs SGLang as an inference layer for local AI, benchmarks models and sizes GPUs, concurrency, context and availability for the real workload.
Companies worldwide trust WZ-IT
The following are trademarks of their respective owners: SGLang (the SGLang project contributors). WZ-IT is an independent service provider and has no business, partnership, or contractual relationship with these companies. We offer independent migration, installation, hosting, and operations services.
Response time and throughput depend on model, quantisation, context length, batch size, concurrency and GPU architecture. A single token figure does not describe business operations.
We test target models with representative inputs and design routing, model cache, metrics, API protection and scaling as part of the platform.
User interface, RAG, identities, governance and business evaluation are supplied by other components. SGLang primarily exposes models through APIs.
GPU type, instance count and target metrics are determined during assessment, so there is no flat standard price.
Model, quantisation, runtime, GPU, context and scheduling form one system that must be benchmarked under realistic load.
Deploy SGLang, drivers, kernels and parameters reproducibly.
Version weights, quantisation, context and APIs.
Optimise KV cache, batching, concurrency and memory limits.
Monitor latency, throughput, failures and GPU use.
WZ-IT operates the runtime; model quality, licensing and permitted use are shared or customer-owned.
| Area | Responsibility | Scope and boundaries |
|---|---|---|
| SGLang runtime | WZ-IT | Deployment, drivers, containers, monitoring and updates. |
| Model artefacts | WZ-IT | Controlled download, storage and launch configuration. |
| Load profile | Shared | We measure while the customer supplies representative prompts. |
| API integration | Shared | Authentication, limits and clients are tested together. |
| Model approval | Customer | Suitability, licences and outputs remain customer-owned. |
| Routing and scaling | Optional WZ-IT service | Multiple models, GPUs and gateways are scoped separately. |
Serve supported language and multimodal models with optimised scheduling and caching.
Benchmark latency, time to first token, token throughput and concurrency with realistic load profiles.
Expose OpenAI-compatible endpoints to approved applications and internal platforms.
Document versions, quantisations, launch parameters and rollout paths transparently.
Protect API authentication, network paths, tenant boundaries and administrative access.
Connect Open WebUI, LiteLLM, agents, RAG services and custom applications to the inference layer.
Assess existing data and configuration, migrate them in a test run and move to managed operations through a controlled cutover.
Secure SSO, roles, administrative paths and external access for the application and existing infrastructure.
Back up all stateful components consistently and document the recovery path for the agreed scope.
Monitor and update the application and its technical dependencies and operate them under the agreed service level.
A clearly defined operating scope instead of an opaque hosting flat fee.
We set up a test instance for you, usually on the next business day. No payment details required. After seven days it is deleted unless you continue.
We combine the right compute size with ongoing operations, backups, monitoring and a service level appropriate for the criticality of SGLang. High availability and recovery targets are designed separately where needed.
We also design custom hosting architectures, integrations and migrations around SGLang. Contact us for a technical assessment.
One managed standard SGLang application is included in the Starter workload. Business and higher levels add a flexible operations allowance for planned work during regular service hours. Select compute, additional applications, storage and the appropriate service level.
A workload is one compute instance with the applications agreed for it.
One standard app per workload is already included. Additional dedicated servers count as separate workloads.
€79.90 per started TB and month, including daily encrypted offsite backup with 7-day retention.
Enquiry
Briefly describe the current state and objective for SGLang. We assess infrastructure, integration, and ongoing operations.
Weights, KV cache, context, batch size, GPU type and TTFT targets must be considered together.
| Usage scenario | Technical starting point | Key factors |
|---|---|---|
| One model and few chats | Model and GPU benchmark | Quantisation, context and interactivity are measured. |
| Several teams or API clients | Concurrency test | Throughput, queueing and response time are measured. |
| Long context or agents | Memory and prefill assessment | KV cache and input load determine capacity. |
| Multiple GPUs or models | Cluster and routing design | Parallelism and failover are planned. |
Pricing follows model, hardware, context, concurrency and target metrics.
The runtime runs on dedicated GPU hardware close to calling applications.
European inference service after benchmarking.
Deployment in your account.
Operation on suitable local GPU servers.
Combine local and external inference endpoints.
Clients use an authenticated, measured API path rather than direct GPU access.
OpenAI-compatible or project-specific clients.
TLS, authentication, quotas and request limits.
Separate keys, models and usage boundaries.
Scheduling, batching, cache and model serving.
Accelerators and kernels for the selected model.
Versioned weights and quantisations.
TTFT, token rate, queue, memory and failures.
Benchmark results only apply to the tested model, runtime, hardware, context and load profile.
Answers about models, GPUs, benchmarks and operations.
It is relevant for performance-focused serving, especially where caching and concurrency matter.
Only after measuring context, input, output and concurrency against target latency.
No. GPU type and quantity are quoted after assessment.
Yes. We deploy it on suitable local GPU hardware.
Yes. We benchmark both with identical models, prompts and concurrency.
As an alternative to managed hosting in the data centre, WZ-IT provides the hardware, configures SGLang, and handles hardening, monitoring, updates, backup and technical support. Access can be limited to the internal network or enabled through VPN and existing identities.
from EUR 349 excl. VAT / month · plus one-time provisioning and initial setup

01.06.2026
Your own ChatGPT, running entirely in-house: a familiar chat interface, a large language model on your own hardware behind it, without a single prompt going...
24.11.2025
OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a...
09.11.2025
Local AI inference means running a language or multimodal model on owned hardware rather than through a public API. Organisations gain control over data paths,...
These solutions are often used together with SGLang
These solutions offer similar functionalities and can be evaluated together
These solutions are direct alternatives with similar use cases
No risk: worst case, you leave with a clearer understanding of your project than before.


“WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.”
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.