Germany → worldwide
WZ-IT Logo

Taking over a Kubernetes cluster: audit and handover checklist

Timo WevelsiepTimo WevelsiepUpdated: 22.07.2026

Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.

Inheriting a cluster or handing over operations? WZ-IT audits, stabilizes and operates existing Kubernetes clusters with a clear scope - monitoring, updates, CVEs, backup and incident response as one lifecycle. See Managed Kubernetes

A Kubernetes cluster rarely changes hands cleanly. A colleague leaves, a provider is replaced, an acquisition brings foreign infrastructure - and suddenly a team is responsible for a system it did not build. Before anyone commits to response times, an honest inventory is needed. This checklist is the operational audit framework for it: not as a security review, but as the question of whether the cluster is even responsibly operable.

Table of contents

Why every takeover begins with an audit

An inherited cluster is a black box until proven otherwise. It may contain undocumented dependencies, long-deprecated components, manually applied configuration that is versioned nowhere, and recovery paths that have never been exercised. Anyone who commits to operations without an inventory is promising availability for a system whose failure modes they do not know.

The audit reverses this order: transparency first, service commitment second. It answers three questions - what is actually running here? What is the biggest risk? What must be stabilized before an SLA even makes sense? Only then are scope, responsibilities and response times agreed.

Not to be confused: an operations audit is not a security audit

A security audit checks hardening, attack surface, pod security, network policies and compliance. That is important - but it does not answer the takeover question. A cluster can be cleanly hardened and still be un-operable because no one fully knows the access paths or the restore was never tested.

The takeover audit asks about operability: do I know all the access paths? Is the version still supported? Does recovery work? Do I know which application depends on which dependency? Security is one dimension of it, not the purpose. This distinction is also why generic "Kubernetes audit checklists" rarely help here: they almost always cover hardening, not handover.

The takeover checklist

Nine dimensions that must be clarified before any operating commitment.

1. Access and inventory. Complete kubeconfig and admin access, access to the underlying infrastructure (Proxmox, provider console, bare metal), to the registry, Git repositories, DNS and certificate management. What matters is not only whether access exists, but whether it is complete - a missing registry or DNS access will block every rollout later.

2. Distribution, version and support skew. Which distribution (kubeadm, k3s, RKE2, Talos or managed)? Which Kubernetes version - and is it still within the support window? Kubernetes only maintains roughly the three most recent minor lines. A cluster several versions behind receives no security patches and can only be upgraded step by step, minor by minor. Also check: the version of the node operating systems and the platform add-ons.

3. Network: CNI, ingress, load balancer, DNS. Which CNI (Cilium, Calico, …) and which network policies? How does traffic get in - Ingress or Gateway API, which controller? Anyone still on ingress-nginx has an urgent topic here, because the community project has been retired. How does load balancing work (provider LB, MetalLB), and how are DNS and TLS automated?

4. Storage, backup and the restore test. Which StorageClasses and CSI drivers, where does persistent data live? And then the decisive question: has a full restore ever been tested? A robust concept protects several layers separately - the control plane's etcd state, the versioned cluster definitions, the persistent volumes and application-consistent database backups. A snapshot alone is none of these layers in full.

5. Identity and RBAC. How do administrators and teams authenticate - static credentials or OIDC/SSO? Who has cluster-admin, and is it traceable? Unused, over-privileged or impersonal access is a common finding.

6. Observability. Is there monitoring (metrics), logging and alerting - and do the alerts produce something actionable or just noise? Without visibility into the control plane, nodes, core components and capacity, no response time can be seriously committed.

7. Workloads and dependencies. Which applications run, which are business-critical, which external dependencies (databases, APIs, message brokers) exist? Which deployment method - GitOps, Helm, manual? A cluster whose applications were rolled out manually and undocumented carries the greatest reconstruction risk.

8. Lifecycle and updates. How were the Kubernetes version, operating systems and add-ons updated so far - planned and tested, or not at all? Is there a maintenance window, are there runbooks? The update backlog is often the single largest item of stabilization.

9. Responsibility boundaries. After the takeover, who owns which layer - control plane, nodes, platform add-ons, workloads, databases, application? Without a clear shared-responsibility mapping, exactly the gaps arise in which an incident falls through.

Red flags: how to spot a critical takeover

Some findings weigh more than others. These five signal that stabilization must come first, before an operating scope can be committed:

  • No tested restore - backups exist, but no one has ever exercised a recovery.
  • Version out of support - several minor versions behind the support window, no security patches.
  • No etcd backup, or it is unknown where and how the control plane is backed up.
  • Manually rolled-out, undocumented workloads without a versioned definition.
  • Unclear or shared admin access without a traceable mapping.

If several of these apply, the honest answer is not "we take over 24/7 operations tomorrow", but "we stabilize first".

From audit to operations

From the audit comes a transition in clear phases: discovery and a risk baseline establish the accepted current state, a stabilization phase brings monitoring, backups, updates and access to a minimum standard, and only the subsequent service definition sets scope, maintenance windows, response and responsibilities. That turns an inherited black box into a system with an agreed service commitment - instead of a promise made in the dark.

Whether the foundation holds at all often depends on the underlying infrastructure; on Proxmox and sovereign infrastructure an inherited cluster can be operated cleanly. And whether a rebuild rather than a takeover is more sensible is answered in Do we need Kubernetes?

Rather have it operated?

You'd rather not run Kubernetes yourself? WZ-IT handles setup, operations and maintenance - GDPR-compliant from Germany.

Frequently Asked Questions

Answers to the most important questions

No. A security audit checks hardening, attack surface and compliance. A takeover or handover checklist clarifies operability: do you know all the access paths, is the version still supported, does the restore work, who owns which layer? Security is part of it, but the purpose is the responsible takeover of live operations - not just an assessment of the security posture.

Because a running, inherited cluster can contain unknown dependencies, outdated components and untested recovery paths. Anyone who promises to operate it without an inventory is committing to response times for a system they do not know. The audit creates transparency and prioritizes the necessary stabilization before a service commitment is defined.

An untested restore. Very often backups or snapshots exist, but no one has ever performed a full recovery - including etcd, persistent data and application-consistent database backups. A backup that was never restored is a guess in an emergency, not a safety net.

Kubernetes only maintains roughly the three most recent minor lines, each for about a year. A cluster several versions behind support receives no security patches and can only be upgraded in several steps, because upgrades usually happen minor by minor. The running version therefore directly determines the effort and risk of the takeover.

No. A VM or volume snapshot is an additional recovery layer, but it does not automatically contain a consistent etcd state, versioned cluster definitions and application-consistent database backups. A robust concept protects several layers separately and tests the recovery of each layer.

It depends on access, documentation, version level, criticality and known issues. After an initial scoping, the audit follows; from it come the effort, stabilization measures and a realistic handover plan. A badly outdated, undocumented cluster first needs a stabilization phase before a full operating scope can responsibly be committed.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

E-Mail
[email protected]

Leading companies trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • Maho Management
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
Timo Wevelsiep & Robin Zins - CEOs of WZ-IT

Timo Wevelsiep & Robin Zins

Managing Directors of WZ-IT

1/3 - Topic Selection33%

What is your inquiry about?

Select one or more areas where we can support you.