Skip to content

Instantly share code, notes, and snippets.

@17twenty
Last active September 13, 2026 06:42
Show Gist options
  • Select an option

  • Save 17twenty/197ed2df9dd7ed63b897464674519b1a to your computer and use it in GitHub Desktop.

Select an option

Save 17twenty/197ed2df9dd7ed63b897464674519b1a to your computer and use it in GitHub Desktop.
CKAD Cheatsheet / Reference

Appendix D - Cilium, eBPF and Gateway API

A runnable networking deep dive for the Kubernetes cookbook.

This appendix assumes you completed:

appendix-c-kubeadm.md

Target versions:

Cilium:      1.20.1
Gateway API: 1.6.1
Kubernetes:  1.35

D.1 What We Are Adding

Our cluster currently looks roughly like:

Application
    |
    v
Kubernetes API
    |
    +--> Deployment controller
    |
    +--> Service objects
    |
    +--> NetworkPolicy
    |
    v
Cilium
    |
    +--> CNI
    |
    +--> NetworkPolicy enforcement
    |
    +--> kube-proxy replacement
    |
    v
eBPF dataplane

We are going to add L7 routing:

Client
  |
  v
Gateway
  |
  v
Cilium Gateway controller
  |
  v
Envoy
  |
  v
HTTPRoute
  |
  v
Service
  |
  v
Pod

Gateway API separates concerns more cleanly than the original Ingress model.

The important resource chain is:

GatewayClass
     |
     v
Gateway
     |
     v
HTTPRoute
     |
     v
Service
     |
     v
Pod

D.2 Why Gateway API

Ingress gives us roughly:

Ingress
   |
   v
Controller
   |
   v
Service

Gateway API makes roles explicit.

Platform owns
----------------
GatewayClass
Gateway


Application team owns
---------------------
HTTPRoute
Service
Deployment

Think:

GatewayClass
   -> Which implementation handles this?

Gateway
   -> Where and how does traffic enter?

HTTPRoute
   -> Where should HTTP requests go?

Service
   -> Which backend Pods receive them?

That separation becomes particularly useful in shared clusters.


D.3 Verify the Starting Cluster

Check nodes:

kubectl get nodes

Check Cilium:

cilium status --wait

Confirm kube-proxy is absent:

kubectl get daemonset kube-proxy \
  -n kube-system

Expected:

Error from server (NotFound)

Confirm replacement mode:

kubectl -n kube-system exec ds/cilium -- \
  cilium-dbg status \
  | grep KubeProxyReplacement

Expected:

KubeProxyReplacement:   True

This matters because Cilium's Gateway API implementation requires kube-proxy replacement.


D.4 Install the Gateway API CRDs

Gateway API is not built into Kubernetes core as a set of automatically available resources.

Install the Gateway API 1.6.1 standard channel:

kubectl apply --server-side \
  -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.1/standard-install.yaml

Discover the new APIs:

kubectl api-resources \
  --api-group=gateway.networking.k8s.io

You should see resources including:

gatewayclasses
gateways
httproutes
grpcroutes
referencegrants

Ask the API:

kubectl explain gateway

Then:

kubectl explain httproute.spec

This should feel familiar from the CKAD cookbook.

We installed new API types.

We have not yet proven that anything implements them.


D.5 Enable Cilium's Gateway API Controller

A normal Cilium Gateway creates a Service of type:

LoadBalancer

On a cloud cluster, the cloud load-balancer integration may give that Service an external address.

Our kubeadm lab deliberately has no cloud load balancer.

Rather than adding another product merely to make the exercise work, we will use Cilium's Gateway host-network mode.

That means:

Gateway listener
      |
      v
Cilium Envoy
      |
      v
node network interface

Use ports above 1023 for this lab.

Find the API server address:

export CONTROL_PLANE_IP="$(
  kubectl get node k8s-control \
    -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'
)"

Check:

echo "$CONTROL_PLANE_IP"

Upgrade the existing Cilium installation:

cilium upgrade 1.20.1 \
  --set kubeProxyReplacement=true \
  --set k8sServiceHost="$CONTROL_PLANE_IP" \
  --set k8sServicePort=6443 \
  --set gatewayAPI.enabled=true \
  --set gatewayAPI.hostNetwork.enabled=true

Wait:

cilium status --wait

Inspect Cilium components:

kubectl get pods \
  -n kube-system \
  -o wide

Look for the Cilium agents, operator and Envoy components.


D.6 Find the GatewayClass

Check:

kubectl get gatewayclass

You should see:

cilium

Inspect:

kubectl describe gatewayclass cilium

The relationship is:

GatewayClass/cilium
        |
        v
Cilium Gateway controller

GatewayClass is cluster-scoped.

It tells Kubernetes which implementation is responsible for Gateways using that class.


D.7 Build Two Backend Versions

Create a namespace for this lab:

kubectl create namespace gateway-lab

Set it as the current default:

kubectl config set-context \
  --current \
  --namespace=gateway-lab

Create version one:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-v1
spec:
  replicas: 2
  selector:
    matchLabels:
      app: api-v1
  template:
    metadata:
      labels:
        app: api-v1
        version: v1
    spec:
      containers:
        - name: web
          image: busybox:1.36
          command:
            - sh
            - -c
            - |
              mkdir -p /www
              echo "hello from api-v1" > /www/index.html
              echo "allowed from api-v1" > /www/allowed
              echo "denied from api-v1" > /www/denied
              httpd -f -p 8080 -h /www
          ports:
            - containerPort: 8080
---
apiVersion: v1
kind: Service
metadata:
  name: api-v1
spec:
  selector:
    app: api-v1
  ports:
    - port: 80
      targetPort: 8080

Save as:

api-v1.yaml

Apply:

kubectl apply -f api-v1.yaml

Create version two:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-v2
spec:
  replicas: 2
  selector:
    matchLabels:
      app: api-v2
  template:
    metadata:
      labels:
        app: api-v2
        version: v2
    spec:
      containers:
        - name: web
          image: busybox:1.36
          command:
            - sh
            - -c
            - |
              mkdir -p /www
              echo "hello from api-v2" > /www/index.html
              httpd -f -p 8080 -h /www
          ports:
            - containerPort: 8080
---
apiVersion: v1
kind: Service
metadata:
  name: api-v2
spec:
  selector:
    app: api-v2
  ports:
    - port: 80
      targetPort: 8080

Save as:

api-v2.yaml

Apply:

kubectl apply -f api-v2.yaml

Check:

kubectl get deployments
kubectl get pods -o wide
kubectl get services

Before touching Gateway API, prove ordinary Kubernetes Services work.

Create a client:

kubectl run client \
  --image=curlimages/curl \
  --restart=Never \
  --command -- \
  sleep 3600

Wait:

kubectl wait \
  --for=condition=Ready \
  pod/client \
  --timeout=90s

Test:

kubectl exec client -- \
  curl -s http://api-v1

Expected:

hello from api-v1

Then:

kubectl exec client -- \
  curl -s http://api-v2

Expected:

hello from api-v2

We have now proved:

Pods
  |
  v
Services
  |
  v
Cilium Service dataplane

works before adding L7 routing.


D.8 Create a Gateway

Create:

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: public
spec:
  gatewayClassName: cilium

  listeners:
    - name: http
      protocol: HTTP
      port: 8080
      hostname: api.example.test

      allowedRoutes:
        namespaces:
          from: Same

Save as:

gateway.yaml

Apply:

kubectl apply -f gateway.yaml

Inspect:

kubectl get gateway public

Then:

kubectl describe gateway public

Look at:

Status
Conditions
Addresses
Listeners

Query conditions:

kubectl get gateway public \
  -o jsonpath='{range .status.conditions[*]}{.type}={.status}{" reason="}{.reason}{"\n"}{end}'

We want conditions indicating the Gateway has been accepted and programmed.

The object path is:

Gateway
   |
   v
Kubernetes API
   |
   v
Cilium Gateway controller
   |
   v
Envoy listener :8080

D.9 Create an HTTPRoute

Create:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api
spec:

  parentRefs:
    - name: public

  hostnames:
    - api.example.test

  rules:
    - backendRefs:
        - name: api-v1
          port: 80

Save as:

route.yaml

Apply:

kubectl apply -f route.yaml

Inspect:

kubectl get httproute api

Then:

kubectl describe httproute api

Query status:

kubectl get httproute api \
  -o jsonpath='{range .status.parents[*].conditions[*]}{.type}={.status}{" reason="}{.reason}{"\n"}{end}'

The complete chain is now:

GatewayClass/cilium
        |
        v
Gateway/public
        |
        v
HTTPRoute/api
        |
        v
Service/api-v1
        |
        v
EndpointSlice
        |
        v
Pods

Notice how much of this is still ordinary Kubernetes API composition.


D.10 Send Traffic Through the Gateway

Because we enabled host-network mode, the listener exists on the node network.

Find node addresses:

kubectl get nodes -o wide

Pick one reachable node IP:

export GATEWAY_NODE_IP="$(
  kubectl get node k8s-worker \
    -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'
)"

If you are using a single-node lab, use k8s-control instead.

Check:

echo "$GATEWAY_NODE_IP"

Request:

curl \
  -H 'Host: api.example.test' \
  "http://${GATEWAY_NODE_IP}:8080/"

Expected:

hello from api-v1

Try without the hostname:

curl \
  "http://${GATEWAY_NODE_IP}:8080/"

The result should not match our route in the same way.

Why?

Our HTTPRoute declares:

api.example.test

Gateway API routing is based on declared listeners and route matches, not merely on a port being open.


D.11 Observe the Resources Cilium Created

Start with Kubernetes:

kubectl get gateway
kubectl get httproute
kubectl get svc

Inspect Cilium:

cilium status

Inspect Envoy-related resources:

kubectl get ciliumenvoyconfigs \
  -A

Depending on Cilium's generated configuration and version, you should see resources representing L7 proxy configuration.

This is an important transition.

The application developer created:

Gateway
HTTPRoute

The implementation created lower-level configuration needed to make those resources real.

Conceptually:

HTTPRoute
    |
    v
Kubernetes API
    |
    v
Cilium controller
    |
    v
Envoy configuration
    |
    v
listener / routes / clusters
    |
    v
actual HTTP traffic

D.12 Break It - Point at a Missing Backend

Patch the route:

kubectl patch httproute api \
  --type=json \
  -p='[
    {
      "op":"replace",
      "path":"/spec/rules/0/backendRefs/0/name",
      "value":"api-does-not-exist"
    }
  ]'

Inspect:

kubectl describe httproute api

Pay attention to route conditions.

You should see that the route cannot completely resolve its references.

Query:

kubectl get httproute api \
  -o jsonpath='{range .status.parents[*].conditions[*]}{.type}={.status}{" reason="}{.reason}{" message="}{.message}{"\n"}{end}'

Try traffic:

curl \
  -H 'Host: api.example.test' \
  "http://${GATEWAY_NODE_IP}:8080/"

Do not start debugging at Envoy.

Start at the API:

Gateway
   |
   v
HTTPRoute
   |
   v
backendRef
   |
   X
Service missing

Fix it:

kubectl patch httproute api \
  --type=json \
  -p='[
    {
      "op":"replace",
      "path":"/spec/rules/0/backendRefs/0/name",
      "value":"api-v1"
    }
  ]'

Wait a moment and inspect again:

kubectl describe httproute api

Then:

curl \
  -H 'Host: api.example.test' \
  "http://${GATEWAY_NODE_IP}:8080/"

Expected:

hello from api-v1

Status is part of the API contract.

Use it.


D.13 Header-Based Routing

Replace the HTTPRoute with:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api
spec:

  parentRefs:
    - name: public

  hostnames:
    - api.example.test

  rules:

    - matches:
        - headers:
            - name: x-api-version
              value: v2
      backendRefs:
        - name: api-v2
          port: 80

    - backendRefs:
        - name: api-v1
          port: 80

Apply:

kubectl apply -f route.yaml

Default request:

curl \
  -H 'Host: api.example.test' \
  "http://${GATEWAY_NODE_IP}:8080/"

Expected:

hello from api-v1

Version-two request:

curl \
  -H 'Host: api.example.test' \
  -H 'x-api-version: v2' \
  "http://${GATEWAY_NODE_IP}:8080/"

Expected:

hello from api-v2

Now the routing decision is:

Request
   |
   +-- x-api-version=v2 --> api-v2
   |
   +-- otherwise --------> api-v1

This is well beyond what a Kubernetes Service selector can express.


D.14 Weighted Backends - A Simple Canary

Replace the route rules with:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api
spec:

  parentRefs:
    - name: public

  hostnames:
    - api.example.test

  rules:
    - backendRefs:

        - name: api-v1
          port: 80
          weight: 90

        - name: api-v2
          port: 80
          weight: 10

Apply:

kubectl apply -f route.yaml

Send several requests:

for i in $(seq 1 30); do
  curl -s \
    -H 'Host: api.example.test' \
    "http://${GATEWAY_NODE_IP}:8080/"
done | sort | uniq -c

You should see requests reaching both versions.

Do not expect exactly 90/10 over a tiny sample.

The important difference from the simple canary in the CKAD cookbook is:

Old approach
------------

9 Pods stable
1 Pod canary

Service selects all 10


Gateway API approach
--------------------

HTTPRoute
  |
  +-- weight 90 -> Service v1
  |
  +-- weight 10 -> Service v2

The routing intent is now explicit in the API.


D.15 Return to a Single Backend

Restore:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api
spec:
  parentRefs:
    - name: public
  hostnames:
    - api.example.test
  rules:
    - backendRefs:
        - name: api-v1
          port: 80

Apply:

kubectl apply -f route.yaml

D.16 Cross-Namespace Backends and ReferenceGrant

Gateway API deliberately makes cross-namespace references explicit.

Create another namespace:

kubectl create namespace payments

Create a backend there:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: payments
  namespace: payments
spec:
  replicas: 1
  selector:
    matchLabels:
      app: payments
  template:
    metadata:
      labels:
        app: payments
    spec:
      containers:
        - name: web
          image: busybox:1.36
          command:
            - sh
            - -c
            - |
              mkdir -p /www
              echo "hello from payments" > /www/index.html
              httpd -f -p 8080 -h /www
          ports:
            - containerPort: 8080
---
apiVersion: v1
kind: Service
metadata:
  name: payments
  namespace: payments
spec:
  selector:
    app: payments
  ports:
    - port: 80
      targetPort: 8080

Save as:

payments.yaml

Apply:

kubectl apply -f payments.yaml

Now point our route at it:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api
  namespace: gateway-lab
spec:

  parentRefs:
    - name: public

  hostnames:
    - api.example.test

  rules:
    - backendRefs:
        - name: payments
          namespace: payments
          port: 80

Apply:

kubectl apply -f route.yaml

Inspect:

kubectl describe httproute api

The reference should not be considered valid merely because the Service exists.

Why?

The route lives in:

gateway-lab

but wants to reference a Service in:

payments

Gateway API requires the target namespace to explicitly permit that reference.

Create:

apiVersion: gateway.networking.k8s.io/v1beta1
kind: ReferenceGrant
metadata:
  name: allow-gateway-lab
  namespace: payments
spec:

  from:
    - group: gateway.networking.k8s.io
      kind: HTTPRoute
      namespace: gateway-lab

  to:
    - group: ""
      kind: Service
      name: payments

Save as:

referencegrant.yaml

Apply:

kubectl apply -f referencegrant.yaml

Inspect the route again:

kubectl describe httproute api

Then test:

curl \
  -H 'Host: api.example.test' \
  "http://${GATEWAY_NODE_IP}:8080/"

Expected:

hello from payments

The security boundary is:

gateway-lab HTTPRoute
        |
        | asks to reference
        v
payments/Service
        |
        X
not allowed by default
        |
        v
ReferenceGrant in payments
        |
        v
reference accepted

The target namespace owns the permission.

That is a powerful multi-team platform primitive.


D.17 Restore the Local Backend

Restore route.yaml:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api
  namespace: gateway-lab
spec:
  parentRefs:
    - name: public
  hostnames:
    - api.example.test
  rules:
    - backendRefs:
        - name: api-v1
          port: 80

Apply:

kubectl apply -f route.yaml

D.18 Look at Cilium Identities

List Cilium endpoints:

kubectl -n kube-system exec ds/cilium -- \
  cilium-dbg endpoint list

Look at identities:

kubectl get ciliumidentities

Cilium does not need to think only in terms of transient Pod IP addresses.

It associates workloads with identities derived from labels.

Conceptually:

Pod labels
    |
    v
Cilium identity
    |
    v
policy / dataplane decisions

That becomes particularly useful when Pods are recreated and their IP addresses change.

The workload identity can remain logically stable even while endpoints change.


D.19 Inspect the eBPF Service Dataplane

List Cilium services:

kubectl -n kube-system exec ds/cilium -- \
  cilium-dbg service list

List the underlying BPF load-balancer state:

kubectl -n kube-system exec ds/cilium -- \
  cilium-dbg bpf lb list

Compare with:

kubectl get services -A

and:

kubectl get endpointslices -A

The layers are:

Kubernetes Service
      |
      v
EndpointSlices
      |
      v
Cilium observes API state
      |
      v
BPF maps
      |
      v
packet forwarding

The important idea is not to memorise BPF map output.

It is to understand that the high-level API is compiled into lower-level dataplane state.


D.20 Standard NetworkPolicy Still Matters

Cilium does not require applications to use Cilium-specific policy objects.

The normal Kubernetes API still works:

networking.k8s.io/v1
NetworkPolicy

Create another in-cluster client:

kubectl run restricted-client \
  --image=curlimages/curl \
  --labels=role=client \
  --restart=Never \
  --command -- \
  sleep 3600

Wait:

kubectl wait \
  --for=condition=Ready \
  pod/restricted-client \
  --timeout=90s

Baseline:

kubectl exec restricted-client -- \
  curl -s http://api-v1

Create default deny for api-v1:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: api-v1-deny
spec:
  podSelector:
    matchLabels:
      app: api-v1
  policyTypes:
    - Ingress

Save as:

networkpolicy-deny.yaml

Apply:

kubectl apply -f networkpolicy-deny.yaml

Retry:

kubectl exec restricted-client -- \
  curl \
  --max-time 3 \
  http://api-v1

It should fail.

The portable API remains:

NetworkPolicy
      |
      v
Kubernetes API
      |
      v
Cilium
      |
      v
actual enforcement

Delete it before continuing:

kubectl delete networkpolicy api-v1-deny

Verify connectivity returns:

kubectl exec restricted-client -- \
  curl -s http://api-v1

This is why the main CKAD cookbook keeps NetworkPolicy and punts the Cilium implementation detail here.


D.21 CiliumNetworkPolicy - Deliberately Cross the Portability Boundary

Sometimes the portable Kubernetes API does not express the policy we want.

Cilium provides:

CiliumNetworkPolicy

as a CRD for additional capabilities.

This is the point where we intentionally choose a vendor/platform-specific API.

We will enforce an HTTP-layer rule.

Create an L7 test server:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: l7-api
spec:
  replicas: 1
  selector:
    matchLabels:
      app: l7-api
  template:
    metadata:
      labels:
        app: l7-api
    spec:
      containers:
        - name: web
          image: busybox:1.36
          command:
            - sh
            - -c
            - |
              mkdir -p /www
              echo "you reached allowed" > /www/allowed
              echo "you reached denied" > /www/denied
              httpd -f -p 8080 -h /www
          ports:
            - containerPort: 8080
---
apiVersion: v1
kind: Service
metadata:
  name: l7-api
spec:
  selector:
    app: l7-api
  ports:
    - port: 80
      targetPort: 8080

Save as:

l7-api.yaml

Apply:

kubectl apply -f l7-api.yaml

Baseline:

kubectl exec restricted-client -- \
  curl -s http://l7-api/allowed

Then:

kubectl exec restricted-client -- \
  curl -s http://l7-api/denied

Both should work.

Now create:

apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: l7-api-policy
spec:

  endpointSelector:
    matchLabels:
      app: l7-api

  ingress:
    - fromEndpoints:
        - matchLabels:
            role: client

      toPorts:
        - ports:
            - port: "8080"
              protocol: TCP

          rules:
            http:
              - method: GET
                path: "/allowed$"

Save as:

cilium-l7-policy.yaml

Apply:

kubectl apply -f cilium-l7-policy.yaml

Allowed:

kubectl exec restricted-client -- \
  curl -i http://l7-api/allowed

Denied:

kubectl exec restricted-client -- \
  curl -i http://l7-api/denied

The policy is now making an L7 decision:

source identity
      |
      v
TCP :8080
      |
      v
HTTP GET
      |
      v
path /allowed
      |
      +--> allow

anything else
      |
      +--> deny

This is functionality beyond the standard Kubernetes NetworkPolicy API.

That power comes with a portability tradeoff.

Delete when finished:

kubectl delete ciliumnetworkpolicy l7-api-policy

D.22 Enable Hubble

Cilium gives us another useful implementation-specific feature:

Hubble

Hubble provides network observability over Cilium-managed endpoints.

Enable Relay:

cilium hubble enable

Wait:

cilium status --wait

You should see Hubble Relay become available.


D.23 Install the Hubble CLI

On Linux:

HUBBLE_VERSION="$(
  curl -s \
  https://raw.githubusercontent.com/cilium/hubble/main/stable.txt
)"

HUBBLE_ARCH=amd64

if [ "$(uname -m)" = "aarch64" ]; then
  HUBBLE_ARCH=arm64
fi

Download:

curl -L --fail --remote-name-all \
  "https://github.com/cilium/hubble/releases/download/${HUBBLE_VERSION}/hubble-linux-${HUBBLE_ARCH}.tar.gz"{,.sha256sum}

Verify:

sha256sum --check \
  "hubble-linux-${HUBBLE_ARCH}.tar.gz.sha256sum"

Install:

sudo tar xzvfC \
  "hubble-linux-${HUBBLE_ARCH}.tar.gz" \
  /usr/local/bin

Cleanup:

rm "hubble-linux-${HUBBLE_ARCH}.tar.gz"{,.sha256sum}

Check:

hubble version

Validate Relay:

hubble status -P

-P tells the CLI to set up the required local port-forward automatically.


D.24 Observe Traffic Instead of Guessing

Generate traffic:

for i in $(seq 1 5); do
  curl -s \
    -H 'Host: api.example.test' \
    "http://${GATEWAY_NODE_IP}:8080/" >/dev/null
done

Observe recent flows:

hubble observe -P \
  --last 20

Observe traffic to the backend namespace:

hubble observe -P \
  --namespace gateway-lab \
  --last 30

You can also filter by Pod:

hubble observe -P \
  --pod gateway-lab/api-v1 \
  --last 20

The debugging model improves from:

"It looks like networking."

to:

What flow happened?

From which identity?

To which destination?

Was it forwarded or dropped?

At which layer?

D.25 Observe a Policy Drop

Reapply the Cilium L7 policy:

kubectl apply -f cilium-l7-policy.yaml

Generate an allowed request:

kubectl exec restricted-client -- \
  curl -s http://l7-api/allowed

Generate a denied request:

kubectl exec restricted-client -- \
  curl -s http://l7-api/denied || true

Inspect Hubble:

hubble observe -P \
  --pod gateway-lab/l7-api \
  --last 30

Now we can tie together:

CiliumNetworkPolicy
       |
       v
policy calculation
       |
       v
Envoy / dataplane enforcement
       |
       v
Hubble flow event

Delete the policy:

kubectl delete ciliumnetworkpolicy l7-api-policy

D.26 Gateway API Debugging Order

When Gateway traffic fails, debug from the API downward.

Do not begin with packet captures.

Use this order:

GatewayClass
     |
     v
Gateway
     |
     v
HTTPRoute
     |
     v
Service
     |
     v
EndpointSlice
     |
     v
Pod readiness
     |
     v
Cilium / Envoy
     |
     v
Hubble
     |
     v
Linux dataplane

Useful commands:

kubectl get gatewayclass
kubectl describe gateway public
kubectl describe httproute api
kubectl get service api-v1
kubectl get endpointslices \
  -l kubernetes.io/service-name=api-v1
kubectl get pods \
  -l app=api-v1 \
  -o wide
cilium status
hubble observe -P --last 30

Only go further down once the upper layer is known to be correct.


D.27 Gateway API Status Is Part of the Contract

Query Gateway conditions:

kubectl get gateway public \
  -o jsonpath='{range .status.conditions[*]}{.type}={.status}{" reason="}{.reason}{"\n"}{end}'

Query HTTPRoute conditions:

kubectl get httproute api \
  -o jsonpath='{range .status.parents[*].conditions[*]}{.type}={.status}{" reason="}{.reason}{"\n"}{end}'

This is the same lesson as:

spec
  -> what I want

status
  -> what the controller observed

from the core cookbook.

Gateway API makes the controller relationship particularly visible.


D.28 Optional - Why We Used Host-Network Mode

By default, Cilium's Gateway controller creates a:

Service type=LoadBalancer

That works naturally when a platform already provides a LoadBalancer implementation.

Our kubeadm lab does not.

Possible real-world solutions include:

cloud load balancer integration
Cilium LB IPAM + L2 announcements
Cilium LB IPAM + BGP
external load balancer
Cilium host-network Gateway mode

For this lab we chose:

host-network mode

because it lets us focus on:

Gateway API
Cilium
Envoy
routing
policy
observability

without first building an entire bare-metal load-balancer control plane.

That is a learning choice, not a universal production recommendation.


D.29 Optional - The LoadBalancer Problem on Bare Metal

A Kubernetes Service can say:

spec:
  type: LoadBalancer

but that does not magically create an external load balancer.

Again:

API object exists

does not imply:

implementation exists

On bare metal someone still needs to answer:

Who allocates the external IP?

Who advertises it onto the network?

Who makes traffic reach the nodes?

Cilium can participate in those layers with:

LB IPAM
L2 Announcements
BGP Control Plane
Node IPAM LB

Those are excellent platform-engineering topics.

They are intentionally outside this first Gateway API lab.


D.30 Clean Up the Lab

Delete Gateway resources:

kubectl delete httproute api \
  --ignore-not-found

kubectl delete gateway public \
  --ignore-not-found

Delete application resources:

kubectl delete deployment \
  api-v1 \
  api-v2 \
  l7-api \
  --ignore-not-found

kubectl delete service \
  api-v1 \
  api-v2 \
  l7-api \
  --ignore-not-found

kubectl delete pod \
  client \
  restricted-client \
  --ignore-not-found

Delete cross-namespace resources:

kubectl delete namespace payments \
  --ignore-not-found

Delete the lab namespace:

kubectl delete namespace gateway-lab

Switch your current context back to default:

kubectl config set-context \
  --current \
  --namespace=default

We normally leave the Gateway API CRDs and Cilium installation in place because they are cluster-level platform components.

For a fully disposable lab, reset the cluster using Appendix C instead.


D.31 The Full Request Path

We can now describe a request from outside the cluster.

curl
 |
 | Host: api.example.test
 v
Linux node :8080
 |
 v
Cilium / Envoy
 |
 | listener selected
 v
Gateway/public
 |
 | route selected
 v
HTTPRoute/api
 |
 | backendRef
 v
Service/api-v1
 |
 | endpoints
 v
EndpointSlice
 |
 v
Pod IP
 |
 v
Cilium dataplane
 |
 v
container

At the same time:

Kubernetes API
      |
      +--> stores Gateway
      |
      +--> stores HTTPRoute
      |
      +--> stores Service
      |
      +--> stores EndpointSlice
      |
      v
controllers observe
      |
      v
lower-level state changes

That is Kubernetes reconciliation expressed as networking.


D.32 The Architecture We Now Understand

At the beginning of the cookbook:

Service -> Pods

was enough.

Now we can expand it:

                         Kubernetes API
                               |
              +----------------+----------------+
              |                |                |
              v                v                v
          Gateway          HTTPRoute        Service
              |                |                |
              +--------+-------+                |
                       |                        |
                       v                        v
                 Cilium controller         EndpointSlice
                       |                        |
                       v                        |
                    Envoy                       |
                       |                        |
                       +-----------+------------+
                                   |
                                   v
                              Cilium agent
                                   |
                                   v
                                eBPF
                                   |
                                   v
                              Linux kernel
                                   |
                                   v
                                  Pod

And with observability:

packet / request
      |
      v
Cilium dataplane
      |
      +----> enforcement
      |
      +----> Hubble
               |
               v
          observable flow

D.33 Portability vs Platform Power

The core cookbook deliberately favoured APIs such as:

Service
NetworkPolicy
Ingress

because they are portable Kubernetes APIs.

This appendix used:

Gateway API

which is portable across conformant Gateway implementations.

Then we deliberately crossed into:

CiliumNetworkPolicy
Cilium identities
Hubble
eBPF maps
Cilium host-network Gateway mode

Those are implementation-specific.

That is not inherently bad.

It is a tradeoff.

Portable API
    |
    +--> easier platform portability
    |
    +--> common Kubernetes mental model


Implementation-specific API
    |
    +--> richer platform capability
    |
    +--> tighter coupling to the implementation

The important thing is to know which side of the boundary you are on.


D.34 The Model to Remember

For networking:

Application intent
      |
      v
Kubernetes API
      |
      v
Controller / network implementation
      |
      v
Envoy / eBPF / Linux
      |
      v
actual traffic

For debugging:

API status
   |
   v
references
   |
   v
Services / endpoints
   |
   v
implementation status
   |
   v
observed flows
   |
   v
dataplane

Do not jump straight to the bottom.

Prove each layer.


D.35 Where You Are Now

The learning path has become:

kind
  |
  v
"Give me Kubernetes"
  |
  v
CKAD cookbook
  |
  v
"Teach me Kubernetes APIs"
  |
  v
kubeadm
  |
  v
"Show me where Kubernetes comes from"
  |
  v
Cilium
  |
  v
"Show me how networking is implemented"
  |
  v
Gateway API
  |
  v
"Give platform and application teams a modern routing API"
  |
  v
Hubble / eBPF
  |
  v
"Show me what the dataplane is actually doing"

At this point the cluster should feel much less magical.

You can follow a concept from:

YAML

all the way to:

Linux packet forwarding

without confusing those layers with one another.


Reference Versions

This appendix was written against:

Kubernetes:  1.35
Cilium:      1.20.1
Gateway API: 1.6.1

Cilium 1.20.1 officially supports Kubernetes 1.35.

Its Gateway API implementation supports Gateway API 1.6.1 and requires:

kubeProxyReplacement=true

with L7 proxy support enabled.

Cilium host-network Gateway mode is used here specifically so the kubeadm lab does not require a separate LoadBalancer implementation.

Always check the matching upstream documentation when moving to a newer Cilium or Gateway API version.

GitOps and Platform Delivery for Kubernetes

A hands-on companion to the Kubernetes cookbook.

The Kubernetes cookbook teaches you how Kubernetes reconciles desired state into running workloads.

This handbook takes the next step:

How does application code safely become desired state in a real cluster?

The goal is not to memorise Argo CD commands.

The goal is to understand a modern delivery system well enough that you can reason about it, debug it, secure it, and eventually build a platform around it.

We will build this path:

developer
   |
   | git push / pull request
   v
application repository
   |
   | CI tests, builds, scans
   v
OCI registry
   |
   | immutable image + digest
   v
promotion pull request
   |
   v
GitOps repository
   |
   | reviewed desired-state change
   v
Argo CD
   |
   | reconciliation
   v
Kubernetes API
   |
   v
Argo Rollouts
   |
   | canary / blue-green / analysis / promotion
   v
running application

By the end, we will have replaced the classic pipeline:

CI job
  |
  | kubectl apply
  | cluster-admin kubeconfig
  v
production

with:

CI
 |
 +-- builds an artifact
 +-- signs / attests it
 +-- proposes a desired-state change
 |
 v
Git pull request
 |
 | human / policy review
 v
Git
 |
 | pulled by controller
 v
cluster

Sections are marked:

  • [DEV] - application developer knowledge
  • [OPS] - operating delivery systems
  • [PLATFORM] - platform engineering and multi-team design
  • [DEEP DIVE] - concepts worth understanding beyond the immediate lab

The recurring teaching loop is the same as the main Kubernetes cookbook:

problem
  |
  v
mental model
  |
  v
small experiment
  |
  v
observe the controllers
  |
  v
change one thing
  |
  v
observe the consequence
  |
  v
break an assumption
  |
  v
explain why

Part I - Git Becomes Desired State

0. Lab Setup [DEV] [OPS]

This handbook assumes you completed enough of the Kubernetes cookbook to be comfortable with:

  • Deployments and Services
  • spec versus status
  • rollouts and ReplicaSets
  • RBAC
  • Kustomize
  • container images
  • kind

The examples assume the existing lab cluster:

kubectl config use-context kind-ckad

Check it:

kubectl get nodes

We will install platform components into their own namespaces rather than the cookbook namespace.

Useful local tools:

git
kubectl
docker
kind

For the complete hosted Git workflow, a GitHub account is convenient.

Optional but useful on macOS:

brew install gh argocd

Argo Rollouts will be installed later.

Two repositories, not one

We will eventually use two repositories:

web-app
  |
  +-- application source
  +-- Dockerfile
  +-- tests
  +-- CI workflow

platform-gitops
  |
  +-- Kubernetes desired state
  +-- environment overlays
  +-- image digest promoted to each environment
  +-- Argo CD Application definitions

That separation is intentional.

The application repository answers:

What software should we build?

The GitOps repository answers:

What exact version of that software should this environment run?

Those are different decisions.


1. Why GitOps Exists [DEV] [OPS]

Suppose CI does this after every merge:

docker build -t registry.example.com/web:latest .
docker push registry.example.com/web:latest
kubectl set image deployment/web \
  web=registry.example.com/web:latest

It works.

It also quietly creates several problems.

CI needs credentials capable of modifying the cluster.

The production state may differ from anything committed to Git.

A mutable tag such as latest does not uniquely identify the bytes being deployed.

A failed CI system can leave the cluster half-mutated.

Auditing becomes:

What is running?
Who changed it?
Which pipeline ran?
Which image did latest mean at that moment?

GitOps changes the direction of authority.

Instead of:

CI ---> cluster

we use:

CI ---> Git
         |
         v
     controller ---> cluster

The cluster-side controller continuously compares:

Git desired state
       |
       | compare
       v
cluster actual state

and reconciles differences.

If this sounds familiar, it should.

Kubernetes already taught us:

spec
 |
 v
controller
 |
 v
actual state

GitOps adds another reconciliation layer:

Git
 |
 v
GitOps controller
 |
 v
Kubernetes spec
 |
 v
Kubernetes controllers
 |
 v
actual state

The important idea is not "YAML in Git".

The important idea is:

A versioned, reviewable source of desired state is continuously reconciled by software running near the target system.


2. Install Argo CD [DEV] [OPS]

Create its namespace:

kubectl create namespace argocd

Install the standard non-HA manifests:

kubectl apply \
  --server-side \
  --force-conflicts \
  -n argocd \
  -f https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml

Wait for the API/UI server:

kubectl rollout status \
  deployment/argocd-server \
  -n argocd \
  --timeout=300s

See what was installed:

kubectl get pods -n argocd

You should see several controllers and supporting components.

Argo CD is itself a Kubernetes application.

That matters because everything we learned earlier still applies:

Pods
Services
ServiceAccounts
RBAC
CRDs
controllers
status

Look at the new API

Argo installed CustomResourceDefinitions.

Find them:

kubectl get crd | grep argoproj

The important one initially is:

applications.argoproj.io

Ask Kubernetes about it:

kubectl explain application.spec

We have extended the Kubernetes API with a new desired-state object.

Optional: open the UI

Forward the Argo CD API server:

kubectl port-forward \
  -n argocd \
  svc/argocd-server \
  8080:443

Retrieve the initial admin password from another terminal:

kubectl -n argocd get secret \
  argocd-initial-admin-secret \
  -o jsonpath='{.data.password}' \
  | base64 -d

echo

Then browse to:

https://127.0.0.1:8080

Username:

admin

The UI is useful, but do not let it hide the API from you.

Throughout this handbook we will keep inspecting the underlying Kubernetes resources.


3. Create the GitOps Repository [DEV]

We need a Git repository Argo CD can read.

For the first lab, make it public so repository authentication does not distract from the reconciliation model.

Later we will discuss private repositories and production credentials.

Set your GitHub username:

export GITHUB_USER="YOUR_GITHUB_USERNAME"

Choose a repository name:

export GITOPS_REPO_NAME="platform-gitops"
export GITOPS_REPO="https://github.com/${GITHUB_USER}/${GITOPS_REPO_NAME}.git"

Create a working directory:

mkdir -p ~/gitops-lab/platform-gitops
cd ~/gitops-lab/platform-gitops

git init -b main

Create the initial structure:

mkdir -p apps/web/base
mkdir -p environments/dev/web

Create the base Deployment:

cat > apps/web/base/deployment.yaml <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
spec:
  replicas: 2
  selector:
    matchLabels:
      app: web
  template:
    metadata:
      labels:
        app: web
    spec:
      containers:
        - name: web
          image: example/web:bootstrap
          ports:
            - containerPort: 80
          readinessProbe:
            httpGet:
              path: /
              port: 80
            initialDelaySeconds: 1
            periodSeconds: 3
EOF

Create the Service:

cat > apps/web/base/service.yaml <<'EOF'
apiVersion: v1
kind: Service
metadata:
  name: web
spec:
  selector:
    app: web
  ports:
    - port: 80
      targetPort: 80
EOF

Create the base Kustomization:

cat > apps/web/base/kustomization.yaml <<'EOF'
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
  - deployment.yaml
  - service.yaml
EOF

Now create the development overlay:

cat > environments/dev/web/kustomization.yaml <<'EOF'
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: web-dev
resources:
  - ../../../apps/web/base
images:
  - name: example/web
    newName: nginx
    newTag: 1.27-alpine
EOF

Render it locally:

kubectl kustomize environments/dev/web

Notice what happened to:

image: example/web:bootstrap

The overlay changed it to:

image: nginx:1.27-alpine

The Git repository now contains enough information to describe the development environment.

Commit it:

git add .
git commit -m "bootstrap web development environment"

Push it

If you use GitHub CLI:

gh auth status

Create a public repository and push:

gh repo create "${GITOPS_REPO_NAME}" \
  --public \
  --source=. \
  --remote=origin \
  --push

Otherwise create the repository through your Git provider and push it normally.

At this point:

Git contains desired state

but

nothing is reconciling it yet

4. Create Your First Argo CD Application [DEV] [OPS]

An Argo CD Application tells Argo:

where desired state lives
       +
what path to render
       +
where to deploy it

Create one:

cat > /tmp/web-dev-application.yaml <<EOF
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: web-dev
  namespace: argocd
spec:
  project: default
  source:
    repoURL: ${GITOPS_REPO}
    targetRevision: main
    path: environments/dev/web
  destination:
    server: https://kubernetes.default.svc
    namespace: web-dev
  syncPolicy:
    syncOptions:
      - CreateNamespace=true
EOF

Apply it:

kubectl apply -f /tmp/web-dev-application.yaml

Inspect it:

kubectl get application web-dev -n argocd

Watch its status:

kubectl get application web-dev \
  -n argocd \
  -w

Because we have not enabled automated sync, Argo should discover that Git and the cluster differ.

Inspect the sync state directly:

kubectl get application web-dev \
  -n argocd \
  -o jsonpath='{.status.sync.status}{"\n"}'

You should see:

OutOfSync

That is the GitOps equivalent of seeing:

spec != actual state

Enable reconciliation

Patch the Application so Argo may automatically sync:

kubectl patch application web-dev \
  -n argocd \
  --type merge \
  -p '{
    "spec": {
      "syncPolicy": {
        "automated": {},
        "syncOptions": ["CreateNamespace=true"]
      }
    }
  }'

Watch the namespace appear:

kubectl get namespace web-dev -w

Then inspect the workload:

kubectl get deployment,service,pods \
  -n web-dev

Argo has now performed the equivalent of:

kubectl apply -k environments/dev/web

but the important difference is ownership of the workflow.

You did not execute that command against the environment.

The controller read Git and reconciled the cluster.


5. Drift: Change the Cluster Behind Git's Back [DEV] [OPS]

Now deliberately violate the model.

Git says:

replicas: 2

Change the live Deployment:

kubectl scale deployment/web \
  -n web-dev \
  --replicas=7

Check:

kubectl get deployment web \
  -n web-dev

We now have:

Git desired state:       2 replicas
cluster actual state:    7 replicas

Ask Argo:

kubectl get application web-dev \
  -n argocd \
  -o jsonpath='{.status.sync.status}{"\n"}'

After reconciliation catches up, it should report:

OutOfSync

But notice that Argo has not necessarily repaired it.

Automated synchronization of new Git revisions and automatic repair of live drift are separate choices.

Enable self-healing

Patch the Application:

kubectl patch application web-dev \
  -n argocd \
  --type merge \
  -p '{
    "spec": {
      "syncPolicy": {
        "automated": {
          "selfHeal": true
        },
        "syncOptions": ["CreateNamespace=true"]
      }
    }
  }'

Watch the Deployment:

kubectl get deployment web \
  -n web-dev \
  -w

It should return to:

2 replicas

We just created the higher-level equivalent of deleting a Pod from a Deployment.

The important hierarchy is now:

Git says 2
   |
   v
Argo CD restores Deployment.spec.replicas=2
   |
   v
Deployment controller restores two Pods

Two controllers are reconciling two different layers of desired state.


6. Pruning: What Happens When Git Deletes Something? [DEV] [OPS]

Reconciliation must answer two questions:

What should exist?

and

What should no longer exist?

Create a ConfigMap in Git:

cd ~/gitops-lab/platform-gitops

cat > apps/web/base/banner.yaml <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: web-banner
data:
  message: hello-from-git
EOF

Add it to the base Kustomization:

python3 - <<'PY'
from pathlib import Path
p = Path("apps/web/base/kustomization.yaml")
s = p.read_text()
if "  - banner.yaml\n" not in s:
    s = s.replace("  - service.yaml\n", "  - service.yaml\n  - banner.yaml\n")
p.write_text(s)
PY

Commit and push:

git add .
git commit -m "add banner config"
git push

Wait for Argo:

kubectl get configmap web-banner \
  -n web-dev \
  -w

Now remove the object from Git:

rm apps/web/base/banner.yaml

python3 - <<'PY'
from pathlib import Path
p = Path("apps/web/base/kustomization.yaml")
s = p.read_text().replace("  - banner.yaml\n", "")
p.write_text(s)
PY

git add .
git commit -m "remove banner config"
git push

Inspect the cluster:

kubectl get configmap web-banner \
  -n web-dev

If automated pruning is disabled, the object may remain even though Git no longer declares it.

The Application becomes out of sync.

Enable pruning

Patch the Application:

kubectl patch application web-dev \
  -n argocd \
  --type merge \
  -p '{
    "spec": {
      "syncPolicy": {
        "automated": {
          "prune": true,
          "selfHeal": true
        },
        "syncOptions": ["CreateNamespace=true"]
      }
    }
  }'

Argo should now remove resources that were previously managed but no longer exist in desired state.

Check:

kubectl get configmap web-banner \
  -n web-dev

Expected:

NotFound

Pruning is powerful.

In production, think carefully before allowing automatic pruning of high-impact resources such as namespaces, storage resources, or shared infrastructure.

Argo also supports requiring explicit confirmation before certain prune/delete operations.


Part II - Build Once, Promote an Immutable Artifact

7. The Application Repository [DEV]

So far our GitOps repository deploys public nginx.

Now we will build our own application.

Create a separate repository:

mkdir -p ~/gitops-lab/web-app
cd ~/gitops-lab/web-app

git init -b main

Create a page:

cat > index.html <<'EOF'
<!doctype html>
<html>
  <body>
    <h1>GitOps demo</h1>
    <p>version: v1</p>
  </body>
</html>
EOF

Create the image:

cat > Dockerfile <<'EOF'
FROM nginx:1.27-alpine
COPY index.html /usr/share/nginx/html/index.html
EOF

Build it locally:

docker build -t web-app:dev .

Run it:

docker run --rm \
  -p 8081:80 \
  web-app:dev

From another terminal:

curl http://127.0.0.1:8081

We now have application source that can become an immutable OCI artifact.

Commit it:

git add .
git commit -m "initial web application"

Create and push the hosted repository if you want to follow the CI lab:

gh repo create web-app \
  --public \
  --source=. \
  --remote=origin \
  --push

8. Tags Are Labels; Digests Identify Bytes [DEV] [OPS]

A container tag is a convenient name:

web:v1
web:sha-2f41c2b
web:stable

A registry may move a tag so that it points at a different artifact.

A digest identifies content:

sha256:9d47...

Kubernetes can deploy either:

image: registry.example.com/team/web:v1

or:

image: registry.example.com/team/web@sha256:9d47...

For a promotion system, the second form is much stronger.

The Git commit then says:

Run these exact image bytes.

not:

Ask the registry what v1 means today.

Recommended pattern

Build one artifact.

Attach useful tags for humans:

sha-<git commit>
v1.4.3
release-2026-09-13

But promote the immutable digest:

registry.example.com/team/web@sha256:...

That gives us both:

human discoverability
       +
immutable deployment identity

Avoid production deployments using:

:latest

because it intentionally hides which artifact is meant.


9. Registry Choices: GHCR for the Lab, Harbor for a Platform [DEV] [PLATFORM]

GitOps does not require a particular OCI registry.

For the easiest hosted lab, GitHub Container Registry gives us CI and registry authentication from the same GitHub account.

A platform often uses something such as Harbor because it adds features useful to operators:

projects
robot accounts
retention
immutability
vulnerability scanning
SBOM generation
signatures / content trust
replication

The delivery architecture remains identical:

CI build
  |
  v
OCI registry
  |
  v
immutable digest
  |
  v
GitOps promotion PR

Harbor production recipe

A sensible Harbor project might be:

apps

Create a project-scoped robot account for CI with only the permissions needed to push artifacts.

Your CI credentials then conceptually become:

HARBOR_USERNAME=robot$web-ci
HARBOR_PASSWORD=...

Authenticate:

echo "$HARBOR_PASSWORD" \
  | docker login harbor.example.com \
      --username "$HARBOR_USERNAME" \
      --password-stdin

Push:

docker tag web-app:dev \
  harbor.example.com/apps/web:sha-$(git rev-parse --short HEAD)

docker push \
  harbor.example.com/apps/web:sha-$(git rev-parse --short HEAD)

Then resolve the pushed digest and promote that digest rather than the tag.

Registry hardening checklist

For a real Harbor project, consider enabling:

immutable release tags
vulnerability scanning
SBOM generation
retention policy
Cosign or Notation signatures
content-trust enforcement
short-lived / scoped robot credentials

A signature stored beside an image is useful evidence.

A policy that verifies signatures before deployment is what turns that evidence into enforcement.


10. CI Builds the Artifact - It Does Not Deploy It [DEV] [OPS]

Now we automate the application repository.

The CI job will:

checkout source
      |
      v
build image
      |
      v
push human-readable tag
      |
      v
capture immutable digest
      |
      v
sign / attest artifact
      |
      v
open GitOps promotion PR

Notice what is missing:

kubectl
cluster kubeconfig
cluster-admin

That is the point.

GitHub + GHCR example

Create:

mkdir -p .github/workflows

Create the build workflow:

cat > .github/workflows/build.yaml <<'EOF'
name: build

on:
  push:
    branches:
      - main

permissions:
  contents: read
  packages: write
  id-token: write

jobs:
  image:
    runs-on: ubuntu-latest

    steps:
      - name: Checkout
        uses: actions/checkout@v6

      - name: Log in to GHCR
        uses: docker/login-action@v3
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}

      - name: Set up Buildx
        uses: docker/setup-buildx-action@v3

      - name: Build and push
        id: build
        uses: docker/build-push-action@v6
        with:
          context: .
          push: true
          tags: ghcr.io/${{ github.repository }}:sha-${{ github.sha }}

      - name: Record immutable reference
        run: |
          echo "image=ghcr.io/${GITHUB_REPOSITORY}" >> "$GITHUB_STEP_SUMMARY"
          echo "digest=${{ steps.build.outputs.digest }}" >> "$GITHUB_STEP_SUMMARY"
EOF

Commit and push:

git add .github/workflows/build.yaml
git commit -m "build and publish OCI image"
git push

Make sure the cluster can pull the image

There is a useful registry boundary hiding here.

A newly-published GHCR package is private by default.

Your GitHub Actions runner can push it because the workflow has package permission.

Your kind node has no such credential.

If you merge the promotion without dealing with that boundary, you will rediscover:

ImagePullBackOff

For the shortest public lab, open the package settings in GitHub and change the container package visibility to Public. Public GHCR container packages can be pulled anonymously.

For a private-registry lab, leave it private and create a read credential. For example, using a token that can read the package:

export GHCR_USER="YOUR_GITHUB_USERNAME"
export GHCR_TOKEN="YOUR_READ_PACKAGES_TOKEN"

kubectl create secret docker-registry ghcr-creds \
  -n web-dev \
  --docker-server=ghcr.io \
  --docker-username="$GHCR_USER" \
  --docker-password="$GHCR_TOKEN"

Attach it to the namespace's default ServiceAccount for this lab:

kubectl patch serviceaccount default \
  -n web-dev \
  --type merge \
  -p '{"imagePullSecrets":[{"name":"ghcr-creds"}]}'

Now Pods using that ServiceAccount can present the registry credential when pulling.

For production, do not manually sprinkle long-lived developer tokens through namespaces. Use a deliberate registry-auth pattern such as scoped robot/service credentials, a secret controller, or a node credential provider where appropriate.

This boundary is the same lesson as the kind image sidequest from the main cookbook:

image exists somewhere
       !=
node is authorised and able to obtain it

The action output from docker/build-push-action includes the registry digest.

That digest is the value we actually want to promote.

Why not rebuild per environment?

Do not do:

merge
  |
  +--> build dev image
  +--> rebuild staging image
  +--> rebuild prod image

Each build can produce different bytes.

Instead:

source commit
     |
     v
ONE image digest
     |
     +--> dev
     |
     +--> staging
     |
     +--> production

Promotion means changing where the already-built artifact is allowed to run.

It should not mean recompiling it.


11. Sign and Attest the Artifact [OPS] [PLATFORM]

A digest answers:

Which bytes?

A signature or provenance attestation helps answer:

Who or what produced these bytes, and under what build identity?

For GitHub CI, Sigstore/Cosign can use GitHub's OIDC identity so the workflow does not need a long-lived signing key.

Add Cosign:

      - name: Install Cosign
        uses: sigstore/cosign-installer@v4

      - name: Sign image
        env:
          IMAGE: ghcr.io/${{ github.repository }}
          DIGEST: ${{ steps.build.outputs.digest }}
        run: |
          cosign sign --yes "${IMAGE}@${DIGEST}"

The workflow needs:

permissions:
  id-token: write

which we already granted.

GitHub also supports artifact provenance attestations using the image digest output.

The exact product choice is less important than the model:

source identity
      |
      v
build system identity
      |
      v
artifact digest
      |
      v
signature / provenance

Verify locally

Install Cosign if needed:

brew install cosign

Then verification takes the general form:

cosign verify \
  ghcr.io/YOUR_USER/web-app@sha256:... \
  --certificate-identity-regexp='^https://github.com/' \
  --certificate-oidc-issuer=https://token.actions.githubusercontent.com

For production, make the expected certificate identity specific to the repository and workflow rather than using a broad regular expression.

Harbor

Harbor can store Cosign signatures as OCI-related artifacts associated with the signed image.

It can also enforce project content-trust policies so unsigned artifacts cannot be pulled.

That makes the chain stronger:

CI signs image
      |
      v
Harbor stores image + signature
      |
      v
Harbor policy rejects unsigned pull

Signing without verification is documentation.

Signing plus enforcement is a security control.


12. Turn the Image Digest into a Promotion Pull Request [DEV] [OPS]

Now we connect the application repository to the GitOps repository.

The application CI should not merge directly into production desired state.

It should propose a change.

Conceptually:

new image digest
      |
      v
branch in GitOps repo
      |
      v
pull request
      |
      +-- CI checks
      +-- policy checks
      +-- human review
      |
      v
merge
      |
      v
Argo reconciliation

What should the PR change?

Our development Kustomization currently contains:

images:
  - name: example/web
    newName: nginx
    newTag: 1.27-alpine

A real promotion should result in something like:

images:
  - name: example/web
    newName: ghcr.io/example/web-app
    digest: sha256:0123456789abcdef...

That diff is boring.

Boring is good.

The deployment decision becomes obvious in code review.

Authentication for the GitOps repository

The build repository needs permission to open a PR in the GitOps repository.

For a lab, you can use a fine-grained token stored as:

GITOPS_TOKEN

Give it only the target GitOps repository permissions it needs for:

contents: write
pull requests: write

For a serious platform, prefer a GitHub App or equivalent workload identity over a developer's personal token.

The identity should represent:

promotion automation

not:

Alice's laptop credential

Promotion job

Add this after the image build/sign steps:

      - name: Propose development promotion
        env:
          GH_TOKEN: ${{ secrets.GITOPS_TOKEN }}
          GITOPS_REPOSITORY: YOUR_USER/platform-gitops
          IMAGE: ghcr.io/${{ github.repository }}
          DIGEST: ${{ steps.build.outputs.digest }}
          SOURCE_SHA: ${{ github.sha }}
        run: |
          set -euo pipefail

          gh auth setup-git
          gh repo clone "$GITOPS_REPOSITORY" gitops
          cd gitops

          branch="promote/web-${SOURCE_SHA:0:12}"
          git checkout -b "$branch"

          python3 - "$IMAGE" "$DIGEST" <<'PY'
          from pathlib import Path
          import sys

          image = sys.argv[1]
          digest = sys.argv[2]
          p = Path("environments/dev/web/kustomization.yaml")
          text = p.read_text()

          start = text.index("images:\n")
          replacement = f"""images:\n  - name: example/web\n    newName: {image}\n    digest: {digest}\n"""
          text = text[:start] + replacement
          p.write_text(text)
          PY

          git config user.name "web promotion bot"
          git config user.email "web-promotion-bot@users.noreply.github.com"

          git add environments/dev/web/kustomization.yaml
          git commit -m "promote web ${SOURCE_SHA:0:12} to dev"
          git push --set-upstream origin "$branch"

          gh pr create \
            --repo "$GITOPS_REPOSITORY" \
            --base main \
            --head "$branch" \
            --title "Promote web ${SOURCE_SHA:0:12} to dev" \
            --body "Image: ${IMAGE}@${DIGEST}

Source commit: ${SOURCE_SHA}"

Replace:

YOUR_USER/platform-gitops

with your GitOps repository.

After your next application merge, CI should build an image and open a GitOps PR.

That PR is the deployment proposal.


13. Review the Promotion Like a Deployment [DEV] [OPS]

Open the generated GitOps PR.

The important diff should effectively be:

 images:
   - name: example/web
-    newName: nginx
-    newTag: 1.27-alpine
+    newName: ghcr.io/example/web-app
+    digest: sha256:...

Before merging, ask:

Did CI pass?
Is the image signed?
Did vulnerability policy pass?
Does the digest correspond to the intended source commit?
Is this environment allowed to consume it?

Merge the PR.

Then watch Argo rather than running kubectl apply:

kubectl get application web-dev \
  -n argocd \
  -w

Watch Kubernetes underneath it:

kubectl get deployment,rs,pods \
  -n web-dev \
  -w

You should recognise the same rollout machinery from the Kubernetes cookbook.

Git changed.

Argo changed the Deployment desired state.

The Deployment controller changed ReplicaSets and Pods.

The hierarchy is:

Git commit
    |
    v
Argo Application
    |
    v
Deployment.spec.template
    |
    v
ReplicaSet
    |
    v
Pods

Prove which image is running

Ask Kubernetes:

kubectl get deployment web \
  -n web-dev \
  -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'

You should see the digest-pinned image reference.

Now Git history and cluster state can be connected directly.


14. Why the Pull Request Is Part of the Control Plane [PLATFORM]

It is tempting to think of the PR as bureaucracy around the "real deployment".

In this model, the PR is part of the deployment control plane.

The pull request is where we can perform controls before desired state changes:

code review
policy checks
security scan result
change ticket reference
ownership approval
release notes
blast-radius review

A production branch might require:

2 reviewers
CODEOWNERS approval
successful policy checks
signed commits
no direct pushes

The deployment itself remains automatic after the desired-state decision is accepted.

That separation is important:

human decides what should happen
        |
        v
controller performs it consistently

Humans should not need to manually reproduce deployment mechanics on every release.


Part III - Environments Are Promotion Boundaries

15. Dev, Staging and Production [DEV] [PLATFORM]

Extend the GitOps repository:

environments/
├── dev/
│   └── web/
├── staging/
│   └── web/
└── prod/
    └── web/

The critical idea is:

Promote the same digest through environments.

Suppose CI built:

ghcr.io/example/web@sha256:abc123

Development:

digest: sha256:abc123

After validation, staging should also use:

digest: sha256:abc123

Production should eventually use:

digest: sha256:abc123

Do not do:

dev     -> sha256:abc123
staging -> rebuild -> sha256:def456
prod    -> rebuild -> sha256:987xyz

You no longer know whether production contains what was tested.

Promotion PRs

A typical flow becomes:

application merge
      |
      v
CI creates immutable image
      |
      v
PR: digest -> dev
      |
      v
dev verification
      |
      v
PR: same digest -> staging
      |
      v
staging verification
      |
      v
PR: same digest -> production

A platform can automate the creation of those PRs while still preserving review gates.


16. Separate Application Change from Environment Promotion [PLATFORM]

This is one of the most useful organisational consequences of the two-repository model.

Application developers can own:

source
unit tests
Dockerfile
application CI

A platform or service team can own:

production rollout policy
replica counts
resources
network exposure
secret references
availability constraints

The image digest is the handshake between them.

application team:
"I produced artifact abc123"

platform/environment:
"production is approved to run abc123"

This does not require a central platform team to approve every deployment.

Repository ownership and branch rules can express whatever autonomy model the organisation chooses.

The point is that responsibilities become explicit.


17. AppProjects: Restrict What Argo May Deploy [OPS] [PLATFORM]

So far our Application uses:

project: default

The default Argo project is intentionally permissive and useful for getting started.

It is not a good final tenancy model.

Create a dedicated project:

cat > /tmp/web-project.yaml <<EOF
apiVersion: argoproj.io/v1alpha1
kind: AppProject
metadata:
  name: web-team
  namespace: argocd
spec:
  description: Web team applications
  sourceRepos:
    - ${GITOPS_REPO}
  destinations:
    - namespace: web-*
      server: https://kubernetes.default.svc
EOF

Apply it:

kubectl apply -f /tmp/web-project.yaml

Move the Application into it:

kubectl patch application web-dev \
  -n argocd \
  --type merge \
  -p '{"spec":{"project":"web-team"}}'

Inspect:

kubectl get appproject web-team \
  -n argocd \
  -o yaml

Now Argo has two layers of authority:

Kubernetes RBAC
      |
      v
what Argo's ServiceAccount can do

AppProject policy
      |
      v
what applications in this project are permitted to request

That distinction matters.

An AppProject can constrain:

trusted source repositories
target clusters
target namespaces
resource kinds
Argo user/team roles

A dangerous boundary

Do not casually permit application projects to deploy into the argocd namespace itself.

An application able to modify Argo's own configuration or credentials can potentially turn application deployment permission into platform administration.

This is the same privilege-escalation reasoning we used in the Kubernetes RBAC chapter.


18. Argo's Kubernetes Permissions Still Matter [OPS] [PLATFORM]

GitOps does not abolish RBAC.

It moves the actor.

With imperative deployment:

developer / CI identity
        |
        v
Kubernetes RBAC

With Argo:

developer
   |
   v
Git permissions
   |
   v
Argo controller identity
   |
   v
Kubernetes RBAC

Ask which ServiceAccounts exist:

kubectl get serviceaccounts \
  -n argocd

Inspect the controller identity:

kubectl get pod \
  -n argocd \
  -l app.kubernetes.io/name=argocd-application-controller \
  -o jsonpath='{.items[0].spec.serviceAccountName}{"\n"}'

Then inspect bindings involving Argo:

kubectl get clusterrolebinding \
  -o yaml \
  | grep -n -C 3 argocd

The standard lab installation is deliberately powerful.

That is convenient for one local cluster.

A production platform should answer explicitly:

Which clusters may this Argo instance manage?
Which namespaces?
Which cluster-scoped resources?
Which teams may change its sources and destinations?

Never assume "GitOps" means "secure" by default.

It gives us a better control model; we still have to configure that model responsibly.


19. Secrets Do Not Magically Become Safe Because They Are in Git [DEV] [OPS]

Do not commit this to a normal Git repository:

apiVersion: v1
kind: Secret
metadata:
  name: database
data:
  password: ...

Base64 is encoding, not encryption.

GitOps needs a separate secret strategy.

Common patterns include:

External Secrets Operator
   Git stores a reference
   secret value lives in Vault / cloud secret manager

SOPS
   encrypted secret material lives in Git
   authorised controller decrypts it

Sealed Secrets
   encrypted object lives in Git
   cluster controller decrypts it

The core design question is:

Can Git contain enough desired state to reference the secret without exposing the secret itself?

A production deployment might contain:

Git:
  "web needs database credential named web-db"

secret manager:
  actual credential bytes

controller:
  materialises Kubernetes Secret

This is another reconciliation loop.

Modern platforms are often a composition of controllers, each owning one slice of desired state.


Part IV - Progressive Delivery

20. Why a Successful Git Sync Is Not the Same as a Safe Release [DEV] [OPS]

Argo CD answers:

Does the cluster match Git?

That does not necessarily answer:

Is this new version safe for 100% of users?

A normal Deployment may replace instances gradually, but it does not inherently evaluate business or reliability signals before continuing.

Progressive delivery adds another control loop:

new desired image
      |
      v
small amount of exposure
      |
      v
observe health
      |
  +---+---+
  |       |
good      bad
  |       |
  v       v
more    abort
traffic

We will use Argo Rollouts for this lab.

Argo CD and Argo Rollouts solve different problems:

Argo CD
  Git -> Kubernetes desired state

Argo Rollouts
  desired release -> controlled transition between versions

They compose well because both are controllers.


21. Install Argo Rollouts [DEV] [OPS]

Create the namespace:

kubectl create namespace argo-rollouts

Install the controller:

kubectl apply \
  -n argo-rollouts \
  -f https://github.com/argoproj/argo-rollouts/releases/latest/download/install.yaml

Wait:

kubectl rollout status \
  deployment/argo-rollouts \
  -n argo-rollouts \
  --timeout=300s

On macOS, install the kubectl plugin:

brew install argoproj/tap/kubectl-argo-rollouts

Check:

kubectl argo rollouts version

Find the new APIs:

kubectl api-resources \
  | grep argoproj

You should now see resources including:

Rollout
AnalysisTemplate
AnalysisRun
Experiment

Again, a product feature has become Kubernetes API objects plus controllers.


22. Migrate a Deployment to Rollouts Without Controllers Fighting [DEV] [OPS]

A tempting migration is:

change kind: Deployment
          to Rollout

For a throwaway manifest that can be enough.

For a running workload, it hides an important ownership problem.

If a Deployment and Rollout temporarily exist with overlapping selectors, both controllers can be active at the same time.

And if Argo CD keeps enforcing a Deployment replica count while Argo Rollouts is trying to scale that Deployment down during migration, two reconcilers can fight over the same field.

That gives us a better lab.

We will migrate using a Rollout workloadRef.

The existing Deployment remains the source of the Pod template.

The Rollout becomes responsible for progressive delivery.

Conceptually:

Git / Kustomize
      |
      v
Deployment Pod template
      |
      | workloadRef
      v
Rollout
      |
      +-- manages progressive ReplicaSets
      |
      +-- scales old Deployment down

First decide who owns Deployment.spec.replicas

Our Argo Application currently has self-healing enabled.

Git says the original Deployment has:

replicas: 2

But during migration, Argo Rollouts needs to scale that Deployment down.

If both controllers insist on owning the same field, we can create a reconciliation tug-of-war:

Git / Argo CD:       replicas must be 2
Argo Rollouts:       replicas must be 0

The fix is not to disable reconciliation globally.

The fix is to define field ownership deliberately.

Tell Argo CD to ignore the replica field of this specific Deployment, and to respect that exclusion during sync:

kubectl patch application web-dev \
  -n argocd \
  --type merge \
  -p '{
    "spec": {
      "ignoreDifferences": [
        {
          "group": "apps",
          "kind": "Deployment",
          "name": "web",
          "namespace": "web-dev",
          "jsonPointers": ["/spec/replicas"]
        }
      ],
      "syncPolicy": {
        "automated": {
          "prune": true,
          "selfHeal": true
        },
        "syncOptions": [
          "CreateNamespace=true",
          "RespectIgnoreDifferences=true"
        ]
      }
    }
  }'

Check the relevant part:

kubectl get application web-dev \
  -n argocd \
  -o yaml \
  | grep -A20 ignoreDifferences

This is an important general GitOps pattern:

If another controller legitimately owns a field, do not make your GitOps controller continuously overwrite it.

HPAs, progressive-delivery controllers, service meshes and other operators can all create this kind of shared ownership.

Add the Rollout

In the GitOps repository:

cd ~/gitops-lab/platform-gitops

Keep the existing Deployment.

Create a Rollout that references it:

cat > apps/web/base/rollout.yaml <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: web
spec:
  replicas: 5
  selector:
    matchLabels:
      app: web
  workloadRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web
    scaleDown: progressively
  strategy:
    canary:
      steps:
        - setWeight: 20
        - pause: {}
        - setWeight: 50
        - pause:
            duration: 30s
        - setWeight: 100
EOF

Add it to the base Kustomization:

python3 - <<'PY2'
from pathlib import Path
p = Path("apps/web/base/kustomization.yaml")
s = p.read_text()
if "  - rollout.yaml\n" not in s:
    s = s.replace("resources:\n", "resources:\n  - rollout.yaml\n")
p.write_text(s)
PY2

Do not remove the Deployment.

The Rollout is using its Pod template through workloadRef.

Render the desired state:

kubectl kustomize environments/dev/web

You should see both:

Deployment/web
Rollout/web

That is intentional during this migration pattern.

Commit and push:

git add .
git commit -m "migrate web delivery to Argo Rollouts"
git push

Watch the two controllers:

kubectl get deployment,rollout,rs,pods \
  -n web-dev \
  -w

Inspect the Rollout:

kubectl argo rollouts get rollout web \
  -n web-dev

The Rollout should establish its own stable ReplicaSet while progressively scaling down the old Deployment-managed workload.

Inspect the original Deployment:

kubectl get deployment web \
  -n web-dev

Its replica count can now be changed by the Rollouts controller without Argo CD immediately restoring the Git value.

Where does the image live now?

Because we used workloadRef, the Pod template still lives on the Deployment:

kubectl get deployment web \
  -n web-dev \
  -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'

Our existing Kustomize image override therefore continues to work exactly where it did before.

When a promotion PR changes the Deployment Pod template image, Argo Rollouts observes the referenced workload change and performs the canary transition.

This gives us a clean division of ownership:

Git / Argo CD
  owns Deployment Pod template
  owns Rollout policy

Argo Rollouts
  owns progressive ReplicaSets
  owns migration scaling

Kubernetes
  owns Pod execution and status

The migration itself has taught us something broader than Argo Rollouts:

Multiple controllers can cooperate safely only when we understand which fields and resources each controller owns.


23. Ship v2 as a Canary [DEV] [OPS]

Change the application page:

cd ~/gitops-lab/web-app

cat > index.html <<'EOF'
<!doctype html>
<html>
  <body>
    <h1>GitOps demo</h1>
    <p>version: v2</p>
  </body>
</html>
EOF

Commit and push:

git add index.html
git commit -m "release v2"
git push

Your CI should:

build v2
push it
capture digest
open GitOps PR

Merge the promotion PR.

Now watch Argo Rollouts:

kubectl argo rollouts get rollout web \
  -n web-dev \
  --watch

The rollout should reach:

20% canary
PAUSED

With five replicas and no dedicated traffic router, the weight is represented approximately by replica count.

Inspect Pods:

kubectl get pods \
  -n web-dev \
  -l app=web \
  -o wide

You should see old and new ReplicaSets coexisting.

This is the same mechanism you learned for Deployments, but now the transition has explicit programmable steps.

Send traffic

Port-forward the Service:

kubectl port-forward \
  -n web-dev \
  service/web \
  8082:80

From another terminal, make several requests:

for i in $(seq 1 20); do
  curl -s http://127.0.0.1:8082 \
    | grep version
  sleep 0.2
done

Depending on connection reuse and Service load balancing you should observe both versions over repeated independent requests.

For exact 5%, 10%, header-based, or mirrored traffic, we need a traffic-management layer rather than relying only on replica ratios.

We will return to that later.


24. Promote the Canary [DEV] [OPS]

The rollout is paused because Git describes the release policy as well as the image.

A human can inspect it:

kubectl argo rollouts get rollout web \
  -n web-dev

Promote it:

kubectl argo rollouts promote web \
  -n web-dev

Watch:

kubectl argo rollouts get rollout web \
  -n web-dev \
  --watch

It should move through the remaining steps until v2 becomes stable.

Notice the separation of decisions:

GitOps PR:
"this is the desired release"

Rollout policy:
"this is how we expose that release"

promotion:
"the observed canary looks acceptable"

Those are three different concerns.


25. Ship a Bad Release and Abort It [DEV] [OPS]

Now make v3 deliberately obvious and undesirable.

cd ~/gitops-lab/web-app

cat > index.html <<'EOF'
<!doctype html>
<html>
  <body>
    <h1>GitOps demo</h1>
    <p>version: v3-BAD</p>
  </body>
</html>
EOF

git add index.html
git commit -m "ship intentionally bad v3"
git push

Let CI create the promotion PR and merge it.

Watch until the canary pauses:

kubectl argo rollouts get rollout web \
  -n web-dev \
  --watch

Suppose testing or metrics say the version is bad.

Abort the rollout immediately:

kubectl argo rollouts abort web \
  -n web-dev

Inspect:

kubectl argo rollouts get rollout web \
  -n web-dev

The stable ReplicaSet should be restored to serve the workload.

But we have a subtle problem.

Git still says:

v3 digest is desired

Argo Rollouts says:

v3 rollout is aborted
stable v2 is serving

The emergency action protected users.

It did not change the source of truth.

This distinction is fundamental.


26. Operational Rollback vs Git Rollback [OPS] [PLATFORM]

We need two different concepts:

ABORT
  stop exposing a bad release now

REVERT
  change desired state back to the known-good artifact

The abort is an operational safety action.

The Git revert makes the desired state honest again.

Revert the promotion

In the GitOps repository, identify the merge commit that promoted v3:

cd ~/gitops-lab/platform-gitops

git log --oneline --decorate -10

Create a revert branch:

git checkout main
git pull

git checkout -b rollback/web-v3

git revert <PROMOTION_MERGE_COMMIT>

Push it:

git push -u origin rollback/web-v3

Open a pull request:

gh pr create \
  --title "Rollback web v3" \
  --body "Restore the previous known-good image digest after canary abort."

Merge it.

Argo CD sees Git return to the previous digest.

Argo Rollouts now has desired state that agrees with the stable version.

The system converges again.

Why not kubectl set image to v2?

Because that would create another hidden live mutation:

cluster says v2
Git says v3

Self-healing GitOps would eventually try to restore v3.

Emergency commands are fine when the system is burning.

But the source of truth must subsequently be corrected.


27. Automate the Decision with Analysis [OPS] [PLATFORM]

Manual promotion is useful while learning.

A mature delivery system usually evaluates signals automatically.

Examples:

HTTP error rate
latency
saturation
queue depth
business conversion
synthetic checks
smoke-test webhook

Argo Rollouts represents this using:

AnalysisTemplate
      |
      v
AnalysisRun
      |
      v
measurement
      |
  +---+---+
  |       |
success failure
  |       |
  v       v
continue abort

A tiny runnable analysis gate

For the lab, create a static JSON health endpoint.

This is deliberately simpler than Prometheus so we can see the mechanism first.

Create:

cat > /tmp/analysis-gate.yaml <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: analysis-gate
  namespace: web-dev
data:
  result.json: |
    {"ok": true}
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: analysis-gate
  namespace: web-dev
spec:
  replicas: 1
  selector:
    matchLabels:
      app: analysis-gate
  template:
    metadata:
      labels:
        app: analysis-gate
    spec:
      containers:
        - name: nginx
          image: nginx:1.27-alpine
          volumeMounts:
            - name: data
              mountPath: /usr/share/nginx/html
      volumes:
        - name: data
          configMap:
            name: analysis-gate
---
apiVersion: v1
kind: Service
metadata:
  name: analysis-gate
  namespace: web-dev
spec:
  selector:
    app: analysis-gate
  ports:
    - port: 80
      targetPort: 80
EOF

kubectl apply -f /tmp/analysis-gate.yaml

Test it:

kubectl run curl-test \
  -n web-dev \
  --rm -i --restart=Never \
  --image=curlimages/curl \
  -- \
  curl -s http://analysis-gate/result.json

Expected:

{"ok": true}

Create an AnalysisTemplate in the GitOps repository:

cd ~/gitops-lab/platform-gitops

cat > apps/web/base/analysis.yaml <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: web-health
spec:
  metrics:
    - name: release-gate
      successCondition: result == true
      failureLimit: 1
      provider:
        web:
          url: http://analysis-gate.web-dev.svc.cluster.local/result.json
          jsonPath: "{$.ok}"
EOF

Add it to Kustomize:

python3 - <<'PY'
from pathlib import Path
p = Path("apps/web/base/kustomization.yaml")
s = p.read_text()
if "  - analysis.yaml\n" not in s:
    s = s.replace("resources:\n", "resources:\n  - analysis.yaml\n")
p.write_text(s)
PY

Now change the Rollout steps so analysis occurs after the 20% canary:

  strategy:
    canary:
      steps:
        - setWeight: 20
        - pause:
            duration: 10s
        - analysis:
            templates:
              - templateName: web-health
        - setWeight: 50
        - pause:
            duration: 20s
        - setWeight: 100

Commit and push the GitOps change.

Then trigger another image promotion.

Inspect generated AnalysisRuns:

kubectl get analysisrun \
  -n web-dev

Describe one:

kubectl describe analysisrun \
  -n web-dev \
  <analysis-run-name>

Make the gate fail

Change the ConfigMap:

kubectl patch configmap analysis-gate \
  -n web-dev \
  --type merge \
  -p '{"data":{"result.json":"{\"ok\": false}\n"}}'

Wait for the projected ConfigMap volume to update, or restart the gate Pod:

kubectl rollout restart deployment/analysis-gate \
  -n web-dev

Trigger another release.

The AnalysisRun should fail and the Rollout should abort rather than progressing.

Production translation

The static gate is only teaching the API.

A real platform would normally query a measurement system such as Prometheus:

success rate >= 99.5%
p95 latency < 300 ms
error rate < 1%

The concept remains identical.


28. Canary Replica Ratios vs Real Traffic Shaping [OPS] [PLATFORM]

Our basic canary uses replica counts.

With five replicas:

1 canary + 4 stable ~= 20%

That is useful but crude.

It cannot express concepts such as:

1% traffic to canary
only users with header X
mirror requests without using canary responses
90/10 traffic while keeping equal replica counts

For that, Argo Rollouts integrates with traffic-management systems.

This is where our Gateway API / Cilium work connects directly.

Conceptually:

Rollout desired weight
       |
       v
Argo Rollouts
       |
       v
Gateway API route weights
       |
       v
Cilium
       |
       v
real network traffic

Modern Argo Rollouts can integrate with Gateway API through its traffic-router plugin system.

That allows the progressive-delivery controller to update Gateway API routing state rather than merely scaling stable/canary replica counts.

For learners following the Cilium Gateway API companion, this is the natural next exercise after mastering the basic Rollout.

The important architectural connection is:

Git
 |
 v
Argo CD
 |
 v
Rollout
 |
 v
Argo Rollouts
 |
 +--> ReplicaSets
 |
 +--> Gateway API route weight
          |
          v
        Cilium

Each controller owns a different concern.

Another controller-ownership trap

A traffic router introduces the same field-ownership issue we saw during Deployment-to-Rollout migration.

If Argo Rollouts dynamically changes route weights while Argo CD insists that the weights must always equal the static values committed in Git, the controllers can report permanent drift or overwrite one another.

The usual GitOps pattern is to keep the route object in Git while explicitly ignoring the fields that the progressive-delivery controller is expected to mutate.

Conceptually:

Git owns:
  route identity
  hostnames
  backends
  rollout policy

Rollouts owns during release:
  dynamic traffic weights

This is not weakening GitOps.

It is defining ownership precisely enough for multiple reconcilers to cooperate.


29. Blue-Green: Build the Replacement Before You Switch Traffic [DEV] [OPS]

Canary is not the only progressive-delivery strategy.

A canary asks:

Can we expose a small amount of real traffic to the new version and increase it gradually?

Blue-green asks a different question:

Can we build the entire replacement, test it separately, and switch production traffic only when we are happy?

The mental model is:

                         +--------------------+
                         | stable ReplicaSet  |
                         |        v3          |
                         +---------+----------+
                                   ^
                                   |
                         active Service
                                   |
                              production


                         +--------------------+
                         | preview ReplicaSet |
                         |        v4          |
                         +---------+----------+
                                   ^
                                   |
                         preview Service
                                   |
                           tests / humans

Both versions exist.

But only one receives production traffic.

When we promote:

before
------

web Service ----------> v3
web-preview Service --> v4


promotion
---------

web Service ----------> v4
web-preview Service --> v4

                         v3 remains briefly
                         available for rollback

This gives blue-green a very useful property:

new version is deployed
        !=
new version is serving production traffic

That distinction is worth experiencing directly.

Canary vs blue-green

The two strategies solve slightly different release problems.

RollingUpdate
  replace instances gradually
  simplest operational model

Canary
  expose some real production traffic
  measure behaviour
  increase exposure gradually

Blue-green
  build a complete replacement
  validate it away from production
  switch traffic when ready

Blue-green is often easier to reason about because there is a hard traffic boundary between the active and preview versions.

It does have a cost.

For at least part of the release we may run two versions simultaneously:

active capacity
+
preview capacity

For an expensive workload that may matter.

Argo Rollouts lets us reduce preview capacity with previewReplicaCount, but the new version must be scaled to the full desired replica count before it becomes active.

When would you choose each strategy?

Blue-green is attractive when:

we can validate a version before production traffic
cutover should happen quickly
old and new versions cannot safely share live traffic
rollback speed matters
extra temporary capacity is acceptable

Canary is attractive when:

real-user behaviour is part of validation
we have useful production metrics
we want gradual blast-radius expansion
our application tolerates multiple versions serving simultaneously
we have a traffic router for precise percentages

Neither is universally better.

A platform can support both.

The release should choose the strategy that matches the application's risk model.


30. Convert Our Rollout from Canary to Blue-Green [DEV] [OPS] [PLATFORM]

We already have:

Git
  |
  v
Argo CD
  |
  v
Rollout/web
  |
  v
ReplicaSets

We are going to keep that delivery chain.

We will change only the progressive-delivery strategy.

The blue-green controller needs two Services:

web
  active production traffic

web-preview
  pre-production validation traffic

Argo Rollouts will dynamically add a ReplicaSet hash to those Service selectors so that each Service points at exactly the intended version.

That creates an important ownership question.

Who owns the Service selector?

Our Service currently lives in Git:

spec:
  selector:
    app: web

For blue-green, Argo Rollouts will turn the live selector into something conceptually like:

spec:
  selector:
    app: web
    rollouts-pod-template-hash: 6d997f5c6

That hash changes as releases change.

If Argo CD insists that the selector must always exactly match the static Git version, we create another reconciliation fight:

Argo Rollouts:
  selector must point at ReplicaSet v4

Argo CD:
  selector must equal Git exactly

We have already seen this class of problem with Deployment.spec.replicas.

The solution is the same:

Define field ownership deliberately.

Git still owns:

Service identity
ports
protocol
application labels

Argo Rollouts owns during blue-green delivery:

the live Service selector used to choose the active/preview ReplicaSet

Give Rollouts ownership of the dynamic selectors

First inspect our current Argo CD exception:

kubectl get application web-dev \
  -n argocd \
  -o yaml \
  | grep -A30 ignoreDifferences

We already ignore the Deployment replica field used during workloadRef migration.

Replace the ignore list with all three controller-owned fields:

kubectl patch application web-dev \
  -n argocd \
  --type merge \
  -p '{
    "spec": {
      "ignoreDifferences": [
        {
          "group": "apps",
          "kind": "Deployment",
          "name": "web",
          "namespace": "web-dev",
          "jsonPointers": ["/spec/replicas"]
        },
        {
          "group": "",
          "kind": "Service",
          "name": "web",
          "namespace": "web-dev",
          "jsonPointers": ["/spec/selector"]
        },
        {
          "group": "",
          "kind": "Service",
          "name": "web-preview",
          "namespace": "web-dev",
          "jsonPointers": ["/spec/selector"]
        }
      ],
      "syncPolicy": {
        "automated": {
          "prune": true,
          "selfHeal": true
        },
        "syncOptions": [
          "CreateNamespace=true",
          "RespectIgnoreDifferences=true"
        ]
      }
    }
  }'

Verify:

kubectl get application web-dev \
  -n argocd \
  -o yaml \
  | grep -A45 ignoreDifferences

The important lesson is not the exact Argo syntax.

It is this:

multiple controllers
        |
        v
must have a coherent ownership model

Restore our analysis gate

The previous chapter deliberately made the analysis endpoint fail.

Set it back to healthy before the next experiment:

kubectl patch configmap analysis-gate \
  -n web-dev \
  --type merge \
  -p '{"data":{"result.json":"{\"ok\": true}\n"}}'

Restart the tiny gate workload so there is no ambiguity about what it serves:

kubectl rollout restart deployment/analysis-gate \
  -n web-dev

Check it:

kubectl run analysis-check \
  -n web-dev \
  --rm -i --restart=Never \
  --image=curlimages/curl \
  -- \
  curl -s http://analysis-gate/result.json

Expected:

{"ok": true}

Add the preview Service to Git

Move to the GitOps repository:

cd ~/gitops-lab/platform-gitops

git checkout main
git pull

git checkout -b platform/blue-green

Create the preview Service:

cat > apps/web/base/web-preview.yaml <<'EOF'
apiVersion: v1
kind: Service
metadata:
  name: web-preview
spec:
  selector:
    app: web
  ports:
    - port: 80
      targetPort: 80
EOF

Add it to the base Kustomization:

python3 - <<'PY'
from pathlib import Path
p = Path("apps/web/base/kustomization.yaml")
s = p.read_text()
if "  - web-preview.yaml\n" not in s:
    s = s.replace("resources:\n", "resources:\n  - web-preview.yaml\n")
p.write_text(s)
PY

Change the Rollout strategy

Replace the canary policy with blue-green:

cat > apps/web/base/rollout.yaml <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: web
spec:
  replicas: 5
  revisionHistoryLimit: 5
  rollbackWindow:
    revisions: 3
  selector:
    matchLabels:
      app: web
  workloadRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web
    scaleDown: progressively
  strategy:
    blueGreen:
      activeService: web
      previewService: web-preview
      previewReplicaCount: 2
      autoPromotionEnabled: false
      scaleDownDelaySeconds: 60
      prePromotionAnalysis:
        templates:
          - templateName: web-health
EOF

There are several important controls here.

activeService

activeService: web

This is the production Service.

Argo Rollouts controls which ReplicaSet it selects.

previewService

previewService: web-preview

This gives us an endpoint for the new version before promotion.

previewReplicaCount

previewReplicaCount: 2

Our Rollout wants five replicas in steady state.

During preview we only need two copies to validate the version.

This saves some temporary capacity:

active:   5 replicas
preview:  2 replicas

Before promotion, Rollouts scales the preview version to the required active size.

autoPromotionEnabled

autoPromotionEnabled: false

Even after the preview is healthy, production does not switch automatically.

The rollout pauses until we explicitly promote it.

prePromotionAnalysis

prePromotionAnalysis:
  templates:
    - templateName: web-health

Our existing AnalysisTemplate runs before the active Service switches.

That gives us:

new version ready
      |
      v
analysis
      |
  +---+---+
  |       |
pass     fail
  |       |
  v       v
pause   abort
  |
manual promotion
  |
active Service switches

scaleDownDelaySeconds

scaleDownDelaySeconds: 60

After promotion, Rollouts keeps the old active ReplicaSet around briefly.

There are two reasons this is useful:

network rule propagation
+
rapid rollback window

The default is already conservative, but a minute makes the behaviour easy to observe in our lab.

rollbackWindow

rollbackWindow:
  revisions: 3

Recent revisions can be fast-tracked if Git moves back to them.

That is particularly useful in a GitOps system:

bad release
   |
Git revert
   |
old digest becomes desired again
   |
Rollouts recognises recent revision
   |
fast rollback

Render before committing

Do not use Git as a YAML syntax checker.

Render the overlay locally:

kubectl kustomize environments/dev/web \
  > /tmp/web-blue-green.yaml

Check that it contains:

grep -nE 'kind: Rollout|kind: Service|name: web-preview|blueGreen:' \
  /tmp/web-blue-green.yaml

Ask the API server to validate it without changing the cluster:

kubectl apply \
  --dry-run=server \
  -f /tmp/web-blue-green.yaml \
  >/dev/null

Now commit the platform change:

git add apps/web/base

git commit -m "add blue-green delivery strategy"

git push -u origin platform/blue-green

Create a pull request:

gh pr create \
  --title "Use blue-green delivery for web" \
  --body "Add a preview Service and move web from canary steps to an Argo Rollouts blue-green strategy."

Review it.

Then merge it:

gh pr merge \
  --merge \
  --delete-branch

Return to main:

git checkout main
git pull

Watch Argo CD reconcile:

kubectl get application web-dev \
  -n argocd \
  -w

In another terminal inspect the Rollout:

kubectl argo rollouts get rollout web \
  -n web-dev

And the Services:

kubectl get service web web-preview \
  -n web-dev

At steady state, both Services may initially point at the current stable ReplicaSet.

The interesting behaviour appears on the next release.


31. Ship a Blue-Green Release: Preview First, Production Later [DEV] [OPS]

Now we will ship a version whose lifecycle is deliberately visible.

The application change still begins in the application repository.

That is important.

We do not edit the GitOps repository by hand to invent a release.

The application CI builds an immutable artifact and proposes its digest for promotion.

Create the next application version

Move to the application repository:

cd ~/gitops-lab/web-app

git checkout main
git pull

Change the page:

cat > index.html <<'EOF'
<!doctype html>
<html>
  <body>
    <h1>GitOps demo</h1>
    <p>version: blue-green-v4</p>
  </body>
</html>
EOF

Commit and push:

git add index.html

git commit -m "release blue-green v4"

git push

The application CI should now perform the chain we built earlier:

commit
  |
  v
build image
  |
  v
push image
  |
  v
resolve immutable digest
  |
  v
open promotion PR against GitOps repository

Inspect the promotion PR:

gh pr list \
  --repo "${GITHUB_USER}/${GITOPS_REPO_NAME}"

Open the diff and verify that the important change is an immutable digest rather than latest:

old digest
    ↓
new digest

Merge the promotion PR.

Watch the preview environment appear

Watch the Rollout:

kubectl argo rollouts get rollout web \
  -n web-dev \
  --watch

The new ReplicaSet should be created and become the preview version.

Because we set:

previewReplicaCount: 2

we expect approximately:

stable ReplicaSet:   5 Pods
preview ReplicaSet:  2 Pods

Inspect them:

kubectl get rs,pods \
  -n web-dev \
  -l app=web

Now look at the Service selectors:

kubectl get service web web-preview \
  -n web-dev \
  -o custom-columns='SERVICE:.metadata.name,HASH:.spec.selector.rollouts-pod-template-hash,APP:.spec.selector.app'

You should see different hashes while the new version is awaiting promotion:

SERVICE       HASH         APP
web           <old-hash>   web
web-preview   <new-hash>   web

This is the blue-green boundary made concrete.

Argo Rollouts has changed routing without changing either Service's identity.

Talk to production

Port-forward the active Service:

kubectl port-forward \
  -n web-dev \
  service/web \
  8082:80

From another terminal:

curl -s http://127.0.0.1:8082 \
  | grep version

You should still see the currently active version.

The new release exists, but production has not moved.

Talk directly to preview

Open another terminal and port-forward the preview Service:

kubectl port-forward \
  -n web-dev \
  service/web-preview \
  8083:80

Now query it:

curl -s http://127.0.0.1:8083 \
  | grep version

Expected:

version: blue-green-v4

We have now personally observed:

active Service  ---> old version
preview Service ---> new version

That is the core of blue-green delivery.

Validate the preview like a real platform would

A human can perform a smoke test:

curl -fsS http://127.0.0.1:8083 >/dev/null \
  && echo "preview responds successfully"

You could also run:

integration tests
synthetic transactions
schema compatibility checks
browser automation
security tests
performance smoke tests

Our Rollout has already run web-health as a pre-promotion AnalysisRun.

Inspect it:

kubectl get analysisrun \
  -n web-dev \
  --sort-by=.metadata.creationTimestamp

Describe the newest one:

LATEST_ANALYSIS=$(kubectl get analysisrun \
  -n web-dev \
  --sort-by=.metadata.creationTimestamp \
  -o jsonpath='{.items[-1:].metadata.name}')

kubectl describe analysisrun \
  -n web-dev \
  "$LATEST_ANALYSIS"

At this point we know:

image built successfully
        |
image was pulled successfully
        |
Pods became Ready
        |
preview endpoint works
        |
automated gate passed
        |
production still serves old version

This is a much stronger decision point than:

CI job went green -> deploy everything

Promote the preview

When satisfied, promote it:

kubectl argo rollouts promote web \
  -n web-dev

Watch:

kubectl argo rollouts get rollout web \
  -n web-dev \
  --watch

During promotion, Rollouts will ensure the new ReplicaSet reaches the full desired size before switching production traffic.

Inspect the Services again:

kubectl get service web web-preview \
  -n web-dev \
  -o custom-columns='SERVICE:.metadata.name,HASH:.spec.selector.rollouts-pod-template-hash'

Both should now point at the new ReplicaSet hash.

Re-run the production request:

curl -s http://127.0.0.1:8082 \
  | grep version

Now production should say:

version: blue-green-v4

The important operation was not replacing the Service.

The Service stayed stable:

web.web-dev.svc.cluster.local

Argo Rollouts changed which ReplicaSet that stable identity selected.

Observe the old version before it disappears

Immediately after promotion:

kubectl get rs \
  -n web-dev \
  -l app=web

The old active ReplicaSet should remain scaled for roughly our configured delay:

scaleDownDelaySeconds: 60

Watch it:

kubectl get rs \
  -n web-dev \
  -l app=web \
  -w

Eventually the old ReplicaSet scales down.

This delay is intentionally different from keeping old ReplicaSet metadata around.

Kubernetes may retain the old ReplicaSet object for revision history even after its replica count reaches zero.

The distinction is:

revision retained
        !=
old application still consuming full compute

32. Break Blue-Green Safely: Fail Before Production Sees It [DEV] [OPS]

Now let's make blue-green earn its keep.

We will deliberately create a release that:

builds successfully
runs successfully
becomes Ready

but fails our release gate.

That is more interesting than a broken container image because Kubernetes itself considers the workload healthy enough to run.

Our delivery policy rejects it.

Make the gate fail first

Set our analysis endpoint to false:

kubectl patch configmap analysis-gate \
  -n web-dev \
  --type merge \
  -p '{"data":{"result.json":"{\"ok\": false}\n"}}'

Restart the tiny gate:

kubectl rollout restart deployment/analysis-gate \
  -n web-dev

Verify:

kubectl run analysis-check \
  -n web-dev \
  --rm -i --restart=Never \
  --image=curlimages/curl \
  -- \
  curl -s http://analysis-gate/result.json

Expected:

{"ok": false}

Build another perfectly runnable release

Move to the application repository:

cd ~/gitops-lab/web-app

Create v5:

cat > index.html <<'EOF'
<!doctype html>
<html>
  <body>
    <h1>GitOps demo</h1>
    <p>version: blue-green-v5-REJECT-ME</p>
  </body>
</html>
EOF

Commit and push:

git add index.html

git commit -m "release blue-green v5 for failed gate lab"

git push

Again, CI should build and push the image, resolve the digest, and open a GitOps promotion PR.

Merge that PR.

Watch the new version appear only behind preview

Watch:

kubectl argo rollouts get rollout web \
  -n web-dev \
  --watch

The preview ReplicaSet should start normally.

Inspect Services:

kubectl get service web web-preview \
  -n web-dev \
  -o custom-columns='SERVICE:.metadata.name,HASH:.spec.selector.rollouts-pod-template-hash'

And inspect the analysis:

kubectl get analysisrun \
  -n web-dev \
  --sort-by=.metadata.creationTimestamp

The new pre-promotion analysis should fail.

Inspect it:

LATEST_ANALYSIS=$(kubectl get analysisrun \
  -n web-dev \
  --sort-by=.metadata.creationTimestamp \
  -o jsonpath='{.items[-1:].metadata.name}')

kubectl describe analysisrun \
  -n web-dev \
  "$LATEST_ANALYSIS"

The Rollout should enter an aborted/degraded state rather than switching the active Service.

Prove production never moved

Query the active Service:

curl -s http://127.0.0.1:8082 \
  | grep version

It should still be the known-good version:

version: blue-green-v4

Query preview:

curl -s http://127.0.0.1:8083 \
  | grep version

Depending on the precise aborted state and cleanup timing, the preview Service may still expose the rejected ReplicaSet long enough to debug it.

The critical fact is:

active Service did not switch

Production never needed a rollback because production never received the rejected release.

That is blue-green's nicest failure mode.

But Git still asks for v5

We have the same desired-state issue we saw during the canary abort.

Operational state says:

v4 is serving safely
v5 was rejected

Git still says:

v5 digest is desired

So the Rollout is correctly unhealthy/degraded.

It has not forgotten what we asked for.

Again:

protecting production
        !=
fixing desired state

Revert the promotion through Git

Move to the GitOps repository:

cd ~/gitops-lab/platform-gitops

git checkout main
git pull

Find the v5 promotion commit:

git log --oneline --decorate -10

Create a rollback branch:

git checkout -b rollback/blue-green-v5

Revert the promotion merge:

git revert <V5_PROMOTION_MERGE_COMMIT>

Push it:

git push -u origin rollback/blue-green-v5

Open the rollback PR:

gh pr create \
  --title "Rollback rejected blue-green v5" \
  --body "Restore the last known-good digest after pre-promotion analysis rejected v5."

Review and merge it.

Then:

git checkout main
git pull

Watch Argo CD and Argo Rollouts converge:

kubectl get application web-dev \
  -n argocd \
  -w

and:

kubectl argo rollouts get rollout web \
  -n web-dev \
  --watch

Because we configured:

rollbackWindow:
  revisions: 3

returning Git to a recent ReplicaSet can be fast-tracked instead of needlessly replaying the entire progressive-delivery process.

Restore the analysis gate

Leave the lab in a healthy state:

kubectl patch configmap analysis-gate \
  -n web-dev \
  --type merge \
  -p '{"data":{"result.json":"{\"ok\": true}\n"}}'

Restart it:

kubectl rollout restart deployment/analysis-gate \
  -n web-dev

33. Rolling, Canary and Blue-Green Are Different Risk Controls [OPS] [PLATFORM]

We now have enough experience to compare the strategies based on behaviour rather than vocabulary.

Strategy What happens to the new version? Production exposure Extra infrastructure Best fit
RollingUpdate Pods gradually replace old Pods increases as replacement proceeds usually low simple stateless workloads
Canary old and new versions coexist gradually increased moderate measurable production risk
Blue-green full/partial preview is created separately none until cutover potentially high strong pre-production validation and fast cutover

The easiest way to remember the difference is:

RollingUpdate:
  replace gradually

Canary:
  expose gradually

Blue-green:
  validate separately, then switch

The same GitOps promotion model supports all three

Notice what did not change between our canary and blue-green exercises:

application developer
      |
      v
application PR
      |
      v
CI build
      |
      v
immutable image digest
      |
      v
promotion PR
      |
      v
Git desired state
      |
      v
Argo CD

Only the release controller's strategy changed after Argo CD delivered the desired artifact to Kubernetes.

That separation is powerful.

A platform can standardise:

build
artifact identity
security evidence
promotion
Git review
cluster reconciliation

while allowing workloads to choose an appropriate rollout policy.

Rollback means different things at different moments

Before blue-green promotion:

new version rejected
production never moved

There is no traffic rollback to perform.

We only need to repair desired state in Git.

Immediately after blue-green promotion:

new version active
old ReplicaSet may still be scaled up

Operational rollback can be extremely fast.

After old capacity has been scaled down:

Git revert
   |
   v
old digest becomes desired
   |
   v
Rollouts restores the previous revision

With a rollback window, recent revisions can be fast-tracked.

For canary:

abort
  |
  v
stable ReplicaSet regains exposure

but Git must still be corrected if it describes the rejected image.

The recurring lesson is:

Operational safety actions and source-of-truth changes are related, but they are not the same operation.

The platform-engineering view

At this point our delivery system contains several reconcilers:

Git
 |
 v
Argo CD
 |
 +------------------------------+
 |                              |
 v                              v
Rollout policy               Services
 |                              ^
 v                              |
Argo Rollouts ------------------+
 |
 v
ReplicaSets
 |
 v
Pods

For canary with a traffic router we add another loop:

Argo Rollouts
      |
      v
Gateway API route weights
      |
      v
Cilium

For blue-green, Rollouts controls active/preview Service selectors instead.

This is why platform engineering is fundamentally about more than installing controllers.

You need to know:

which controller owns which resource
which controller owns which field
what the source of truth is
which mutations are temporary operational state
which mutations must go back through Git

That is the difference between a platform composed of controllers and a pile of controllers fighting each other.


Part V - Platform Engineering

34. CI Should Have Registry and Git Permissions - Not Cluster Admin [PLATFORM]

Our final developer pipeline has approximately these permissions:

application repository
  read source

OCI registry
  push application artifact

GitOps repository
  create branch
  create pull request

It does not need:

Kubernetes API credential
production kubeconfig
cluster-admin
SSH access to nodes

This dramatically reduces what a compromised CI runner can do directly.

That does not mean CI compromise is harmless.

An attacker able to build and promote an image may still be able to propose malicious software.

The defence becomes layered:

branch protection
review
artifact signing
vulnerability policy
admission policy
Argo project constraints
Kubernetes RBAC
runtime isolation

Good architecture reduces authority at each step rather than trusting one enormous credential.


35. The GitOps Repository Is Production Infrastructure [PLATFORM]

Treat the GitOps repository accordingly.

Do not think of it as a random directory of YAML.

It contains production authority.

Useful controls include:

branch protection
CODEOWNERS
required pull-request reviews
status checks
signed commits where appropriate
restricted automation credentials
audit logs
protected environment directories

For example:

/environments/dev/**
  application team may approve

/environments/prod/**
  service owner + platform policy approval

/platform/**
  platform team only

Git permissions are now part of the infrastructure permission model.

Earlier we asked:

Can Alice delete Nodes through Kubernetes RBAC?

GitOps introduces another question:

Can Alice merge desired state that causes Argo to create or delete powerful resources?

Indirect privilege is still privilege.


36. Repository Layout Is an API Design Problem [PLATFORM]

There is no universally correct GitOps repository shape.

The structure communicates ownership and promotion boundaries.

One reasonable layout is:

platform-gitops/
├── apps/
│   └── web/
│       └── base/
├── environments/
│   ├── dev/
│   │   └── web/
│   ├── staging/
│   │   └── web/
│   └── prod/
│       └── web/
└── platform/
    ├── argocd/
    ├── rollouts/
    └── gateway/

Another platform may use one repository per environment or business unit.

The useful questions are:

Who may approve this path?
What blast radius does changing this directory have?
Can one repository compromise every cluster?
How is an artifact promoted?
How do we audit the change?

Repository topology is not merely aesthetic.

It is part of your security and operating model.


37. ApplicationSet: When One Application Becomes Hundreds [OPS] [PLATFORM]

Creating one Argo Application by hand is fine.

Creating 500 nearly-identical Application objects manually is not.

Argo CD provides ApplicationSet to generate Applications from data.

Conceptually:

list of clusters / directories / tenants
            |
            v
       ApplicationSet
            |
            v
    many Applications
            |
            v
       many reconciliations

A simplified list example:

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: web-environments
  namespace: argocd
spec:
  generators:
    - list:
        elements:
          - env: dev
            namespace: web-dev
          - env: staging
            namespace: web-staging
  template:
    metadata:
      name: 'web-{{env}}'
    spec:
      project: web-team
      source:
        repoURL: https://github.com/example/platform-gitops.git
        targetRevision: main
        path: 'environments/{{env}}/web'
      destination:
        server: https://kubernetes.default.svc
        namespace: '{{namespace}}'
      syncPolicy:
        automated:
          prune: true
          selfHeal: true

Do not use ApplicationSet merely because it exists.

Use it when the repetition itself is data-driven.

Typical platform examples include:

one app per cluster
one app per tenant
one app per region
one environment directory per service

This is the GitOps version of moving from hand-created Pods to controllers.


38. App of Apps, ApplicationSet and Platform Bootstrapping [DEEP DIVE]

Eventually you may want GitOps to manage the platform components that enable GitOps.

That sounds circular because it is.

A bootstrap process often looks like:

create cluster
    |
    v
install minimal Argo CD
    |
    v
point Argo at platform bootstrap repository
    |
    v
Argo installs / manages
  - policies
  - ingress / Gateway API
  - observability
  - operators
  - team Applications

Once bootstrapped, most ongoing platform state can flow through Git.

You still need an answer for:

Who creates the cluster?
Who installs the first Argo instance?
Who provides its Git credentials?
Who upgrades Argo itself?

GitOps moves the bootstrap boundary.

It does not eliminate it.

This is where tools such as Cluster API, hosted control planes, Terraform/OpenTofu, MAAS/NiCO, vCluster, or Kamaji may sit beneath the application platform.

A useful full-stack model is:

infrastructure desired state
        |
        v
clusters / nodes / networks
        |
        v
GitOps bootstrap
        |
        v
platform services
        |
        v
application GitOps
        |
        v
workloads

39. Supply-Chain Policy: Build Evidence Must Reach Deployment Policy [PLATFORM]

We built and signed an artifact.

That is only half the story.

A mature platform connects build evidence to deployment policy.

For example:

CI produces:
  image digest
  signature
  provenance
  SBOM
  vulnerability result

promotion policy requires:
  trusted builder identity
  no forbidden vulnerabilities
  approved source repository
  immutable digest

admission policy verifies:
  image is allowed to start

This closes a gap that Git review alone cannot.

A perfectly reviewed Git commit could still point to:

unsigned malicious image

if nothing verifies the artifact.

Likewise, a perfectly signed artifact can still be accidentally promoted to the wrong environment if Git permissions are too broad.

The controls reinforce one another.


40. Harbor as a Platform Registry [PLATFORM]

Harbor becomes particularly useful when an organisation wants the registry itself to enforce platform rules.

A production pattern might be:

Harbor project: payments

CI robot:
  push
  read

runtime identity:
  pull only

platform admin:
  configure retention
  immutability
  scanning
  trust policy

Useful policy decisions:

Make release tags immutable

Allow:

sha-abc123
v1.4.2

but prevent those tags from being overwritten once published.

This does not replace digest pinning.

It makes human-readable references less surprising.

Generate SBOMs

Harbor can integrate SBOM generation with its scanner.

The SBOM gives visibility into the packages inside an artifact.

Enforce signatures

Harbor can associate Cosign/Notation signatures with OCI artifacts and can enforce content trust on a project.

The desired path becomes:

unsigned image
   X
cannot be consumed

signed trusted image
   |
   v
runtime may pull it

Use robot accounts

Automations should not log in using a platform administrator's personal username and password.

Create separate identities for:

CI publisher
replication
scanner
runtime pull

with the smallest permissions each needs.

This mirrors Kubernetes ServiceAccounts and RBAC.

Machine identity should be explicit.


41. Observability for the Delivery System [OPS] [PLATFORM]

The deployment system is production software too.

Monitor it.

Useful questions include:

Is Argo able to read Git?
Are Applications OutOfSync?
Are Applications Degraded?
How long do syncs take?
Are Rollouts paused or aborted?
Are AnalysisRuns failing?
Can nodes pull images?
Did registry scanning fail?
Are promotion PRs stuck?

A useful event chain for one release is:

source commit
    |
    v
CI run
    |
    v
image digest
    |
    v
promotion PR
    |
    v
Git merge commit
    |
    v
Argo sync revision
    |
    v
Rollout revision
    |
    v
ReplicaSet
    |
    v
Pods

Preserving those identifiers makes incident response dramatically easier.

Good platform metadata might include:

source git SHA
image digest
GitOps commit SHA
service name
environment
owner
release timestamp

42. Failure Drills [OPS]

A delivery system is not understood until you have watched it fail.

Run these deliberately in the lab.

Drill 1 - Git changes to an invalid image digest

Set the GitOps image to a nonexistent digest.

Observe:

Argo: desired state accepted
Kubernetes: ImagePullBackOff
Application health: degraded / progressing

Question:

Which controller is functioning correctly, and where is reconciliation blocked?

Drill 2 - Delete a managed Pod

kubectl delete pod \
  -n web-dev \
  -l app=web

Observe the Rollout/ReplicaSet restore it without any Git change.

Drill 3 - Scale live state manually

Change replica count with kubectl.

Observe Argo self-heal return it to Git.

Drill 4 - Remove a manifest from Git

Observe prune behaviour.

Drill 5 - Break repository access

Temporarily point the Application at a repository/path Argo cannot read.

Observe:

kubectl describe application web-dev \
  -n argocd

Drill 6 - Abort a canary but do not revert Git

Observe the tension between:

stable runtime state

and:

new desired Git state

Then repair it properly with a revert.

Drill 7 - Stop Argo CD

Scale the application controller down in the lab.

Observe that running workloads continue.

GitOps is a reconciliation mechanism, not the runtime dataplane.

Restore the controller and observe reconciliation resume.


43. Anti-Patterns That Look Reasonable [DEV] [OPS] [PLATFORM]

CI runs kubectl apply

It works, but pushes cluster credentials into the CI system and makes Git less authoritative.

Prefer:

CI -> artifact + PR
Argo -> cluster

Deploy :latest

You cannot reliably answer which bytes Git intended.

Prefer an immutable digest.

Rebuild the image for production

You are no longer deploying what staging tested.

Promote the same digest.

Let automation merge its own production PR immediately

Then the PR is theatre.

If no review or policy decision is required, be explicit about that rather than pretending there is a gate.

Give every Argo Application the default project forever

The default project is deliberately broad.

Create explicit source and destination boundaries.

Put plaintext Secrets in Git because "GitOps"

GitOps requires a secret-management pattern, not wishful thinking.

Sign images but never verify signatures

You collected evidence but do not enforce it.

Abort the rollout but leave Git pointing at the failed version

You fixed runtime exposure but not desired state.

Abort first if needed; revert Git next.

Let developers mutate production manually "just this once"

Sometimes incidents require imperative changes.

The important rule is:

Reconcile the source of truth immediately afterwards.

Otherwise emergency drift becomes permanent architecture.


44. A Production Delivery Reference Architecture [PLATFORM]

Putting the pieces together:

                         DEVELOPER
                             |
                             | pull request
                             v
                    APPLICATION REPOSITORY
                             |
                    review / tests / merge
                             |
                             v
                            CI
                  +----------+----------+
                  |                     |
               build                 test/scan
                  |                     |
                  +----------+----------+
                             |
                             v
                       OCI REGISTRY
                     Harbor / GHCR
                             |
                  +----------+----------+
                  |                     |
              image digest          signature
              SBOM                  provenance
                  |                     |
                  +----------+----------+
                             |
                             v
                     PROMOTION AUTOMATION
                             |
                             | opens PR
                             v
                       GITOPS REPOSITORY
                             |
                   review / policy / merge
                             |
                             v
                          ARGO CD
                             |
                             v
                     KUBERNETES API
                             |
                  +----------+----------+
                  |                     |
             Argo Rollouts         other controllers
                  |
        +---------+---------+
        |                   |
   stable ReplicaSet    canary ReplicaSet
        |                   |
        +---------+---------+
                  |
        Gateway API / Cilium
                  |
                  v
                USERS

Underneath this application-delivery plane may be another platform layer:

cluster lifecycle
      |
      +-- managed cloud Kubernetes
      +-- Cluster API
      +-- Kamaji
      +-- vCluster
      +-- metal provisioning
      +-- networking / storage

GitOps is not the entire platform.

It is one extremely useful reconciliation boundary inside the platform.


45. From Kubernetes Learner to Platform Engineer [DEV] [OPS] [PLATFORM]

The progression through these cookbooks now looks approximately like this:

LEVEL 1 - Kubernetes user

Pod
Deployment
Service
ConfigMap
Secret
PVC
probes

        |
        v

LEVEL 2 - Kubernetes developer

rollouts
resources
scheduling
RBAC
Kustomize
Helm
Gateway API

        |
        v

LEVEL 3 - Kubernetes internals

spec / status
controllers
CRDs
ownerReferences
finalizers
custom Go reconciler

        |
        v

LEVEL 4 - delivery engineer

OCI images
immutable digests
CI
GitOps
Argo CD
promotion PRs
progressive delivery
rollback

        |
        v

LEVEL 5 - platform engineer

AppProjects
multi-tenancy
hosted control planes
Gateway API / Cilium
registry policy
artifact trust
secret systems
fleet management
ApplicationSets
observability
policy

        |
        v

LEVEL 6 - platform designer

Who owns desired state?
Where are trust boundaries?
Which controller owns each resource?
How does software move between environments?
How is privilege constrained?
How do we prove what is running?
How does the platform fail safely?

The tools will change.

Those questions age much more slowly.


46. Production Checklist [PLATFORM]

Before calling a GitOps delivery platform production-ready, be able to answer all of these.

Artifact

[ ] Is every release associated with an immutable digest?
[ ] Is the same digest promoted through environments?
[ ] Are release tags immutable where practical?
[ ] Are artifacts vulnerability scanned?
[ ] Is an SBOM available?
[ ] Are artifacts signed / attested?
[ ] Is signature/provenance verification enforced somewhere?

CI

[ ] Does CI avoid general Kubernetes credentials?
[ ] Are registry credentials scoped?
[ ] Is GitOps write access scoped to the required repository/path?
[ ] Is automation represented by a machine identity, not a person's token?
[ ] Can the build identity be audited?

Git

[ ] Are production branches protected?
[ ] Are CODEOWNERS / reviewers appropriate?
[ ] Is direct push disabled where required?
[ ] Can every production digest be traced to a reviewed change?
[ ] Is rollback performed through source-of-truth changes?

Argo CD

[ ] Are AppProjects explicit rather than relying on default?
[ ] Are source repositories restricted?
[ ] Are destination namespaces/clusters restricted?
[ ] Is Argo's own Kubernetes RBAC understood?
[ ] Is access to the argocd namespace tightly controlled?
[ ] Are prune and self-heal policies deliberate?
[ ] Are Argo upgrades and backups planned?

Progressive delivery

[ ] Is rollout strategy appropriate for the service?
[ ] Are canary signals meaningful?
[ ] Can a release be aborted quickly?
[ ] Is there a documented Git rollback path?
[ ] Are automated analyses observable and auditable?
[ ] Is traffic shaping coarse replica ratio or real routed traffic by design?

Secrets

[ ] Are plaintext production secrets kept out of ordinary Git history?
[ ] Is secret rotation supported?
[ ] Can workloads access only the secrets they need?

Operations

[ ] Can you identify source commit -> image digest -> GitOps commit -> running Pods?
[ ] Are reconciliation failures alerted?
[ ] Are OutOfSync and Degraded applications visible?
[ ] Are stuck/aborted Rollouts visible?
[ ] Have failure drills actually been run?

If several answers are "we assume so", keep building.


47. Cleanup [DEV]

Remove the Argo-managed application first:

kubectl delete application web-dev \
  -n argocd

Depending on finalizer/prune configuration, inspect whether its managed resources remain:

kubectl get all \
  -n web-dev

Delete the lab namespace if required:

kubectl delete namespace web-dev

Remove Argo Rollouts:

kubectl delete namespace argo-rollouts

Remove Argo CD:

kubectl delete namespace argocd

The CRDs are cluster-scoped and may remain after deleting namespaces.

Inspect:

kubectl get crd \
  | grep argoproj

For a disposable kind lab, the cleanest full reset is often simply deleting and recreating the cluster.

Do not blindly use that advice on a cluster containing anything you care about.


Appendix A - The Whole Reconciliation Stack

A useful final mental model:

SOURCE CODE
   |
   | developer changes application behaviour
   v
APPLICATION GIT
   |
   | CI reconciles source into an artifact
   v
OCI IMAGE DIGEST
   |
   | promotion automation proposes desired deployment
   v
GITOPS GIT
   |
   | Argo CD reconciles Git into Kubernetes resources
   v
ROLLOUT / DEPLOYMENT SPEC
   |
   | workload controller reconciles replicas
   v
REPLICASETS / PODS
   |
   | kubelet reconciles Pod specs on nodes
   v
CONTAINERS

With progressive delivery:

ROLLOUT
   |
   +--> stable ReplicaSet
   |
   +--> canary ReplicaSet
   |
   +--> AnalysisRun
   |
   +--> traffic routing policy

With Gateway API:

HTTPRoute desired weight
        |
        v
Gateway controller / Cilium
        |
        v
network dataplane

With secrets:

ExternalSecret
      |
      v
secret controller
      |
      v
Kubernetes Secret

With a multi-tenant platform:

tenant Git
   |
   v
Argo permissions
   |
   v
tenant API / namespace boundary
   |
   v
worker isolation

A modern Kubernetes platform is not one giant program.

It is a set of reconciliation loops with deliberately-designed ownership and trust boundaries.


Appendix B - Tag vs Digest Cheat Sheet

Use tags for humans:

web:v1.4.2
web:sha-4f91d20

Use digests when exact bytes matter:

web@sha256:abc123...

A useful CI output is both:

Tag:
  ghcr.io/acme/web:sha-4f91d20

Digest:
  sha256:abc123...

The GitOps repository promotes:

ghcr.io/acme/web@sha256:abc123...

The PR body can show the friendly tag/source commit for humans.


Appendix C - When to Use Argo CD Image Updater, Renovate or Similar

In this handbook, application CI opens the promotion PR itself.

That is not the only valid automation model.

Another controller or bot can watch registries and propose Git changes.

Examples include:

Argo CD Image Updater
Renovate
Flux image automation
custom release automation

The important requirement is not which bot edits Git.

The requirement is that the final desired-state decision remains explicit and auditable.

A useful distinction is:

DISCOVERY
"a new artifact exists"

PROMOTION
"this environment should now run it"

Do not accidentally turn discovery into uncontrolled production promotion unless that is explicitly the policy you want.

For example:

automatically discover every new dev image
    -> probably fine

automatically deploy every new image to prod
    -> much stronger decision

Appendix D - Further Reading

Prefer primary documentation because these projects evolve quickly.

Argo CD

Argo Rollouts

Kubernetes

Harbor

Sigstore


Final Takeaway

Bad CI/CD often gives one automation system enormous authority:

source
  |
  v
CI
  |
  | build
  | mutate production
  | own cluster credentials
  v
cluster

A better delivery platform separates concerns:

source
  |
  v
CI
  |
  | produces immutable evidence
  v
artifact
  |
  v
promotion PR
  |
  | changes reviewed desired state
  v
Git
  |
  v
Argo CD
  |
  | reconciles cluster intent
  v
Argo Rollouts
  |
  | controls exposure
  v
runtime

At each boundary we can ask:

What is the desired state?
Who may change it?
Which controller reconciles it?
What identity does that controller use?
How do we observe failure?
How do we return to a known-good state?

If you can answer those questions, you are no longer merely deploying to Kubernetes.

You are designing a platform.

Kubernetes for Application Developers

A runnable Kubernetes cookbook for application developers and CKAD candidates.

The goal is not to memorise YAML.

The goal is to build a mental model of Kubernetes, use the API deliberately, observe what the control plane did, break things on purpose, and work out why they broke.

Target: Kubernetes 1.35 and the current CKAD curriculum.

Sections are marked:

  • [CKAD] - directly relevant to CKAD
  • [DEV] - practical application developer knowledge
  • [DEEP DIVE] - controllers, operators and platform engineering

This cookbook is intentionally cumulative. We will keep reusing the same resources so that later concepts explain earlier behaviour rather than appearing as unrelated YAML fragments.

The recurring teaching loop is:

problem
  |
  v
mental model
  |
  v
small experiment
  |
  v
observe Kubernetes
  |
  v
change one thing
  |
  v
observe the consequence
  |
  v
break an assumption
  |
  v
explain why

When a section does not need every step, we will not force it into a template.


Part I - Learn the Control Loop

0. Lab Setup [CKAD]

Use one namespace for the whole cookbook so that commands stay short and cleanup is easy.

kubectl create namespace cookbook

Make it the default namespace for the current context:

kubectl config set-context \
  --current \
  --namespace=cookbook

Check:

kubectl config view --minify \
  -o jsonpath='{..namespace}{"\n"}'

Expected:

cookbook

Check that the cluster is reachable:

kubectl version
kubectl get nodes

Useful throughout the cookbook:

kubectl get all

Despite the name, all does not mean every Kubernetes API resource.

For actual API discovery:

kubectl api-resources

A note about the lab

Most examples work on any ordinary Kubernetes cluster.

A few exercises depend on optional cluster capabilities:

  • PersistentVolumeClaims need a storage provisioner or an existing PersistentVolume.
  • NetworkPolicy enforcement needs a CNI that implements NetworkPolicy.
  • Ingress routing needs an Ingress controller.
  • CRD exercises need permission to create cluster-scoped CustomResourceDefinitions.

When one of those assumptions matters, the recipe will say so.


1. The Kubernetes Control Loop [CKAD] [DEV]

Before learning all of Kubernetes' resource types, learn the one idea that connects nearly all of them.

Kubernetes is not primarily a system for running commands on servers.

It is a system for declaring what you want and continuously trying to make reality match it.

Conceptually:

you declare what you want
        |
        v
   Kubernetes API
        |
        v
     controller
    /    |     \
observe compare  act
    \    |     /
        reality
        |
        +-------- repeat

This repeating process is called reconciliation.

Let's make it happen rather than just defining it.

Start with one application

Create an nginx Deployment with one replica:

kubectl create deployment api \
  --image=nginx:1.27-alpine \
  --replicas=1

Look at the Deployment:

kubectl get deployment api

And the Pod it caused Kubernetes to create:

kubectl get pods -l app=api

You asked Kubernetes for one copy of nginx.

A Pod now exists.

The important part is how Kubernetes represents that request.

spec is what you asked for

Ask the API how many replicas the Deployment should have:

kubectl get deployment api \
  -o jsonpath='{.spec.replicas}{"\n"}'

Expected:

1

That value lives under something like:

spec:
  replicas: 1

spec describes desired state.

It is your instruction to Kubernetes:

I want one replica of this application.

Now ask Kubernetes how many replicas are actually ready:

kubectl get deployment api \
  -o jsonpath='{.status.readyReplicas}{"\n"}'

Once nginx is ready, you should see:

1

That value comes from status:

status:
  readyReplicas: 1

A useful first approximation is:

spec    -> what you want
status  -> what Kubernetes currently sees

You normally change spec.

Kubernetes and its controllers populate status.

Change what you want

Suppose one replica is no longer enough.

Tell Kubernetes you want three:

kubectl scale deployment api --replicas=3

This command did not directly start two containers.

It changed the desired state stored in the Kubernetes API.

Check:

kubectl get deployment api \
  -o jsonpath='{.spec.replicas}{"\n"}'

Now:

3

Watch the Pods:

kubectl get pods -l app=api -w

Two more Pods should appear.

Press Ctrl-C once all three are running.

Now compare desired state with observed state:

kubectl get deployment api \
  -o jsonpath='{.spec.replicas}{" desired, "}{.status.readyReplicas}{" ready\n"}'

Eventually:

3 desired, 3 ready

For a short time the system may have looked like:

spec.replicas:          3
status.readyReplicas:   1

Desired state and observed state disagreed.

Controllers noticed the difference and acted.

Eventually:

spec.replicas:          3
status.readyReplicas:   3

Reality caught up with intent.

That is reconciliation.

Change reality instead

Now do the opposite.

Instead of changing what we want, interfere with what actually exists.

Delete one Pod:

kubectl delete pod \
  $(kubectl get pods \
    -l app=api \
    -o jsonpath='{.items[0].metadata.name}')

Immediately watch:

kubectl get pods -l app=api -w

The Pod you deleted disappears.

Then another Pod appears.

This distinction matters:

The old Pod did not restart. Kubernetes created a replacement.

Why?

Deleting a Pod changed actual state, but it did not change the Deployment's desired state.

The Deployment still says, in effect:

spec:
  replicas: 3

Kubernetes therefore sees:

desired replicas: 3
actual replicas:  2

and reconciles the difference:

desired: 3
    |
    v
observe: 2
    |
    v
difference: -1
    |
    v
create another Pod
    |
    v
actual: 3

This is why Kubernetes applications should normally be designed around replaceable workloads rather than individual machines or individual containers.

There is more than one controller involved

We have simplified the story slightly.

A Deployment does not directly create Pods.

The relationship is approximately:

Deployment
    |
    v
ReplicaSet
    |
    +-- Pod
    +-- Pod
    +-- Pod

The Deployment controller manages ReplicaSets.

The ReplicaSet controller makes sure the correct number of matching Pods exists.

See the hierarchy:

kubectl get deployment api
kubectl get replicasets
kubectl get pods -l app=api

Inspect Pod ownership:

kubectl get pods \
  -l app=api \
  -o custom-columns='POD:.metadata.name,OWNER:.metadata.ownerReferences[0].name'

Do not worry about memorising ReplicaSets yet.

For now, remember the pattern:

declare
   |
   v
observe
   |
   v
compare
   |
   v
reconcile
   |
   +------ repeat

That pattern is Kubernetes.

How do we know the controller saw our change?

Many controller-managed objects expose another useful pair of fields.

kubectl get deployment api \
  -o jsonpath='{.metadata.generation}{" desired generation, "}{.status.observedGeneration}{" observed generation\n"}'

metadata.generation changes when the desired configuration changes.

status.observedGeneration tells us which generation the controller has observed.

You do not need to memorise those fields yet.

The useful idea is that Kubernetes APIs can expose both:

what was requested

and:

how far the controller has progressed towards it

That becomes very useful when debugging.

Useful detour: talk to the application

We know nginx is running.

Prove it with a local port-forward:

kubectl port-forward deployment/api 8080:80

In another terminal:

curl http://127.0.0.1:8080

You should receive nginx's HTML response.

kubectl port-forward is extremely useful while developing and debugging because it creates a temporary tunnel from your machine to a workload in the cluster.

It is not application networking.

your laptop
    |
localhost:8080
    |
 kubectl
    |
API server
    |
    v
   Pod

When the command stops, the tunnel stops.

Later we will create a Service, which solves a different problem: giving replaceable Pods a stable network identity inside Kubernetes.

What to keep from this chapter

Do not memorise the commands yet.

Remember the model:

spec
 |
 | what should exist
 v
controller
 |
 | continuously reconciles
 v
actual state
 |
 | reported back as
 v
status

Changing spec changes your intent.

Changing reality without changing spec causes Kubernetes to try to repair reality.

Almost everything else in this cookbook builds on that idea.


2. API Objects and Changing Desired State [CKAD] [DEV]

The previous chapter used commands such as:

kubectl create deployment ...
kubectl scale deployment ...

It is tempting to think of those as special operations implemented by Kubernetes.

A better mental model is:

kubectl is mostly a client for the Kubernetes API.

Different commands give us different ways to create or modify API objects.

The common shape of a Kubernetes object

Most manifests have the same top-level structure:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: example
spec:
  replicas: 3

Think:

apiVersion -> which API schema?
kind       -> what type of object?
metadata   -> which object?
spec       -> what do I want?
status     -> what does Kubernetes observe?

status is normally omitted from manifests you write.

Controllers populate it after the object exists.

Ask Kubernetes for the object we already created

kubectl get deployment api -o yaml

There is a lot of output.

Do not try to read it all.

Find these areas:

metadata:
spec:
status:

The object in the API is richer than the small amount of configuration we initially supplied because Kubernetes defaults fields and controllers add status.

Generate YAML instead of writing it from memory

Generate a Deployment manifest without creating anything:

kubectl create deployment generated \
  --image=nginx:1.27-alpine \
  --replicas=2 \
  --dry-run=client \
  -o yaml

Save it:

kubectl create deployment generated \
  --image=nginx:1.27-alpine \
  --replicas=2 \
  --dry-run=client \
  -o yaml > generated.yaml

Now the manifest is editable source rather than something you had to remember from scratch.

Several commands, one underlying idea

These commands look different:

kubectl scale
kubectl set image
kubectl edit
kubectl patch
kubectl apply

But they all ultimately modify desired state stored in the API.

kubectl scale

Purpose-built mutation:

kubectl scale deployment api --replicas=4

It changes:

.spec.replicas

Return to three replicas:

kubectl scale deployment api --replicas=3

kubectl set image

Another purpose-built mutation:

kubectl set image deployment/api \
  nginx=nginx:1.27-alpine

It changes the container image in the Deployment's Pod template.

kubectl edit

Open the live object in an editor:

kubectl edit deployment api

Useful during an exam or incident.

Less useful as a repeatable deployment process because the change only exists as an API mutation unless you also update your source manifest.

kubectl patch

Change a small part of an object without rewriting the whole manifest:

kubectl patch deployment api \
  --type=merge \
  -p '{"spec":{"replicas":2}}'

Check:

kubectl get deployment api \
  -o jsonpath='{.spec.replicas}{"\n"}'

Then restore:

kubectl scale deployment api --replicas=3

kubectl apply

apply is the usual declarative workflow.

You keep the desired configuration in a file and ask Kubernetes to make the live object agree with it.

For example:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: declarative-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: declarative-demo
  template:
    metadata:
      labels:
        app: declarative-demo
    spec:
      containers:
        - name: nginx
          image: nginx:1.27-alpine

Save as declarative-demo.yaml and apply:

kubectl apply -f declarative-demo.yaml

Change:

replicas: 2

Apply again:

kubectl apply -f declarative-demo.yaml

Observe:

kubectl get deployment declarative-demo

Cleanup:

kubectl delete -f declarative-demo.yaml
rm -f declarative-demo.yaml generated.yaml

Imperative and declarative are not opposing religions

For application developers, both are useful.

Imperative commands are excellent for:

  • exploration
  • debugging
  • CKAD speed
  • one-off mutations
  • generating starter YAML

Declarative files are excellent for:

  • review
  • version control
  • repeatability
  • GitOps
  • production change management

The important thing is understanding what API field you are changing and why.


3. API Discovery and Querying [CKAD]

Kubernetes has a large API.

Do not guess resource names or YAML fields when the cluster can tell you.

Discover resource types

kubectl api-resources

Find Deployments:

kubectl api-resources | grep -i deployment

You should see the apps API group.

Check supported API versions:

kubectl api-versions

Ask for the schema

kubectl explain deployment

Dig deeper:

kubectl explain deployment.spec
kubectl explain deployment.spec.template
kubectl explain deployment.spec.template.spec.containers

For Pods:

kubectl explain pod.spec.containers.resources

This is usually faster and safer than guessing YAML.

Query individual fields

Deployment name:

kubectl get deployment api \
  -o jsonpath='{.metadata.name}{"\n"}'

Desired replicas:

kubectl get deployment api \
  -o jsonpath='{.spec.replicas}{"\n"}'

Ready replicas:

kubectl get deployment api \
  -o jsonpath='{.status.readyReplicas}{"\n"}'

Container image:

kubectl get deployment api \
  -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'

Pod IPs:

kubectl get pods \
  -l app=api \
  -o custom-columns='NAME:.metadata.name,IP:.status.podIP'

This is querying the Kubernetes API.

It is not executing commands inside a container.

Learn to switch output formats

Human summary:

kubectl get pods

More columns:

kubectl get pods -o wide

Full object:

kubectl get pod <pod-name> -o yaml

JSON:

kubectl get pod <pod-name> -o json

Selected fields:

kubectl get pods \
  -o custom-columns='NAME:.metadata.name,PHASE:.status.phase,NODE:.spec.nodeName'

A strong Kubernetes workflow is often:

get
 |
 v
describe
 |
 v
query exact fields
 |
 v
logs/events when needed

Part II - Pods and Workload Controllers

4. Pods: The Unit Kubernetes Schedules [CKAD]

A Pod is the smallest workload Kubernetes schedules onto a node.

A Pod contains one or more containers that share:

  • an IP address
  • a network namespace
  • localhost
  • optional volumes
  • scheduling fate

The most important practical property is this:

Pods are replaceable.

Do not treat a Pod name or Pod IP as a durable application endpoint.

The container image is an input to Kubernetes

Kubernetes schedules and runs container images; it does not normally build your application image for you.

A minimal image definition might look like:

FROM nginx:1.27-alpine
COPY index.html /usr/share/nginx/html/index.html

Build it with an OCI-compatible image tool such as Docker or Podman:

docker build -t example/web:v1 .

or:

podman build -t example/web:v1 .

The cluster then needs to be able to pull or otherwise access that image. In a real workflow that usually means pushing it to a registry the cluster can reach.

The Pod spec references the resulting image:

spec:
  containers:
    - name: web
      image: example/web:v1

Keep the responsibility boundary clear:

Dockerfile / Containerfile
      |
      v
build an OCI image
      |
      v
registry / cluster image store
      |
      v
Pod spec references image
      |
      v
kubelet asks container runtime to run it

Most exercises use public images so we can concentrate on Kubernetes itself. In Chapter 9 we will deliberately build our own image and discover what “the cluster needs to be able to access that image” actually means.

Create a standalone Pod

The Deployment from Chapter 1 is controller-managed.

Create a separate Pod so we can inspect Pod behaviour without a controller replacing it:

kubectl run pod-lab \
  --image=nginx:1.27-alpine \
  --restart=Never \
  --labels=app=pod-lab

Observe:

kubectl get pod pod-lab
kubectl get pod pod-lab -o wide

describe connects configuration to runtime events

kubectl describe pod pod-lab

Pay attention to:

Node
IP
Containers
Conditions
Events

get -o yaml gives the complete API object.

describe gives a human-oriented operational summary.

Both are useful for different reasons.

Logs are container output

kubectl logs pod-lab

NGINX may have little to show until it receives traffic.

Port-forward temporarily:

kubectl port-forward pod/pod-lab 8081:80

In another terminal:

curl http://127.0.0.1:8081

Now check logs again:

kubectl logs pod-lab

exec runs a process inside the existing container

kubectl exec pod-lab -- nginx -v

Try:

kubectl exec pod-lab -- ps

For an interactive shell:

kubectl exec -it pod-lab -- sh

Exit with:

exit

Container command and args

Kubernetes can override an image's configured entrypoint and arguments.

A rough mapping is:

container image ENTRYPOINT -> command
container image CMD        -> args

Example:

apiVersion: v1
kind: Pod
metadata:
  name: command-demo
spec:
  restartPolicy: Never
  containers:
    - name: shell
      image: busybox:1.36
      command: ["sh", "-c"]
      args:
        - echo "hello from command-demo"; sleep 30

Apply:

kubectl apply -f command-demo.yaml

Logs:

kubectl logs command-demo

Cleanup:

kubectl delete pod command-demo --ignore-not-found

What happens without a controller?

Delete the standalone Pod:

kubectl delete pod pod-lab

Then:

kubectl get pod pod-lab

It is gone.

No controller declared that pod-lab should continue to exist.

Compare that with deleting one of the api Deployment Pods: the controller recreates it.

That difference is why production applications are normally managed by workload controllers rather than naked Pods.


5. Labels, Selectors and Annotations [CKAD]

Kubernetes needs a way for independent API objects to find groups of other objects.

It uses labels and selectors heavily for this.

Labels are small key/value identifiers such as:

app=api
environment=dev
release=stable

See the labels on our application

kubectl get pods -l app=api --show-labels

The Deployment's Pod template contains:

metadata:
  labels:
    app: api

The ReplicaSet creates Pods carrying those labels.

Create some disposable labelled Pods

kubectl run label-a \
  --image=busybox:1.36 \
  --restart=Never \
  --labels='app=demo,environment=dev' \
  --command -- sleep 3600

kubectl run label-b \
  --image=busybox:1.36 \
  --restart=Never \
  --labels='app=demo,environment=test' \
  --command -- sleep 3600

Select both:

kubectl get pods -l app=demo

Select only dev:

kubectl get pods -l app=demo,environment=dev

Set-based selector:

kubectl get pods -l 'environment in (dev,test)'

Change a label:

kubectl label pod label-a environment=test --overwrite

Now:

kubectl get pods -l environment=dev
kubectl get pods -l environment=test

Why selectors matter

Selectors connect otherwise independent resources.

Later we will see:

Service selector
      |
      v
matching Pod labels

and:

ReplicaSet selector
      |
      v
matching Pod labels

A selector mismatch is therefore not cosmetic.

It changes behaviour.

Annotations are different

Annotations are metadata that is not normally used for selection.

Example:

kubectl annotate pod label-a \
  example.com/description='temporary label experiment'

Inspect:

kubectl get pod label-a \
  -o jsonpath='{.metadata.annotations}{"\n"}'

Think:

labels      -> identity and selection
annotations -> additional metadata

Cleanup:

kubectl delete pod label-a label-b

6. Requests, Limits and Scheduling [CKAD]

A container image tells Kubernetes what to run.

The scheduler also needs to know what resources it needs.

That is where requests and limits enter.

A useful first approximation is:

request -> capacity considered when scheduling me
limit   -> runtime ceiling imposed on me

CPU and memory behave differently when a limit is exceeded:

  • CPU is throttled.
  • Memory can result in the container being OOM-killed.

Create a Pod with resource requirements

Save as limited.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: limited
spec:
  containers:
    - name: web
      image: nginx:1.27-alpine
      resources:
        requests:
          cpu: 100m
          memory: 64Mi
        limits:
          cpu: 500m
          memory: 128Mi

Apply:

kubectl apply -f limited.yaml

Inspect:

kubectl describe pod limited

Or query exactly what you asked for:

kubectl get pod limited \
  -o jsonpath='{.spec.containers[0].resources}{"\n"}'

If Metrics Server exists:

kubectl top pod limited

Break scheduling on purpose

The scheduler cannot place a Pod whose resource request cannot be satisfied by any eligible node.

Save as unschedulable.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: unschedulable
spec:
  containers:
    - name: sleeper
      image: busybox:1.36
      command: ["sh", "-c", "sleep 3600"]
      resources:
        requests:
          memory: 1000Gi

Apply:

kubectl apply -f unschedulable.yaml

Observe:

kubectl get pod unschedulable

It should remain Pending.

Ask why:

kubectl describe pod unschedulable

Look at the Events section.

The useful debugging lesson is:

Pending Pod
   |
   v
scheduler could not bind it to a node
   |
   v
inspect scheduling events

Cleanup:

kubectl delete pod limited unschedulable
rm -f limited.yaml unschedulable.yaml

ResourceQuota and LimitRange

Requests and limits are workload configuration.

Namespaces can also have policy around resource consumption.

Inspect any existing quota:

kubectl get resourcequota
kubectl get limitrange

A ResourceQuota can cap aggregate namespace consumption.

A LimitRange can constrain or default per-container or per-Pod values.

You do not need to confuse these with requests and limits:

request / limit -> what this workload asks for
LimitRange      -> policy/defaults around individual workloads
ResourceQuota   -> aggregate namespace budget

Scheduling controls placement, not API permission

Later we will meet node selectors, affinity, taints and tolerations in broader platform contexts.

Keep one distinction in mind now:

scheduler controls
    -> where a Pod is eligible to run

RBAC
    -> what API operations an identity may perform

A NoSchedule taint on a control-plane node can keep ordinary workloads away from that node. It does not decide who may edit Node objects through the API, and it does not control who may SSH into the machine.

We will bring those boundaries together in Chapter 23.


7. Deployments and ReplicaSets [CKAD]

We already used a Deployment before formally studying it because it gave us the smallest useful reconciliation experiment.

Now we can unpack the machinery.

A Deployment is appropriate for replaceable application replicas where individual Pod identity does not matter.

The ownership chain is:

Deployment
    |
    v
ReplicaSet
    |
    +-- Pod
    +-- Pod
    +-- Pod

Inspect the Deployment

kubectl get deployment api

ReplicaSet:

kubectl get replicasets -l app=api

Pods:

kubectl get pods -l app=api

Follow ownership through the API

Pod -> ReplicaSet:

kubectl get pods \
  -l app=api \
  -o custom-columns='POD:.metadata.name,OWNER:.metadata.ownerReferences[0].name'

ReplicaSet -> Deployment:

kubectl get rs \
  -l app=api \
  -o custom-columns='RS:.metadata.name,OWNER:.metadata.ownerReferences[0].name'

The ownership chain is not hidden state.

It is represented in the API.

Scaling is reconciliation, not cloning

Scale to five:

kubectl scale deployment api --replicas=5

Watch:

kubectl get pods -l app=api -w

Then inspect:

kubectl get deployment api \
  -o jsonpath='{.spec.replicas}{" desired, "}{.status.readyReplicas}{" ready\n"}'

Scale back:

kubectl scale deployment api --replicas=3

Scaling does not create a new Deployment revision because it does not change the Pod template.

The Pod template is the important boundary

Inspect:

kubectl get deployment api \
  -o jsonpath='{.spec.template}{"\n"}'

A Deployment says, approximately:

I want N Pods that look like this template.

Changing replicas changes how many Pods.

Changing .spec.template changes what the Pods should be, which leads directly to rolling updates.


8. Rolling Updates, Revisions and Rollback [CKAD]

Suppose we want a new application version.

Replacing every Pod at once would create unnecessary downtime.

A Deployment can instead progressively move from one ReplicaSet to another.

Check the current image

kubectl get deployment api \
  -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'

Change the Pod template

kubectl set image deployment/api \
  nginx=nginx:1.28-alpine

Watch rollout progress:

kubectl rollout status deployment/api

Inspect ReplicaSets:

kubectl get rs -l app=api

You should now see an older ReplicaSet and a newer one.

Why did Kubernetes create a new ReplicaSet this time but not when we scaled?

Because the Pod template changed.

change .spec.replicas
      |
      v
same Pod template
      |
      v
same Deployment revision

change .spec.template
      |
      v
new desired Pod definition
      |
      v
new ReplicaSet / revision

Rollout history

kubectl rollout history deployment/api

Roll back

kubectl rollout undo deployment/api

Watch:

kubectl rollout status deployment/api

Check the image again:

kubectl get deployment api \
  -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'

RollingUpdate strategy

Inspect the strategy:

kubectl get deployment api \
  -o jsonpath='{.spec.strategy}{"\n"}'

And ask the schema what the fields mean:

kubectl explain deployment.spec.strategy
kubectl explain deployment.spec.strategy.rollingUpdate

The two important controls are:

maxUnavailable -> how many desired replicas may be unavailable
maxSurge       -> how many extra replicas may exist during rollout

Do not memorise defaults when kubectl explain can tell you what the current API supports.


9. Failed Rollouts: Debugging and Local Images [CKAD] [DEV]

A failed rollout is more educational than a successful one.

We will fail one in two different ways.

The first failure is obvious: we ask for an image that does not exist. That gives us a controlled way to learn the debugging path and rollback.

The second failure is more interesting: we build an image ourselves, prove that it exists, and Kubernetes still cannot run it.

Those two failures can produce the same Pod status for very different reasons.

That is exactly why good debugging starts with evidence rather than memorising status names.

Experiment 1 - deploy an image that does not exist

Change the Deployment to an image tag that is deliberately invalid:

kubectl set image deployment/api \
  nginx=nginx:this-tag-does-not-exist

Watch Pods:

kubectl get pods -l app=api -w

You should eventually see a new Pod move through states such as:

ErrImagePull
ImagePullBackOff

Press Ctrl-C once you have seen the failure.

Do not just memorise what ImagePullBackOff means.

Follow the object chain.

Is the Deployment healthy?

kubectl get deployment api

Try waiting for the rollout, but use a short timeout so the lab does not wait for the Deployment's full progress deadline:

kubectl rollout status deployment/api --timeout=30s

Which ReplicaSet is new?

kubectl get rs -l app=api

Which Pod is failing?

kubectl get pods -l app=api

Describe the failing Pod:

kubectl describe pod <failing-pod>

The Events section should explain the image pull failure.

Cluster events can also help:

kubectl get events \
  --sort-by=.metadata.creationTimestamp

The debugging path was:

Deployment
    |
    v
ReplicaSet
    |
    v
Pod
    |
    v
container image
    |
    v
image pull event

The high-level status told us where reconciliation was stuck.

Events told us why.

Why are the old Pods still running?

This is a useful consequence of the rolling update strategy from the previous chapter.

Inspect the Deployment and ReplicaSets together:

kubectl get deployment,rs,pods -l app=api

You may see three healthy old Pods plus one broken new Pod.

Conceptually:

Deployment api
    |
    +-- old ReplicaSet
    |      +-- Running Pod
    |      +-- Running Pod
    |      +-- Running Pod
    |
    +-- new ReplicaSet
           +-- ImagePullBackOff

Kubernetes has not forgotten that you asked for three replicas.

But the desired state now means more than just a number:

replicas = 3
AND
image = nginx:this-tag-does-not-exist

Reality currently satisfies the replica availability requirement using the old version, but it cannot satisfy the new Pod template.

A rolling update therefore stalls instead of immediately throwing away every healthy old Pod.

Fix the desired state

Undo the bad revision:

kubectl rollout undo deployment/api

Verify:

kubectl rollout status deployment/api
kubectl get pods -l app=api

This is the reconciliation model again.

We did not repair individual broken Pods.

We corrected desired state and let Kubernetes converge.


Sidequest - the image exists, so why can't Kubernetes run it?

The previous failure was unsurprising.

The image did not exist.

Now let's create an image that definitely does exist and see why Kubernetes may still report the same failure.

This is a useful developer workflow as well as a debugging exercise.

Build a tiny application image

Create a temporary working directory:

mkdir -p /tmp/k8s-web
cd /tmp/k8s-web

Create a simple page:

cat > index.html <<'EOF'
<!doctype html>
<html>
  <body>
    <h1>Hello from our own image</h1>
    <p>version: v1</p>
  </body>
</html>
EOF

Create a Dockerfile:

cat > Dockerfile <<'EOF'
FROM nginx:1.27-alpine
COPY index.html /usr/share/nginx/html/index.html
EOF

Build it:

docker build -t example/web:v1 .

Confirm that Docker can see it:

docker image inspect example/web:v1 \
  --format '{{.RepoTags}}'

We now have:

source files
    |
    v
docker build
    |
    v
example/web:v1

So the image exists.

The next question is where it exists.

Before changing the Deployment, inspect the container name

Run:

kubectl get deployment api \
  -o jsonpath='{range .spec.template.spec.containers[*]}{.name}{" -> "}{.image}{"\n"}{end}'

You should see something similar to:

nginx -> nginx:1.27-alpine

There are three separate identities here:

Deployment name: api
container name:  nginx
image name:      nginx:1.27-alpine

They are not interchangeable.

kubectl set image uses:

kubectl set image <resource> <container-name>=<new-image>

So this is correct:

kubectl set image deployment/api \
  nginx=example/web:v1

This would not be correct if the container were not named web:

kubectl set image deployment/api web=example/web:v1
                                 ^^^
                                 container name

Kubernetes is not identifying the container from the image name.

Watch the rollout

kubectl get pods -l app=api -w

On the kind lab cluster, you will probably see something like:

NAME                  READY   STATUS             RESTARTS
api-6645d7d87-24nw7   1/1     Running            0
api-6645d7d87-6wbjr   1/1     Running            0
api-6645d7d87-wcwpx   1/1     Running            0
api-8fc59985f-4nhhm   0/1     ErrImagePull       0
api-8fc59985f-4nhhm   0/1     ImagePullBackOff   0

Press Ctrl-C once the failure appears.

This looks very similar to our deliberately broken rollout.

But this time we know:

example/web:v1 exists

So "image does not exist" cannot be the whole explanation.

Use the debugging path again

Check the rollout:

kubectl rollout status deployment/api --timeout=30s

Inspect the ReplicaSets and Pods:

kubectl get deployment,rs,pods -l app=api

Describe the failing Pod:

kubectl describe pod <failing-pod>

Read the Events at the bottom.

The node is trying to obtain:

example/web:v1

and cannot.

The same symptom now has a different root cause.

Docker's image store is not the kind node's image store

Check the current Kubernetes context:

kubectl config current-context

For the lab used in this cookbook you may see:

kind-ckad

kind means Kubernetes IN Docker.

The Kubernetes node itself runs as a container and has its own container runtime.

Our image currently exists in the image store used by Docker on the host:

host
 |
 +-- Docker image store
 |      |
 |      +-- example/web:v1
 |
 +-- kind node container
        |
        +-- containerd image store
               |
               +-- example/web:v1 is missing

These are different image stores.

When the kubelet asks the node's container runtime to start:

example/web:v1

and the image is not present locally, the runtime attempts to obtain it from a registry according to the Pod's image pull policy.

We built the image locally, but we never pushed it to a registry.

So the pull fails.

This is the boundary that a happy-path public image hides:

docker build succeeding on your workstation does not mean every Kubernetes node can access that image.

In a normal remote or production workflow the path is usually:

developer / CI
      |
      v
build image
      |
      v
push to registry
      |
      v
Kubernetes node pulls image
      |
      v
container starts

For our disposable local kind cluster, there is a shortcut.

Load the image into kind

The context name is normally:

kind-<cluster-name>

So for:

kind-ckad

the kind cluster name is:

ckad

Load the image into its node:

kind load docker-image example/web:v1 \
  --name ckad

Now the path is:

host Docker image store
        |
        | kind load docker-image
        v
kind node image store
        |
        v
kubelet can start example/web:v1

Kubernetes may recover on its next image-pull retry.

If the existing Pod is already backing off and you want to make the next observation immediate, delete only the failing Pod:

kubectl delete pod <failing-pod>

Do not modify the Deployment.

The new ReplicaSet still wants a Pod matching the current template, so deleting the failed Pod causes another one to appear.

Watch:

kubectl get pods -l app=api -w

This time the replacement should be able to start from the image now present on the node.

Watch reconciliation continue

kubectl rollout status deployment/api

Then inspect the whole ownership chain:

kubectl get deployment,rs,pods -l app=api

Eventually the old ReplicaSet should be scaled to zero and the new ReplicaSet should own all three running Pods.

Notice what we did not do:

restart Kubernetes
recreate the Deployment
manually create three Pods

The desired state was already correct.

We fixed the external condition preventing the node from satisfying it.

The existing control loops carried on from there.

Prove that our application is running

Port-forward the Deployment:

kubectl port-forward deployment/api 8080:80

From another terminal:

curl http://127.0.0.1:8080

You should see HTML containing:

Hello from our own image
version: v1

Press Ctrl-C in the port-forward terminal when finished.

Do the workflow correctly with v2

Now repeat the process with the order understood.

Change the page:

cat > index.html <<'EOF'
<!doctype html>
<html>
  <body>
    <h1>Hello from our own image</h1>
    <p>version: v2</p>
  </body>
</html>
EOF

Build a new image tag:

docker build -t example/web:v2 .

Make the image available to the kind node before requesting it:

kind load docker-image example/web:v2 \
  --name ckad

Then change desired state:

kubectl set image deployment/api \
  nginx=example/web:v2

Watch the rollout:

kubectl rollout status deployment/api

The local development loop is now explicit:

edit source
    |
    v
build image
    |
    v
make image available to nodes
    |
    v
change Deployment spec
    |
    v
controllers reconcile
    |
    v
new Pods become ready

For kind:

docker build -t example/web:v2 .
kind load docker-image example/web:v2 --name ckad
kubectl set image deployment/api nginx=example/web:v2
kubectl rollout status deployment/api

On a real multi-node or remote cluster, the middle step is normally a registry rather than kind load.

Why use a new image tag?

During local development it is tempting to rebuild the same tag repeatedly:

example/web:v1

with different contents.

Prefer immutable or at least unique version tags while learning this workflow:

example/web:v1
example/web:v2
example/web:v3

Then the desired image reference tells you which build Kubernetes was asked to run.

Production systems often go further and deploy immutable image digests.

The important principle is the same:

know exactly which image desired state refers to

What the two failures taught us

Both experiments produced an image pull failure.

But the causes were different.

Failure 1

requested image
      |
      v
image does not exist
      |
      v
pull fails

The fix was to correct desired state:

kubectl rollout undo

Failure 2

requested image
      |
      v
image exists on developer machine
      |
      v
image absent from Kubernetes node
      |
      v
registry cannot provide it
      |
      v
pull fails

The desired state was valid.

The fix was to make the requested artifact available to the node:

kind load docker-image

That distinction is much more useful than memorising:

ImagePullBackOff = image problem

A better mental model is:

status tells you where the system is stuck
        |
        v
Events and object relationships tell you why

And underneath both examples is still the same Kubernetes control loop:

desired Pod template
        |
        v
controller creates replacement workload
        |
        v
node attempts to realise it
        |
        +---- cannot obtain image ----> status + Events expose failure
        |
        v
condition fixed
        |
        v
reconciliation continues

Part III - Stable Networking for Replaceable Pods

10. Services: Stable Identity for Replaceable Pods [CKAD]

Our Deployment Pods are intentionally disposable.

That creates a networking problem.

A Pod can disappear and its replacement can have a different IP address.

Clients therefore need something more stable than a Pod IP.

A Service provides a stable virtual endpoint over a selected set of backends.

Conceptually:

Client
  |
  v
Service
  |
  v
EndpointSlice
  |
  +-- Pod
  +-- Pod
  +-- Pod

Expose the Deployment

kubectl expose deployment api \
  --name=api \
  --port=80 \
  --target-port=80

Inspect:

kubectl get service api
kubectl describe service api

The important distinction is:

port       -> port clients use on the Service
targetPort -> port the application receives traffic on

How does the Service find Pods?

Query its selector:

kubectl get service api \
  -o jsonpath='{.spec.selector}{"\n"}'

You should see something equivalent to:

map[app:api]

Check matching Pods:

kubectl get pods -l app=api --show-labels

The Service does not contain a permanent list of those Pod IPs.

It declares a selector.

Another controller continuously maintains the backend endpoint data.

There is our reconciliation model again.


11. EndpointSlices and DNS [CKAD] [DEV]

A common simplified drawing is:

Service -> Pods

Useful, but incomplete.

For Services with selectors, Kubernetes normally maintains EndpointSlices containing matching backend endpoints.

Inspect EndpointSlices

kubectl get endpointslices

Filter for our Service:

kubectl get endpointslices \
  -l kubernetes.io/service-name=api

Show addresses:

kubectl get endpointslices \
  -l kubernetes.io/service-name=api \
  -o jsonpath='{.items[*].endpoints[*].addresses[*]}{"\n"}'

Compare with Pod IPs:

kubectl get pods \
  -l app=api \
  -o custom-columns='NAME:.metadata.name,IP:.status.podIP'

The addresses should correspond.

The relationship is closer to:

Service selector
      |
      v
matching Pods
      |
      v
EndpointSlice controller
      |
      v
EndpointSlices
      |
      v
networking implementation

Create a client Pod

We need something inside the cluster to test cluster networking.

kubectl run client \
  --image=busybox:1.36 \
  --restart=Never \
  --labels=app=client \
  --command -- sleep 3600

Wait until it is ready:

kubectl wait \
  --for=condition=Ready \
  pod/client \
  --timeout=60s

Service DNS

Resolve the Service:

kubectl exec client -- nslookup api

Call it:

kubectl exec client -- wget -qO- http://api

Try the namespace-qualified name:

kubectl exec client -- \
  wget -qO- http://api.cookbook

A conventional fully qualified Service name is:

api.cookbook.svc.cluster.local

You normally use the shortest name that is unambiguous.

Inside the same namespace:

api

is usually enough.

Compare Service and port-forward

A Service is persistent cluster configuration:

Pod replacement
     |
     v
endpoint set changes
     |
     v
Service identity remains

A port-forward is a temporary developer/debug tunnel.

Do not confuse them.


12. Break a Service and Follow the Network Path [CKAD]

Now deliberately break the selector.

First prove the application works:

kubectl exec client -- wget -qO- http://api

Inspect current endpoints:

kubectl get endpointslices \
  -l kubernetes.io/service-name=api

Break the selector

kubectl patch service api \
  --type=merge \
  -p '{"spec":{"selector":{"app":"broken"}}}'

Now ask for endpoint addresses:

kubectl get endpointslices \
  -l kubernetes.io/service-name=api \
  -o jsonpath='{.items[*].endpoints[*].addresses[*]}{"\n"}'

Try the request:

kubectl exec client -- \
  wget -T 2 -qO- http://api

It should fail.

Why?

Service selector = app=broken

Pods = app=api

No match
   |
   v
No ready backend endpoints

Fix desired state

kubectl patch service api \
  --type=merge \
  -p '{"spec":{"selector":{"app":"api"}}}'

Verify:

kubectl exec client -- wget -qO- http://api

A useful Service debugging order

When DNS resolves but traffic fails, walk the path rather than trying random commands:

Service
  |
  | selector and ports correct?
  v
Pod labels
  |
  | selector matches?
  v
EndpointSlice
  |
  | ready addresses present?
  v
Pod readiness
  |
  | endpoint considered ready?
  v
application port
  |
  | process actually listening?
  v
application

We will revisit this after adding readiness probes.


Part IV - Configuration and Health

13. ConfigMaps: Configuration Without Rebuilding the Image [CKAD]

Our current nginx-based image has application content baked into it.

Suppose that content or configuration needs to vary between environments.

Rebuilding an image for every small configuration change is often the wrong abstraction.

A ConfigMap stores non-secret configuration in the Kubernetes API.

Create configuration

kubectl create configmap api-content \
  --from-literal=index.html='hello from kubernetes'

Inspect:

kubectl get configmap api-content -o yaml

Query only the value:

kubectl get configmap api-content \
  -o jsonpath='{.data.index\.html}{"\n"}'

Mount the ConfigMap into the application

Patch the Deployment:

kubectl patch deployment api --type=strategic -p '
spec:
  template:
    spec:
      containers:
        - name: nginx
          volumeMounts:
            - name: content
              mountPath: /usr/share/nginx/html
      volumes:
        - name: content
          configMap:
            name: api-content
'

Because .spec.template changed, this creates a new Deployment revision.

Wait:

kubectl rollout status deployment/api

Call the Service:

kubectl exec client -- wget -qO- http://api

Expected:

hello from kubernetes

You have connected:

ConfigMap
    |
    v
Volume
    |
    v
Pod filesystem
    |
    v
Application

Change configuration

Update the ConfigMap declaratively from an imperative generator:

kubectl create configmap api-content \
  --from-literal=index.html='configuration changed' \
  --dry-run=client \
  -o yaml | kubectl apply -f -

ConfigMap-backed volumes are eventually updated by the kubelet.

After a short delay, retry:

kubectl exec client -- wget -qO- http://api

You should eventually see:

configuration changed

A subtle but important caveat:

A ConfigMap mounted using subPath does not receive the normal projected-volume updates.

That is one reason this example mounts the ConfigMap as the directory rather than mounting one key with subPath.

ConfigMap as environment variables

ConfigMaps can also populate environment variables.

The difference matters operationally:

ConfigMap volume
    -> projected into filesystem
    -> updates can appear later

ConfigMap environment variable
    -> value captured when container starts
    -> Pod must be recreated to see a new value

14. Secrets: Sensitive Configuration [CKAD]

Some configuration is sensitive enough that it should not be mixed casually with ordinary ConfigMaps.

Kubernetes provides the Secret API type.

Create one:

kubectl create secret generic api-secret \
  --from-literal=username=developer \
  --from-literal=password=correct-horse

Inspect metadata:

kubectl get secret api-secret

View YAML:

kubectl get secret api-secret -o yaml

Values under data are base64 encoded.

Decode one:

kubectl get secret api-secret \
  -o jsonpath='{.data.username}' | base64 -d

echo

Base64 is encoding, not encryption.

The security value of a Secret comes from Kubernetes access controls and how the cluster protects Secret data, not from base64.

stringData

Declarative Secrets can avoid manual base64 encoding:

apiVersion: v1
kind: Secret
metadata:
  name: example-secret
type: Opaque
stringData:
  token: example-value

The API server converts stringData into the encoded data representation.

Consume a Secret as an environment variable

Patch the Deployment:

kubectl patch deployment api --type=strategic -p '
spec:
  template:
    spec:
      containers:
        - name: nginx
          env:
            - name: API_USERNAME
              valueFrom:
                secretKeyRef:
                  name: api-secret
                  key: username
'

Wait for rollout:

kubectl rollout status deployment/api

Inspect one Pod:

POD=$(kubectl get pods -l app=api -o jsonpath='{.items[0].metadata.name}')
kubectl exec "$POD" -- printenv API_USERNAME

Expected:

developer

Use real secret-management controls in production rather than committing plaintext Secret manifests to Git.


15. The Downward API: Let the Workload See Its Kubernetes Identity [CKAD]

Applications sometimes need metadata about the Pod they are running inside.

Hard-coding it would defeat the purpose of replaceable Pods.

The Downward API can expose selected Pod fields as environment variables or files.

Save as downward.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: downward
  labels:
    app: downward-demo
spec:
  containers:
    - name: shell
      image: busybox:1.36
      command: ["sh", "-c", "sleep 3600"]
      env:
        - name: POD_NAME
          valueFrom:
            fieldRef:
              fieldPath: metadata.name
        - name: POD_NAMESPACE
          valueFrom:
            fieldRef:
              fieldPath: metadata.namespace
        - name: POD_IP
          valueFrom:
            fieldRef:
              fieldPath: status.podIP

Apply:

kubectl apply -f downward.yaml

Query from inside the container:

kubectl exec downward -- printenv POD_NAME
kubectl exec downward -- printenv POD_NAMESPACE
kubectl exec downward -- printenv POD_IP

The application learned runtime identity from the platform rather than from a baked image or hand-written config file.

Cleanup:

kubectl delete pod downward
rm -f downward.yaml

16. Probes: Running Is Not the Same as Healthy [CKAD]

A process can exist without being useful.

Kubernetes therefore asks several different health questions.

Readiness

Should this Pod receive traffic right now?

A failed readiness probe does not normally restart the container.

It makes the Pod unready so that Services can stop routing normal traffic to it.

Liveness

Is this container unhealthy enough that kubelet should restart it?

A failed liveness probe can restart the container.

Startup

Has this slow-starting application successfully started yet?

A startup probe protects slow applications from liveness/readiness behaviour until startup succeeds.

Think:

startup   -> have you started?
readiness -> should you receive traffic?
liveness  -> should the container be restarted?

Add readiness and liveness to nginx

Patch the Deployment:

kubectl patch deployment api --type=strategic -p '
spec:
  template:
    spec:
      containers:
        - name: nginx
          readinessProbe:
            httpGet:
              path: /
              port: 80
            initialDelaySeconds: 2
            periodSeconds: 5
          livenessProbe:
            httpGet:
              path: /
              port: 80
            initialDelaySeconds: 5
            periodSeconds: 10
'

Wait:

kubectl rollout status deployment/api

Inspect:

kubectl get pods -l app=api

Break readiness on purpose

Change readiness to a path that does not exist:

kubectl patch deployment api --type=strategic -p '
spec:
  template:
    spec:
      containers:
        - name: nginx
          readinessProbe:
            httpGet:
              path: /definitely-not-here
              port: 80
'

Watch:

kubectl get pods -l app=api -w

The Pods can be:

STATUS: Running

while showing:

READY: 0/1

That is an important distinction.

Running means the Pod's containers are running.

It does not mean Kubernetes considers the application ready for traffic.

Follow the consequence into networking

Inspect EndpointSlice readiness:

kubectl get endpointslices \
  -l kubernetes.io/service-name=api \
  -o yaml

Try the Service:

kubectl exec client -- \
  wget -T 2 -qO- http://api

The request should fail once all backends are unready.

This connects two ideas we learned separately:

readiness probe fails
      |
      v
Pod Ready=False
      |
      v
endpoint not ready for normal Service traffic
      |
      v
client loses a usable backend

Now the purpose of readiness is concrete rather than definitional.

Fix readiness

kubectl patch deployment api --type=strategic -p '
spec:
  template:
    spec:
      containers:
        - name: nginx
          readinessProbe:
            httpGet:
              path: /
              port: 80
'

Wait:

kubectl rollout status deployment/api

Verify:

kubectl exec client -- wget -qO- http://api

Part V - Containers Working Together

17. Multi-Container Pods [CKAD]

A Pod can contain multiple containers.

That does not mean a Pod should become a miniature virtual machine full of unrelated services.

Containers belong in the same Pod when they need very tight lifecycle, networking or storage coupling.

Containers in one Pod share the same network namespace, so they can talk over localhost.

They can also mount the same volumes.

Shared-volume experiment

Save as multi.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: multi
spec:
  containers:
    - name: writer
      image: busybox:1.36
      command:
        - sh
        - -c
        - |
          while true; do
            date > /data/index.html
            sleep 2
          done
      volumeMounts:
        - name: shared
          mountPath: /data

    - name: web
      image: nginx:1.27-alpine
      volumeMounts:
        - name: shared
          mountPath: /usr/share/nginx/html

  volumes:
    - name: shared
      emptyDir: {}

Apply:

kubectl apply -f multi.yaml

Wait:

kubectl wait \
  --for=condition=Ready \
  pod/multi \
  --timeout=60s

Inspect the containers:

kubectl get pod multi \
  -o jsonpath='{.spec.containers[*].name}{"\n"}'

Read the shared file from nginx:

kubectl exec multi -c web -- \
  cat /usr/share/nginx/html/index.html

Wait a few seconds and repeat.

The writer updates the file.

The web container sees the same volume.

Shared networking

From the writer container, call nginx over localhost:

kubectl exec multi -c writer -- \
  wget -qO- http://127.0.0.1

No Service is required between containers in the same Pod.

They already share a network namespace.

Container-specific logs and exec

With multiple containers, specify the container when needed:

kubectl logs multi -c writer
kubectl logs multi -c web
kubectl exec multi -c writer -- date

Cleanup:

kubectl delete pod multi
rm -f multi.yaml

18. Init Containers and Native Sidecars [CKAD] [DEV]

Different supporting containers solve different lifecycle problems.

The useful question is not:

How many containers can a Pod contain?

It is:

What lifecycle relationship does this supporting process have with the application?

Init container: do work before the application starts

Use an init container when setup must complete successfully before normal application containers begin.

Save as init-demo.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: init-demo
spec:
  initContainers:
    - name: prepare
      image: busybox:1.36
      command:
        - sh
        - -c
        - echo "generated by init container" > /data/index.html
      volumeMounts:
        - name: data
          mountPath: /data

  containers:
    - name: web
      image: nginx:1.27-alpine
      volumeMounts:
        - name: data
          mountPath: /usr/share/nginx/html

  volumes:
    - name: data
      emptyDir: {}

Apply:

kubectl apply -f init-demo.yaml

Inspect:

kubectl describe pod init-demo

Read the generated content:

kubectl exec init-demo -- \
  cat /usr/share/nginx/html/index.html

The sequence was:

init container starts
      |
      v
prepares shared data
      |
      v
init container completes
      |
      v
application container starts

Inspect the separate status lists:

kubectl get pod init-demo \
  -o jsonpath='{range .status.initContainerStatuses[*]}init:{.name}={.state.terminated.reason}{"\n"}{end}{range .status.containerStatuses[*]}app:{.name}={.state.running.startedAt}{"\n"}{end}'

The init container terminated successfully before nginx began its normal lifetime.

Cleanup:

kubectl delete pod init-demo
rm -f init-demo.yaml

Native sidecar: start in init ordering, then stay alive

Native sidecars are restartable init containers.

They use:

restartPolicy: Always

Unlike an ordinary init container, the sidecar does not need to finish before the application can keep running.

Let's prove that.

Save as sidecar-demo.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: sidecar-demo
spec:
  initContainers:
    - name: log-forwarder
      image: busybox:1.36
      restartPolicy: Always
      command:
        - sh
        - -c
        - |
          touch /logs/app.log
          tail -F /logs/app.log
      volumeMounts:
        - name: logs
          mountPath: /logs

  containers:
    - name: app
      image: busybox:1.36
      command:
        - sh
        - -c
        - |
          i=0
          while true; do
            i=$((i + 1))
            echo "application message $i" >> /logs/app.log
            sleep 2
          done
      volumeMounts:
        - name: logs
          mountPath: /logs

  volumes:
    - name: logs
      emptyDir: {}

Apply and wait:

kubectl apply -f sidecar-demo.yaml
kubectl wait \
  --for=condition=Ready \
  pod/sidecar-demo \
  --timeout=60s

Now read the sidecar logs:

kubectl logs sidecar-demo \
  -c log-forwarder \
  --tail=5

You should see the application messages even though the log-forwarder was declared under initContainers.

Inspect its state:

kubectl get pod sidecar-demo \
  -o jsonpath='{range .status.initContainerStatuses[*]}{.name}{" running="}{.state.running.startedAt}{" restarts="}{.restartCount}{"\n"}{end}'

The slightly surprising result is:

spec.initContainers
        |
        +-- ordinary init container -> eventually terminates
        |
        +-- restartPolicy: Always   -> remains running as a sidecar

That is why native sidecars can participate in init ordering while still living for the Pod lifetime.

Good uses include:

  • log forwarding
  • local proxying
  • configuration synchronisation
  • security helpers

Use a sidecar when the supporting functionality genuinely belongs to the same Pod lifecycle.

Do not group unrelated services into one Pod merely because Kubernetes allows multiple containers.

Cleanup:

kubectl delete pod sidecar-demo
rm -f sidecar-demo.yaml

Part VI - Finite, Stateful and Node-Scoped Workloads

19. Jobs and CronJobs [CKAD]

A Deployment represents work that should keep running.

Some work should finish.

That is the problem a Job solves.

Job: run finite work to completion

Create:

kubectl create job hello \
  --image=busybox:1.36 \
  -- echo hello-from-job

Inspect:

kubectl get jobs
kubectl get pods -l job-name=hello

Logs:

kubectl logs job/hello

Completion status:

kubectl get job hello \
  -o jsonpath='{.status.succeeded}{"\n"}'

Ownership:

kubectl get pods \
  -l job-name=hello \
  -o custom-columns='POD:.metadata.name,OWNER:.metadata.ownerReferences[0].name'

A Job controller is still reconciling desired state.

Its desired state is just different:

Deployment -> keep N replicas running
Job        -> achieve N successful completions

Cleanup:

kubectl delete job hello

CronJob: create Jobs on a schedule

Create:

kubectl create cronjob clock \
  --image=busybox:1.36 \
  --schedule='*/2 * * * *' \
  -- date

Inspect:

kubectl get cronjobs

Rather than waiting, manually create a Job from its template:

kubectl create job \
  --from=cronjob/clock \
  clock-now

Follow the ownership model:

CronJob
   |
   v
Job
   |
   v
Pod

Logs:

kubectl logs job/clock-now

Useful CronJob controls include:

schedule
suspend
concurrencyPolicy
startingDeadlineSeconds
successfulJobsHistoryLimit
failedJobsHistoryLimit

Discover them rather than guessing:

kubectl explain cronjob.spec

Cleanup:

kubectl delete cronjob clock
kubectl delete job clock-now

20. Storage: Pod Lifetime vs Data Lifetime [CKAD]

Containers and Pods are replaceable.

Data is often not.

Kubernetes therefore separates workload lifetime from storage lifetime.

emptyDir: data that belongs to a Pod

We already used emptyDir in the multi-container example.

Its lifetime is tied to the Pod:

container restart
      |
      v
emptyDir remains

Pod deletion
      |
      v
emptyDir disappears

This makes it useful for:

  • scratch space
  • caches
  • sharing files between containers in one Pod

It is not durable application storage.

PersistentVolumeClaim: ask the cluster for durable storage

A Pod should not usually care whether storage is backed by EBS, Ceph, local disks, NFS or another implementation.

It asks for storage through a PersistentVolumeClaim.

First inspect StorageClasses:

kubectl get storageclass

Save as pvc.yaml:

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: data
spec:
  accessModes:
    - ReadWriteOnce
  resources:
    requests:
      storage: 1Gi

Apply:

kubectl apply -f pvc.yaml

Inspect:

kubectl get pvc data

If your cluster has a default dynamic provisioner, the claim should become Bound.

If it remains Pending, inspect:

kubectl describe pvc data

The cluster may not have a default StorageClass or suitable PersistentVolume.

That is a platform capability issue rather than a YAML syntax issue.

Prove that data can outlive a Pod

Assuming the claim is Bound, create a writer Pod.

Save as pvc-writer.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: pvc-writer
spec:
  containers:
    - name: writer
      image: busybox:1.36
      command: ["sh", "-c", "echo persistent-data > /data/message; sleep 3600"]
      volumeMounts:
        - name: data
          mountPath: /data
  volumes:
    - name: data
      persistentVolumeClaim:
        claimName: data

Apply:

kubectl apply -f pvc-writer.yaml

Wait:

kubectl wait \
  --for=condition=Ready \
  pod/pvc-writer \
  --timeout=60s

Check:

kubectl exec pvc-writer -- cat /data/message

Delete the Pod:

kubectl delete pod pvc-writer

Now create a different Pod using the same claim.

Save as pvc-reader.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: pvc-reader
spec:
  containers:
    - name: reader
      image: busybox:1.36
      command: ["sh", "-c", "sleep 3600"]
      volumeMounts:
        - name: data
          mountPath: /data
  volumes:
    - name: data
      persistentVolumeClaim:
        claimName: data

Apply:

kubectl apply -f pvc-reader.yaml

Read:

kubectl exec pvc-reader -- cat /data/message

Expected:

persistent-data

The first Pod is gone.

The claim remained.

That is the useful boundary:

Pod lifetime != persistent data lifetime

Cleanup:

kubectl delete pod pvc-reader --ignore-not-found
kubectl delete pvc data
rm -f pvc.yaml pvc-writer.yaml pvc-reader.yaml

21. StatefulSets: Stable Replica Identity [CKAD]

Deployments intentionally treat replicas as interchangeable.

Sometimes the application cares which replica is which.

Databases, clustered systems and ordered members may need stable identity.

That is the problem StatefulSets solve.

A StatefulSet can provide:

  • stable Pod names
  • ordered creation and termination
  • stable network identity when paired with a headless Service
  • per-replica persistent storage through volume claim templates

Small identity experiment

Save as stateful.yaml:

apiVersion: v1
kind: Service
metadata:
  name: stateful-web
spec:
  clusterIP: None
  selector:
    app: stateful-web
  ports:
    - port: 80
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: stateful-web
spec:
  serviceName: stateful-web
  replicas: 3
  selector:
    matchLabels:
      app: stateful-web
  template:
    metadata:
      labels:
        app: stateful-web
    spec:
      containers:
        - name: nginx
          image: nginx:1.27-alpine

Apply:

kubectl apply -f stateful.yaml

Observe:

kubectl get statefulsets
kubectl get pods -l app=stateful-web

Pod names are predictable:

stateful-web-0
stateful-web-1
stateful-web-2

Delete one:

kubectl delete pod stateful-web-1

Observe:

kubectl get pods -l app=stateful-web -w

The replacement is still:

stateful-web-1

Compare with Deployment-generated Pod names.

The point is not that StatefulSet Pods are immortal.

They are still replaceable.

The point is that their identity is stable across replacement.

Cleanup:

kubectl delete -f stateful.yaml
rm -f stateful.yaml

22. DaemonSets: One Workload Per Matching Node [CKAD]

A Deployment asks for an arbitrary number of replicas.

Some software instead needs to run on every relevant node.

Examples include:

  • CNI agents
  • storage agents
  • log collectors
  • node monitoring

That is what a DaemonSet expresses.

Save as daemonset.yaml:

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-demo
spec:
  selector:
    matchLabels:
      app: node-demo
  template:
    metadata:
      labels:
        app: node-demo
    spec:
      containers:
        - name: sleeper
          image: busybox:1.36
          command: ["sh", "-c", "sleep 3600"]

Apply:

kubectl apply -f daemonset.yaml

Observe:

kubectl get daemonset node-demo
kubectl get pods -l app=node-demo -o wide
kubectl get nodes

On a simple untainted lab cluster you will often see roughly one Pod per eligible node.

The important distinction is:

Deployment
    -> desired replica count

DaemonSet
    -> desired eligible-node coverage

Cleanup:

kubectl delete daemonset node-demo
rm -f daemonset.yaml

Part VII - Identity, Authorization and Runtime Security

23. Who Is Allowed to Do What? Users, ServiceAccounts, RBAC and Admission [CKAD] [DEV]

Until now we have mostly used kubectl as a highly privileged lab administrator.

That is useful for learning, but it hides an important production question:

Who should be allowed to do what?

There are several separate decisions in the Kubernetes API request path.

request
   |
   v
authentication
   |
   | who are you?
   v
authorization
   |
   | may that identity perform this action?
   v
admission
   |
   | for writes: is this object acceptable, or should it be mutated?
   v
Kubernetes API state

Keeping those stages separate avoids a lot of security confusion.

Humans and workloads use different kinds of identity

Kubernetes commonly deals with two broad identity types:

human / external client          workload inside Kubernetes
          |                                |
          v                                v
     User / Group                    ServiceAccount
          |                                |
          +---------------+----------------+
                          |
                          v
                         RBAC

A ServiceAccount is a Kubernetes API object.

A normal human User is not.

Kubernetes does not provide a User resource that you create with:

kubectl create user alice

Instead, human authentication normally comes from something outside the Kubernetes object model, such as:

client certificate
OIDC / SSO identity
cloud IAM integration
authentication proxy

After authentication, the API server has identity information such as:

username: alice
groups:
  - developers

RBAC then decides what that identity may do.

Lab: give Alice namespace-scoped developer access

We do not need to configure a real identity provider just to learn authorization.

kubectl can ask the API server to evaluate a request as another identity using impersonation.

--as=alice does not create Alice. It asks the API server to evaluate the request as the username alice. Your current identity must itself be allowed to impersonate users. The administrator credentials used by our kind lab normally are.

Create a namespace-scoped developer role.

Save as alice-rbac.yaml:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: developer
  namespace: cookbook
rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["apps"]
    resources: ["deployments"]
    verbs: ["get", "list", "watch", "create", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: alice-developer
  namespace: cookbook
subjects:
  - kind: User
    name: alice
    apiGroup: rbac.authorization.k8s.io
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: developer

Apply it:

kubectl apply -f alice-rbac.yaml

Ask whether Alice can read Pods in our namespace:

kubectl auth can-i list pods \
  --as=alice \
  -n cookbook

Expected:

yes

Can she change a Deployment?

kubectl auth can-i patch deployments \
  --as=alice \
  -n cookbook

Expected:

yes

Can she read Secrets?

kubectl auth can-i get secrets \
  --as=alice \
  -n cookbook

Expected:

no

Can she administer another namespace?

kubectl auth can-i patch deployments \
  --as=alice \
  -n default

Expected:

no

Can she delete a cluster-scoped Node object?

kubectl auth can-i delete nodes \
  --as=alice

Expected:

no

This is the permission boundary we wanted:

Alice
  |
  v
Kubernetes API
  |
  +-- read Pods in cookbook             yes
  +-- change Deployments in cookbook    yes
  +-- read Secrets in cookbook          no
  +-- change Deployments in default     no
  +-- delete Nodes                       no

You can ask for a broader view of the permissions Kubernetes calculates:

kubectl auth can-i --list \
  --as=alice \
  -n cookbook

RBAC permissions are additive

Kubernetes RBAC grants permissions.

It does not contain explicit deny rules.

Think:

matching allow rule exists
        -> allowed

no matching allow rule
        -> not allowed

That means you need to consider all RoleBindings and ClusterRoleBindings attached to an identity.

A narrow RoleBinding does not protect Alice if some other binding also gives her broad cluster permissions.

A permission can have indirect effects

Alice cannot directly create Pods with the Role above.

Check:

kubectl auth can-i create pods \
  --as=alice \
  -n cookbook

Expected:

no

But Alice can create a Deployment.

A Deployment controller can then create ReplicaSets and Pods on her behalf.

Alice
  |
  | create Deployment allowed
  v
Deployment
  |
  v
Deployment controller
  |
  v
ReplicaSet
  |
  v
Pods

Authorization is evaluated against the API request Alice makes.

You therefore need to reason about what a permitted object can cause controllers to do, not merely about the object's name.

This becomes particularly important with powerful workload features such as privileged containers, host mounts and scheduling controls.

ServiceAccount: workload identity

Now do the same exercise for an application rather than a human.

Create a ServiceAccount:

kubectl create serviceaccount api-sa

Inspect it:

kubectl get serviceaccount api-sa -o yaml

Modern Kubernetes normally gives Pods short-lived projected ServiceAccount credentials rather than relying on automatically created permanent token Secrets.

Request a temporary token when the cluster allows it:

kubectl create token api-sa

Assign the ServiceAccount to our Deployment:

kubectl set serviceaccount deployment/api api-sa

Wait:

kubectl rollout status deployment/api

Check:

kubectl get deployment api \
  -o jsonpath='{.spec.template.spec.serviceAccountName}{"\n"}'

Expected:

api-sa

Give that workload identity read-only Pod access.

Save as workload-rbac.yaml:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: pod-reader
  namespace: cookbook
rules:
  - apiGroups: [""]
    resources: ["pods"]
    verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: api-sa-pod-reader
  namespace: cookbook
subjects:
  - kind: ServiceAccount
    name: api-sa
    namespace: cookbook
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: pod-reader

Apply:

kubectl apply -f workload-rbac.yaml

Ask whether that identity may list Pods:

kubectl auth can-i list pods \
  --as=system:serviceaccount:cookbook:api-sa \
  -n cookbook

Expected:

yes

Ask whether it may delete them:

kubectl auth can-i delete pods \
  --as=system:serviceaccount:cookbook:api-sa \
  -n cookbook

Expected:

no

Human and workload authorization now look almost identical after authentication:

User/alice --------------------+
                               |
ServiceAccount/api-sa ---------+
                               |
                               v
                              RBAC
                               |
                         allowed verbs
                         on resources
                         in a scope

Role, ClusterRole, RoleBinding and ClusterRoleBinding

RBAC has four main API objects.

Role
  -> permission rules defined for one namespace

ClusterRole
  -> reusable permission rules
  -> can also describe cluster-scoped resources

RoleBinding
  -> grants a Role or ClusterRole inside one namespace

ClusterRoleBinding
  -> grants a ClusterRole across the cluster

A useful relationship is:

subject
  |
  | User / Group / ServiceAccount
  v
binding
  |
  v
role containing rules
  |
  v
verbs + resources

For example:

alice
  |
  v
RoleBinding/cookbook
  |
  v
Role/developer
  |
  +-- get/list/watch Pods
  +-- create/update/patch Deployments

Be especially careful with ClusterRoleBinding.

This:

Alice can administer one namespace

and this:

Alice can administer the whole cluster

can differ by only the binding used.

A namespace is a scope, not an automatic security boundary

Namespaces are extremely useful administrative boundaries.

But merely placing two teams in different namespaces does not automatically isolate them.

namespace alone
    !=
permission boundary

You normally combine namespaces with controls such as:

RBAC
ResourceQuota / LimitRange
NetworkPolicy
Pod security / admission policy
storage policy

The exact isolation you need depends on whether the tenants trust one another.

API permission, workload placement and machine access are different controls

This distinction matters particularly around control-plane nodes.

There are at least three independent questions:

1. API authorization
   "Can Alice delete or modify this Node object?"

2. workload placement
   "Can this Pod be scheduled onto this node?"

3. machine access
   "Can Alice SSH into or otherwise administer the actual host?"

They are enforced by different layers:

                         control-plane machine
                                  |
          +-----------------------+-----------------------+
          |                       |                       |
          v                       v                       v
   Kubernetes API            scheduler target          Linux / VM / metal
          |                       |                       |
         RBAC              taints / tolerations       IAM / SSH / firewall
                          affinity / selectors        OS permissions

For example, control-plane nodes are commonly tainted so ordinary workloads do not schedule there:

node-role.kubernetes.io/control-plane:NoSchedule

That is a scheduling control.

It is not the same thing as denying API access to Node objects, and it is not the same thing as denying SSH access to the machine.

It is also not, by itself, a strong tenant security boundary: a workload that is allowed to specify a matching toleration can become eligible for the node.

Likewise:

kubectl auth can-i delete nodes --as=alice

answers an API authorization question.

It tells you nothing about whether Alice has infrastructure credentials for the underlying VM or bare-metal host.

What about the kubelet itself? [DEV] [DEEP DIVE]

Nodes are API clients too.

A kubelet commonly authenticates with an identity resembling:

system:node:worker-01

Kubernetes has a special-purpose Node authorizer that can constrain kubelet API access based on the Pods assigned to that node.

The NodeRestriction admission plugin adds additional restrictions around what kubelets may modify.

Conceptually:

human / application identities
        -> RBAC

kubelet node identities
        -> Node authorizer
        -> NodeRestriction admission

You do not need to configure these for CKAD, but knowing that node identity has its own authorization path prevents the misleading idea that every Kubernetes permission problem is just a RoleBinding.

RBAC is not the same thing as multi-tenancy [DEV] [DEEP DIVE]

RBAC can give multiple teams restricted access to one Kubernetes API:

Alice ----+
          |
Bob ------+--> one kube-apiserver
          |        |
          |       RBAC
          |        |
          +--> namespace-scoped views

That can be entirely appropriate for trusted teams.

But stronger tenancy may instead give each tenant its own Kubernetes API/control-plane boundary:

Alice ---> tenant A API

Bob -----> tenant B API

                |
                v
        provider infrastructure

Projects such as vCluster and Kamaji operate in this design space, although they implement it differently.

RBAC still exists inside each tenant cluster. The difference is that the tenant boundary no longer depends only on permissions inside one shared API server.

The companion multitenancy-appendix.md continues this model and compares shared-cluster RBAC, vCluster and Kamaji without turning the CKAD path into a platform-engineering course.

Admission control

Authorization is not necessarily the final decision for a write request.

After authentication and authorization, admission can validate or mutate an incoming object before it is persisted.

Examples include:

  • Pod security requirements
  • quotas
  • policy rules
  • injected defaults
  • validating or mutating webhooks

This explains a useful failure class:

I am authenticated
      |
      v
I am authorized
      |
      v
write request still rejected
      |
      v
check admission or policy error

One subtle distinction: normal read operations such as get, list and watch do not pass through admission control in the same way write requests do.

The permission model to keep

When something is denied, ask which boundary you are actually debugging:

Who am I?
  -> authentication

May I make this API request?
  -> authorization / RBAC

Is this write acceptable?
  -> admission

May this workload land on that node?
  -> scheduling controls

May this workload talk to another workload?
  -> NetworkPolicy / network controls

May this person administer the actual machine?
  -> infrastructure IAM / SSH / OS controls

Those controls cooperate, but they are not substitutes for one another.

Cleanup the lab RBAC objects, but keep the ServiceAccount because the Deployment currently uses it:

kubectl delete -f alice-rbac.yaml
kubectl delete -f workload-rbac.yaml
rm -f alice-rbac.yaml workload-rbac.yaml

24. SecurityContext and Container Privilege [CKAD]

Container runtime security should be part of the workload definition rather than an undocumented node-side convention.

A securityContext can exist at Pod level and container level.

Before building a secure Pod, deliberately ask Kubernetes for an impossible combination.

Break it: require non-root without choosing a non-root user

Save as root-forbidden.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: root-forbidden
spec:
  securityContext:
    runAsNonRoot: true

  containers:
    - name: shell
      image: busybox:1.36
      command: ["sh", "-c", "sleep 3600"]

Apply:

kubectl apply -f root-forbidden.yaml

Inspect:

kubectl get pod root-forbidden
kubectl describe pod root-forbidden

The image normally runs as UID 0, but the Pod says that root is forbidden.

The kubelet therefore cannot construct the requested container safely.

You should see a failure such as:

CreateContainerConfigError

with an Event explaining that runAsNonRoot conflicts with a root runtime identity.

This is a useful distinction:

image says
run as root
    |
    X
Pod securityContext says
must not run as root

Kubernetes did not silently weaken the requested security policy to make the container start.

Delete the failed Pod:

kubectl delete pod root-forbidden
rm -f root-forbidden.yaml

Fix it: choose the runtime identity deliberately

Save as secure-demo.yaml:

apiVersion: v1
kind: Pod
metadata:
  name: secure-demo
spec:
  securityContext:
    runAsNonRoot: true
    seccompProfile:
      type: RuntimeDefault

  containers:
    - name: shell
      image: busybox:1.36
      command: ["sh", "-c", "sleep 3600"]
      securityContext:
        runAsUser: 10001
        allowPrivilegeEscalation: false
        readOnlyRootFilesystem: true
        capabilities:
          drop:
            - ALL

Apply and wait:

kubectl apply -f secure-demo.yaml
kubectl wait \
  --for=condition=Ready \
  pod/secure-demo \
  --timeout=60s

Inspect identity:

kubectl exec secure-demo -- id

You should see UID 10001 rather than root.

Prove the root filesystem is read-only:

kubectl exec secure-demo -- \
  sh -c 'touch /tmp/should-fail'

The write should fail.

Inspect the configured controls:

kubectl get pod secure-demo \
  -o jsonpath='{.spec.securityContext}{"\n"}{.spec.containers[0].securityContext}{"\n"}'

Useful controls include:

runAsNonRoot
runAsUser
runAsGroup
fsGroup
allowPrivilegeEscalation
readOnlyRootFilesystem
capabilities
seccompProfile

Do not treat these as synonyms.

For example:

runAsNonRoot
    -> refuse a root runtime identity

runAsUser
    -> choose a numeric runtime UID

allowPrivilegeEscalation: false
    -> process cannot gain more privileges than its parent

capabilities.drop
    -> remove specific Linux capabilities

readOnlyRootFilesystem
    -> make the container root filesystem read-only

seccompProfile: RuntimeDefault
    -> apply the runtime's default syscall filter

The broader lesson is the same as elsewhere in Kubernetes:

security intent in spec
        |
        v
runtime tries to satisfy it
        |
        +-- possible   -> container runs
        |
        +-- impossible -> visible failure

Cleanup:

kubectl delete pod secure-demo
rm -f secure-demo.yaml

Part VIII - Network Policy and External HTTP Routing

25. NetworkPolicy: Which Pods May Talk? [CKAD]

A Service answers:

Where should traffic go?

A NetworkPolicy answers a different question:

Which network flows should be allowed?

NetworkPolicy enforcement depends on the cluster's CNI plugin.

If your CNI does not enforce NetworkPolicy, the objects can exist without changing packet flow.

Confirm current connectivity

kubectl exec client -- wget -qO- http://api

Deny ingress to the API Pods

Save as deny-api.yaml:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: deny-api
spec:
  podSelector:
    matchLabels:
      app: api
  policyTypes:
    - Ingress

Apply:

kubectl apply -f deny-api.yaml

If your CNI enforces policy, this should eventually fail:

kubectl exec client -- \
  wget -T 2 -qO- http://api

Why?

The selected API Pods now have ingress isolation, but no ingress rule permits the client.

Allow only labelled clients

Label the client:

kubectl label pod client access=api

Replace the policy with allow-api.yaml:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-api
spec:
  podSelector:
    matchLabels:
      app: api
  policyTypes:
    - Ingress
  ingress:
    - from:
        - podSelector:
            matchLabels:
              access: api
      ports:
        - protocol: TCP
          port: 80

Apply:

kubectl delete networkpolicy deny-api
kubectl apply -f allow-api.yaml

Retry:

kubectl exec client -- wget -qO- http://api

This is a selector story again:

NetworkPolicy podSelector
    -> which destination Pods are governed?

from.podSelector
    -> which source Pods are allowed?

Cleanup:

kubectl delete networkpolicy allow-api --ignore-not-found
kubectl label pod client access-
rm -f deny-api.yaml allow-api.yaml

26. Ingress: The Frozen HTTP API and the Road to Gateway API [CKAD] [DEV]

A ClusterIP Service gives an application a stable endpoint inside the cluster.

Users outside the cluster often need HTTP routing to that Service.

Historically, Kubernetes modelled that with Ingress.

An important distinction is:

An Ingress object is configuration. An Ingress controller is the software that implements it.

That distinction is worth observing directly.

Check whether anything implements Ingress

Ask the cluster for its available implementations:

kubectl get ingressclass

A stock kind cluster may return no classes at all.

That means the API server understands Ingress, but nothing is currently responsible for turning those objects into working HTTP listeners.

This is the same pattern we will later see with CRDs and controllers:

API object exists
      |
      v
controller watches it
      |
      v
real behaviour appears

No controller means the middle step is missing.

Create the API object anyway

Save as ingress.yaml:

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: api
spec:
  rules:
    - host: api.cookbook.local
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: api
                port:
                  number: 80

Apply:

kubectl apply -f ingress.yaml

Inspect:

kubectl get ingress api
kubectl describe ingress api

Inspect status directly:

kubectl get ingress api \
  -o jsonpath='{.status.loadBalancer.ingress}{"\n"}'

If there is no controller, that status will normally remain empty and no HTTP listener will magically appear.

That is useful behaviour to understand:

Ingress stored in API       yes
routing implementation      no
external traffic path       no

If your cluster already has a maintained Ingress controller, set the matching class:

spec:
  ingressClassName: <class-name>

and follow that controller's documented exposure method.

The abstract path is still:

HTTP request
   |
   v
Ingress controller
   |
   | host/path rule
   v
Service
   |
   v
ready backend Pods

Ingress does not replace a Service.

It routes to one.

Why we do not install an Ingress controller just for this lab

The Kubernetes project now recommends Gateway API instead of Ingress.

Ingress remains a stable API and is not being removed, but the API is frozen and no longer gaining features.

Also avoid old tutorials that tell you to install ingress-nginx: that project was retired in March 2026 and no longer receives fixes or security updates.

So the main cookbook keeps Ingress because it is still important Kubernetes and CKAD knowledge, but we do not introduce a legacy controller merely to make this one exercise route traffic.

Instead, the networking deep dive takes the modern path.

Continue with cillium-gateay-appendix.md, where we actually install and exercise Cilium's Gateway API implementation:

GatewayClass
     |
     v
Gateway
     |
     v
HTTPRoute
     |
     v
Service
     |
     v
Pod

That appendix sends real traffic, breaks backend references, inspects status, exercises header routing, weighted backends, cross-namespace ReferenceGrant, and follows the implementation down through Envoy, Cilium and eBPF.

The important progression is:

Ingress
  -> simple, stable, frozen HTTP routing API

Gateway API
  -> role-oriented, extensible service-networking APIs

For a broader tour of the Gateway API resource model and its routing features, Roman Glushko's deep dive is also excellent:

https://www.romaglushko.com/blog/k8s-gateway-api/

Cleanup:

kubectl delete ingress api
rm -f ingress.yaml

Part IX - Deployment Strategies and Packaging

27. Blue-Green and Canary Deployments [CKAD]

A Deployment's rolling update strategy is not the only way to release software.

Two common patterns are blue-green and canary.

The interesting Kubernetes lesson is that both can be built from primitives we already understand: Deployments, labels and Services.

Blue-green: switch the Service selector

Create two versions with explicit Pod labels.

Save as blue-green.yaml:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: shop-blue
spec:
  replicas: 2
  selector:
    matchLabels:
      app: shop
      release: blue
  template:
    metadata:
      labels:
        app: shop
        release: blue
    spec:
      containers:
        - name: nginx
          image: nginx:1.27-alpine
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: shop-green
spec:
  replicas: 2
  selector:
    matchLabels:
      app: shop
      release: green
  template:
    metadata:
      labels:
        app: shop
        release: green
    spec:
      containers:
        - name: nginx
          image: nginx:1.28-alpine
---
apiVersion: v1
kind: Service
metadata:
  name: shop
spec:
  selector:
    app: shop
    release: blue
  ports:
    - port: 80
      targetPort: 80

Apply:

kubectl apply -f blue-green.yaml

See which Pods are selected:

kubectl get pods -l app=shop --show-labels
kubectl get endpointslices \
  -l kubernetes.io/service-name=shop

Initially the Service selects blue.

Switch to green:

kubectl patch service shop \
  --type=merge \
  -p '{"spec":{"selector":{"app":"shop","release":"green"}}}'

Inspect endpoints again:

kubectl get endpointslices \
  -l kubernetes.io/service-name=shop

The release switch was a selector change.

That is blue-green in its simplest Kubernetes form.

Cleanup:

kubectl delete -f blue-green.yaml
rm -f blue-green.yaml

Canary: two versions behind one Service

A crude Kubernetes-only canary can run two Deployments with the same Service selector.

Conceptually:

stable Deployment: 9 replicas
canary Deployment: 1 replica

both Pods:
  app=shop

Service selector:
  app=shop

Roughly one tenth of the available backend Pods would be canary Pods.

This is not precise request weighting.

A plain Service does not promise exact percentages, user affinity, header matching or request-level policy.

Those features generally require an ingress controller, gateway, service mesh or another higher-level traffic-routing system.

The lesson is to understand what the primitive actually guarantees rather than reading more into it.


28. Helm: Package Kubernetes Resources [CKAD]

Raw YAML is useful, but real applications often contain many related resources and environment-specific values.

Helm packages Kubernetes resources into a Chart.

Think:

Chart templates + values
          |
          v
rendered Kubernetes manifests
          |
          v
Kubernetes API

Helm is not a replacement for Kubernetes objects.

It generates and manages them.

Create a local chart

Assuming the helm CLI is installed:

helm create cookbook-api

Inspect:

find cookbook-api -maxdepth 2 -type f

Render before installing

helm template test cookbook-api

Override a value:

helm template test cookbook-api \
  --set replicaCount=2

Rendering is a powerful debugging technique because it lets you inspect the actual Kubernetes manifests before they reach the API server.

Install

helm install cookbook-example cookbook-api

Inspect:

helm list
kubectl get all -l app.kubernetes.io/instance=cookbook-example

Change a value through an upgrade:

helm upgrade cookbook-example cookbook-api \
  --set replicaCount=2

History:

helm history cookbook-example

Uninstall:

helm uninstall cookbook-example
rm -rf cookbook-api

A useful failure workflow is:

values
  |
  v
helm template
  |
  v
inspect rendered YAML
  |
  v
kubectl explain / server validation

29. Kustomize: Modify YAML Without a Template Language [CKAD]

Kustomize solves a different packaging problem.

Instead of templating YAML, it starts with ordinary Kubernetes manifests and layers transformations over them.

Think:

base resources
      +
overlay changes
      |
      v
rendered Kubernetes manifests

kubectl has built-in Kustomize support.

Create a base

mkdir -p kustomize/base
mkdir -p kustomize/overlays/dev

Generate a Deployment:

kubectl create deployment k-api \
  --image=nginx:1.27-alpine \
  --dry-run=client \
  -o yaml > kustomize/base/deployment.yaml

Create kustomize/base/kustomization.yaml:

resources:
  - deployment.yaml

Create an overlay

Create kustomize/overlays/dev/kustomization.yaml:

resources:
  - ../../base

namePrefix: dev-

replicas:
  - name: k-api
    count: 2

Render:

kubectl kustomize kustomize/overlays/dev

Apply:

kubectl apply -k kustomize/overlays/dev

Inspect:

kubectl get deployment dev-k-api

Cleanup:

kubectl delete -k kustomize/overlays/dev
rm -rf kustomize

A useful distinction:

Helm
    -> package + template/value system + release management

Kustomize
    -> transform/compose ordinary Kubernetes YAML

They can coexist in real systems.


Part X - Observability and Debugging

30. API Versions and Deprecation [CKAD]

Kubernetes evolves.

Old manifests found in blogs and repositories can reference API versions the current cluster no longer serves.

Do not blindly reuse them.

Ask the cluster.

Supported resources:

kubectl api-resources

Supported API versions:

kubectl api-versions

Deployment schema:

kubectl explain deployment

Deployment strategy:

kubectl explain deployment.spec.strategy

Ingress paths:

kubectl explain ingress.spec.rules.http.paths

Validate before changing the cluster

Client-side dry run checks local construction:

kubectl apply \
  --dry-run=client \
  -f manifest.yaml

Server-side dry run asks the API server to process the request without persisting it:

kubectl apply \
  --dry-run=server \
  -f manifest.yaml

Server-side validation is especially useful when you want current cluster schema and admission behaviour to participate.

A reliable workflow is:

old example found online
      |
      v
kubectl api-resources / api-versions
      |
      v
kubectl explain
      |
      v
server-side dry run
      |
      v
apply

31. A Systematic Debugging Workflow [CKAD] [DEV]

Random commands make Kubernetes feel mysterious.

Most failures become easier when you follow the relationship between objects.

Workload failure

Follow ownership downward:

Deployment
    |
    v
ReplicaSet
    |
    v
Pod
    |
    v
Container
    |
    v
Process

Commands:

kubectl get deployment api
kubectl get rs -l app=api
kubectl get pods -l app=api
kubectl describe pod <pod>
kubectl logs <pod>

Previous crashed container instance:

kubectl logs <pod> --previous

Events:

kubectl get events \
  --sort-by=.metadata.creationTimestamp

Questions to ask in order:

Did the controller create the expected child resource?
Did the Pod schedule?
Did the image pull?
Did the container start?
Did the process stay alive?
Did readiness succeed?

Networking failure

Follow the request path:

DNS
 |
 v
Service
 |
 v
selector
 |
 v
EndpointSlice
 |
 v
ready Pod
 |
 v
container port
 |
 v
process

Useful commands:

kubectl exec client -- nslookup api
kubectl get service api -o yaml
kubectl get pods -l app=api --show-labels
kubectl get endpointslices \
  -l kubernetes.io/service-name=api
kubectl describe pod <pod>

Then test the application directly from inside the cluster when possible.

Configuration failure

Follow references:

Pod spec
  |
  +-- ConfigMap name/key
  |
  +-- Secret name/key
  |
  +-- volume name
  |
  +-- mount path

Useful commands:

kubectl describe pod <pod>
kubectl get configmap <name> -o yaml
kubectl get secret <name> -o yaml

A misspelled Secret key can prevent a container from starting even though the Secret object itself exists.

Authorization failure

Ask Kubernetes instead of guessing:

kubectl auth can-i get pods
kubectl auth can-i create deployments

For another identity, when impersonation is permitted:

kubectl auth can-i list pods \
  --as=system:serviceaccount:cookbook:api-sa

The general debugging habit is:

Follow the API relationships until you find the first place where observed state stops matching your expectation.


32. Debugging Minimal Containers with Ephemeral Containers [DEV]

Production images may intentionally not contain:

bash
curl
dig
tcpdump
ps

That is often desirable.

A production image does not need a complete incident-response toolbox merely to make debugging convenient.

Let's hit that problem before solving it.

First try ordinary exec

Pick one API Pod:

POD=$(kubectl get pod -l app=api \
  -o jsonpath='{.items[0].metadata.name}')

CONTAINER=$(kubectl get pod "$POD" \
  -o jsonpath='{.spec.containers[0].name}')

printf 'pod=%s container=%s\n' "$POD" "$CONTAINER"

Try to use curl inside the application container:

kubectl exec "$POD" -c "$CONTAINER" -- \
  curl -s http://127.0.0.1

Our nginx-based image does not normally contain curl, so the command should fail with an executable-not-found error.

That is not necessarily an image defect.

It can be a deliberate production-image choice.

Add temporary tooling instead

Attach an ephemeral debugging container:

kubectl debug -it "$POD" \
  --image=busybox:1.36 \
  --target="$CONTAINER" \
  -- sh

Inside the debug container, call the application over the Pod's shared network namespace:

wget -qO- http://127.0.0.1

You can also inspect processes:

ps

Then exit:

exit

Inspect what Kubernetes added:

kubectl get pod "$POD" \
  -o jsonpath='{range .spec.ephemeralContainers[*]}{.name}{" image="}{.image}{" target="}{.targetContainerName}{"\n"}{end}'

The original application image did not change.

The Pod now has temporary debugging tooling attached to it.

Conceptually:

minimal production container
        |
        | exec lacks tooling
        v
   debugging blocked
        |
        | kubectl debug
        v
ephemeral container joins Pod
        |
        +-- same Pod network
        +-- optional process targeting
        +-- extra tools

Ephemeral containers are intentionally different from normal application containers:

  • they are added to an existing Pod for troubleshooting
  • they are not part of the normal workload template
  • they do not restart like ordinary workload containers
  • they cannot define normal container resources such as ports or probes

Depending on the runtime and security configuration, process visibility and debugging capabilities can differ.

The model to keep is:

production container stays minimal
        +
temporary debug tooling when needed

rather than:

ship every debugging utility in every production image forever

Part XI - Extending Kubernetes

33. CRDs and Custom Resources [CKAD] [DEV]

Kubernetes' built-in API contains types such as:

Pod
Deployment
Service
Secret
Job

But platform teams often want domain-specific APIs.

For example:

kind: PreviewEnvironment
spec:
  image: nginx:1.27-alpine
  replicas: 2

A CustomResourceDefinition (CRD) teaches the Kubernetes API server about a new resource type.

This exercise requires permission to create cluster-scoped CRDs.

Create a CRD

Save as preview-crd.yaml:

apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: previewenvironments.platform.example.com
spec:
  group: platform.example.com
  scope: Namespaced

  names:
    plural: previewenvironments
    singular: previewenvironment
    kind: PreviewEnvironment
    shortNames:
      - preview

  versions:
    - name: v1alpha1
      served: true
      storage: true

      subresources:
        status: {}

      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                image:
                  type: string
                replicas:
                  type: integer
                  minimum: 1
              required:
                - image
                - replicas

            status:
              type: object
              properties:
                readyReplicas:
                  type: integer
                url:
                  type: string

Apply:

kubectl apply -f preview-crd.yaml

Discover it:

kubectl api-resources | grep -i preview

Ask for its schema:

kubectl explain previewenvironments
kubectl explain previewenvironments.spec
kubectl explain previewenvironments.status

The CRD extended API discovery just like a built-in type.

The status subresource also gives a future controller somewhere separate to report observed state without pretending that status is user intent.

Create a Custom Resource

Save as preview.yaml:

apiVersion: platform.example.com/v1alpha1
kind: PreviewEnvironment
metadata:
  name: pr-482
spec:
  image: nginx:1.27-alpine
  replicas: 2

Apply:

kubectl apply -f preview.yaml

Query:

kubectl get previewenvironments

Or use the short name:

kubectl get preview

Inspect:

kubectl get preview pr-482 -o yaml

Now notice what did not happen.

There are no application Pods for pr-482.

Why?

CRD
 |
 v
Kubernetes understands the data type

but

No controller
 |
 v
Nothing implements its behaviour

This distinction is fundamental:

A CRD extends the API. It does not, by itself, implement a control loop.

Leave preview-crd.yaml and preview.yaml in place.

The next chapter gives them behaviour.


34. Build a Tiny Go Controller [DEV] [DEEP DIVE]

Chapter 1 told us that Kubernetes is a control system.

Now we are going to write one of those control loops ourselves.

Our custom API says:

kind: PreviewEnvironment
spec:
  image: nginx:1.27-alpine
  replicas: 2

We want that to cause:

PreviewEnvironment
       |
       +-- Deployment
       |
       +-- Service

and we want the custom resource to report:

status:
  readyReplicas: 2
  url: http://pr-482.cookbook.svc.cluster.local

That requires a controller.

The smallest useful reconciliation loop

A production controller normally uses watches, informers, a work queue, retries and often leader election.

We are deliberately starting with something smaller:

every 2 seconds
     |
     v
list PreviewEnvironments
     |
     v
for each one
     |
     +-- apply desired Deployment
     +-- apply desired Service
     +-- read observed Deployment status
     +-- update PreviewEnvironment status

Polling is not the architecture we would choose for a serious controller.

It is useful here because the reconciliation logic stays visible.

The controller uses server-side apply for its child resources. Re-running the same desired definition therefore converges instead of creating another Deployment every loop.

Create the Go project

You need Go installed for this deep dive.

Check:

go version

Create a workspace:

mkdir -p /tmp/preview-controller
cd /tmp/preview-controller

go mod init example.com/preview-controller

go get \
  k8s.io/apimachinery@v0.35.0 \
  k8s.io/client-go@v0.35.0

client-go uses matching v0.X.Y versions for Kubernetes v1.X.Y releases, so v0.35.0 aligns with our Kubernetes 1.35 target.

Save as main.go:

package main

import (
	"context"
	"encoding/json"
	"fmt"
	"log"
	"os"
	"time"

	metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
	"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
	"k8s.io/apimachinery/pkg/runtime/schema"
	"k8s.io/apimachinery/pkg/types"
	"k8s.io/client-go/dynamic"
	"k8s.io/client-go/rest"
	"k8s.io/client-go/tools/clientcmd"
)

var (
	previewGVR = schema.GroupVersionResource{Group: "platform.example.com", Version: "v1alpha1", Resource: "previewenvironments"}
	deployGVR  = schema.GroupVersionResource{Group: "apps", Version: "v1", Resource: "deployments"}
	serviceGVR = schema.GroupVersionResource{Group: "", Version: "v1", Resource: "services"}
)

type controller struct {
	namespace string
	client    dynamic.Interface
}

func main() {
	cfg, err := kubeConfig()
	if err != nil {
		log.Fatal(err)
	}

	client, err := dynamic.NewForConfig(cfg)
	if err != nil {
		log.Fatal(err)
	}

	namespace := os.Getenv("NAMESPACE")
	if namespace == "" {
		namespace = "cookbook"
	}

	c := &controller{namespace: namespace, client: client}
	ctx := context.Background()
	ticker := time.NewTicker(2 * time.Second)
	defer ticker.Stop()

	log.Printf("reconciling PreviewEnvironments in %q", namespace)

	for {
		if err := c.reconcileAll(ctx); err != nil {
			log.Printf("reconcile: %v", err)
		}
		<-ticker.C
	}
}

func kubeConfig() (*rest.Config, error) {
	if cfg, err := rest.InClusterConfig(); err == nil {
		return cfg, nil
	}

	return clientcmd.NewNonInteractiveDeferredLoadingClientConfig(
		clientcmd.NewDefaultClientConfigLoadingRules(),
		&clientcmd.ConfigOverrides{},
	).ClientConfig()
}

func (c *controller) reconcileAll(ctx context.Context) error {
	previews := c.client.Resource(previewGVR).Namespace(c.namespace)
	list, err := previews.List(ctx, metav1.ListOptions{})
	if err != nil {
		return err
	}

	for i := range list.Items {
		preview := &list.Items[i]
		if preview.GetDeletionTimestamp() != nil {
			continue
		}
		if err := c.reconcile(ctx, preview); err != nil {
			log.Printf("%s: %v", preview.GetName(), err)
		}
	}
	return nil
}

func (c *controller) reconcile(ctx context.Context, preview *unstructured.Unstructured) error {
	name := preview.GetName()
	image, _, _ := unstructured.NestedString(preview.Object, "spec", "image")
	replicas, _, _ := unstructured.NestedInt64(preview.Object, "spec", "replicas")
	if image == "" || replicas < 1 {
		return fmt.Errorf("spec.image and spec.replicas are required")
	}

	labels := map[string]any{
		"app.kubernetes.io/name":       name,
		"platform.example.com/preview": name,
	}
	owner := []any{map[string]any{
		"apiVersion": "platform.example.com/v1alpha1",
		"kind":       "PreviewEnvironment",
		"name":       name,
		"uid":        string(preview.GetUID()),
		"controller": true,
	}}

	deployment := &unstructured.Unstructured{Object: map[string]any{
		"apiVersion": "apps/v1",
		"kind":       "Deployment",
		"metadata": map[string]any{
			"name":            name,
			"namespace":       c.namespace,
			"ownerReferences": owner,
		},
		"spec": map[string]any{
			"replicas": replicas,
			"selector": map[string]any{"matchLabels": labels},
			"template": map[string]any{
				"metadata": map[string]any{"labels": labels},
				"spec": map[string]any{
					"containers": []any{map[string]any{
						"name":  "web",
						"image": image,
						"ports": []any{map[string]any{"containerPort": int64(80)}},
					}},
				},
			},
		},
	}}

	service := &unstructured.Unstructured{Object: map[string]any{
		"apiVersion": "v1",
		"kind":       "Service",
		"metadata": map[string]any{
			"name":            name,
			"namespace":       c.namespace,
			"ownerReferences": owner,
		},
		"spec": map[string]any{
			"selector": labels,
			"ports": []any{map[string]any{
				"name":       "http",
				"port":       int64(80),
				"targetPort": int64(80),
			}},
		},
	}}

	if err := c.apply(ctx, deployGVR, deployment); err != nil {
		return err
	}
	if err := c.apply(ctx, serviceGVR, service); err != nil {
		return err
	}

	current, err := c.client.Resource(deployGVR).Namespace(c.namespace).
		Get(ctx, name, metav1.GetOptions{})
	if err != nil {
		return err
	}
	ready, _, _ := unstructured.NestedInt64(current.Object, "status", "readyReplicas")

	return c.updateStatus(ctx, preview, ready)
}

func (c *controller) apply(ctx context.Context, gvr schema.GroupVersionResource, obj *unstructured.Unstructured) error {
	body, err := json.Marshal(obj.Object)
	if err != nil {
		return err
	}

	force := true
	_, err = c.client.Resource(gvr).Namespace(c.namespace).Patch(
		ctx,
		obj.GetName(),
		types.ApplyPatchType,
		body,
		metav1.PatchOptions{FieldManager: "preview-controller", Force: &force},
	)
	return err
}

func (c *controller) updateStatus(ctx context.Context, preview *unstructured.Unstructured, ready int64) error {
	url := fmt.Sprintf("http://%s.%s.svc.cluster.local", preview.GetName(), c.namespace)
	oldReady, _, _ := unstructured.NestedInt64(preview.Object, "status", "readyReplicas")
	oldURL, _, _ := unstructured.NestedString(preview.Object, "status", "url")
	if oldReady == ready && oldURL == url {
		return nil
	}

	updated := preview.DeepCopy()
	_ = unstructured.SetNestedField(updated.Object, ready, "status", "readyReplicas")
	_ = unstructured.SetNestedField(updated.Object, url, "status", "url")

	_, err := c.client.Resource(previewGVR).Namespace(c.namespace).
		UpdateStatus(ctx, updated, metav1.UpdateOptions{})
	return err
}

Format and resolve dependencies:

gofmt -w main.go
go mod tidy

Run the controller from your laptop first

The program first tries in-cluster credentials. If it is not running in Kubernetes, it falls back to your normal kubeconfig.

Make sure you are still pointed at the cookbook lab:

kubectl config current-context
kubectl config view --minify \
  -o jsonpath='{..namespace}{"\n"}'

Then run:

go run .

Leave it running.

In another terminal:

kubectl get preview,deployment,service,pods

The custom resource from Chapter 33 should now cause a Deployment and Service to appear.

Wait for the child Deployment to be created, then for its rollout:

kubectl wait   --for=create   deployment/pr-482   --timeout=30s

kubectl rollout status deployment/pr-482

Then inspect custom status:

kubectl get preview pr-482 \
  -o jsonpath='{.status.readyReplicas}{" ready -> "}{.status.url}{"\n"}'

You should eventually see:

2 ready -> http://pr-482.cookbook.svc.cluster.local

The path is now real:

PreviewEnvironment.spec
        |
        v
our Go controller
        |
        +-- server-side apply Deployment
        |
        +-- server-side apply Service
        |
        v
Deployment.status
        |
        v
PreviewEnvironment.status

Change desired state

Change the custom resource rather than the generated Deployment:

kubectl patch preview pr-482 \
  --type=merge \
  -p '{"spec":{"replicas":3}}'

Watch:

kubectl get preview,deployment,pods -w

The controller sees the new desired state and changes the Deployment.

Create drift on purpose

Now fight the controller.

Scale its child Deployment directly:

kubectl scale deployment pr-482 --replicas=1

Check immediately:

kubectl get deployment pr-482

Then check again a few seconds later:

sleep 3
kubectl get deployment pr-482

It should return to three replicas.

Why?

PreviewEnvironment.spec.replicas = 3
              |
              v
controller observes child replicas = 1
              |
              v
server-side apply desired Deployment
              |
              v
child replicas = 3

This is Chapter 1's control loop implemented by code we wrote ourselves.

Delete a child

Delete the generated Deployment:

kubectl delete deployment pr-482

Watch the labelled Deployment set rather than asking for the temporarily missing object by name:

kubectl get deployment   -l platform.example.com/preview=pr-482   -w

Within a reconciliation cycle, it should reappear.

Again, the controller is not issuing an imperative "restart" command.

It is repeatedly asserting:

A Deployment named pr-482 should exist with this spec.

Press Ctrl-C in the terminal running go run . before the next step.

Run the controller as a Kubernetes workload

Running from your laptop proved the control loop.

Now make the controller obey the same platform rules as every other workload.

Create a Dockerfile:

FROM golang:1.25-alpine AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY main.go ./
RUN CGO_ENABLED=0 GOOS=linux go build -o /controller .

FROM scratch
COPY --from=build /controller /controller
USER 65532:65532
ENTRYPOINT ["/controller"]

Build it:

docker build -t example/preview-controller:v1 .

Our lab is kind, so reuse the image-distribution lesson from Chapter 9:

kind load docker-image \
  example/preview-controller:v1 \
  --name ckad

Give it only the API permissions it needs

Save as controller-rbac.yaml:

apiVersion: v1
kind: ServiceAccount
metadata:
  name: preview-controller
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: preview-controller
rules:
  - apiGroups: ["platform.example.com"]
    resources:
      - previewenvironments
      - previewenvironments/status
    verbs:
      - get
      - list
      - watch
      - update
      - patch

  - apiGroups: ["apps"]
    resources:
      - deployments
    verbs:
      - get
      - list
      - watch
      - create
      - update
      - patch

  - apiGroups: [""]
    resources:
      - services
    verbs:
      - get
      - list
      - watch
      - create
      - update
      - patch
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: preview-controller
subjects:
  - kind: ServiceAccount
    name: preview-controller
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: preview-controller

Apply:

kubectl apply -f controller-rbac.yaml

Prove the scope before running anything:

kubectl auth can-i patch deployments \
  --as=system:serviceaccount:cookbook:preview-controller

kubectl auth can-i update previewenvironments/status \
  --api-group=platform.example.com \
  --as=system:serviceaccount:cookbook:preview-controller

kubectl auth can-i delete nodes \
  --as=system:serviceaccount:cookbook:preview-controller

The first two should be allowed.

Deleting Nodes should not be.

That connects our controller directly back to Chapter 23:

controller needs API access
        |
        v
ServiceAccount identity
        |
        v
Role grants minimum namespace permissions

Deploy the controller

Save as controller-deployment.yaml:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: preview-controller
spec:
  replicas: 1

  selector:
    matchLabels:
      app: preview-controller

  template:
    metadata:
      labels:
        app: preview-controller

    spec:
      serviceAccountName: preview-controller

      securityContext:
        runAsNonRoot: true
        seccompProfile:
          type: RuntimeDefault

      containers:
        - name: controller
          image: example/preview-controller:v1
          imagePullPolicy: IfNotPresent

          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities:
              drop:
                - ALL

          env:
            - name: NAMESPACE
              valueFrom:
                fieldRef:
                  fieldPath: metadata.namespace

Apply:

kubectl apply -f controller-deployment.yaml
kubectl rollout status deployment/preview-controller

Follow its logs:

kubectl logs deployment/preview-controller -f

The controller now uses:

  • a ServiceAccount from the RBAC chapter
  • namespace discovery from the Downward API chapter
  • a hardened SecurityContext from Chapter 24
  • a locally built image loaded into kind as in Chapter 9
  • a custom API from Chapter 33
  • server-side apply to express child desired state

This is why the earlier chapters matter.

They compose.

Leave the controller, CRD and PreviewEnvironment/pr-482 running.

Chapter 35 will inspect the ownership and status relationships we just created.

What we deliberately left out

Our controller polls every two seconds because that keeps the first implementation understandable.

A production controller normally evolves toward:

watch / informer
      |
      v
work queue
      |
      v
reconcile(key)
      |
      +-- retry with backoff
      +-- status / conditions
      +-- metrics
      +-- leader election when replicated

Libraries such as client-go and controller-runtime provide those building blocks.

The essential idea does not change:

Reconciliation should be idempotent. Running it repeatedly should converge the system toward the same desired state, not create more side effects every time.


35. Ownership, Finalizers, Status and Operators [DEV] [DEEP DIVE]

Once you understand controllers, several Kubernetes mechanisms become easier to place.

They are not arbitrary advanced features.

They help controllers express responsibility, cleanup and observed state.

We now have our own controller running, so we can inspect these mechanisms on something we built rather than only discussing them in the abstract.

Ownership: which object is responsible for this child?

First inspect the Deployment created for our preview:

kubectl get deployment pr-482 \
  -o jsonpath='{.metadata.ownerReferences}{"\n"}'

Then the Service:

kubectl get service pr-482 \
  -o jsonpath='{.metadata.ownerReferences}{"\n"}'

Both should point at:

PreviewEnvironment/pr-482

The relationship is:

PreviewEnvironment/pr-482
       |
       +-- Deployment/pr-482
       |        |
       |        v
       |     ReplicaSet
       |        |
       |        v
       |       Pods
       |
       +-- Service/pr-482

Compare that with the built-in chain we saw earlier:

Deployment
    |
    v
ReplicaSet
    |
    v
Pod

Our controller is using the same Kubernetes ownership mechanism as built-in controllers.

Labels and ownership are not the same thing:

label
  -> which objects match this query/selector?

ownerReference
  -> which object controls this dependent resource?

Controller responsibility: delete a child

Delete the generated Service:

kubectl delete service pr-482

Watch the labelled Service set:

kubectl get service   -l platform.example.com/preview=pr-482   -w

Within a reconciliation cycle, the Service should reappear.

That happened because our controller still sees:

PreviewEnvironment/pr-482 exists
        |
        v
Service/pr-482 should exist

Ownership did not recreate the Service.

Reconciliation did.

That distinction matters.

Owner references describe relationships and enable garbage collection; controllers are the actors that continuously restore desired state.

Status: report observed state through the API

Chapter 33 gave the CRD a status subresource.

Chapter 34 made the controller populate it.

Inspect it:

kubectl get preview pr-482 \
  -o jsonpath='{.spec.replicas}{" desired, "}{.status.readyReplicas}{" ready, "}{.status.url}{"\n"}'

You should see something like:

3 desired, 3 ready, http://pr-482.cookbook.svc.cluster.local

We have now implemented the same pattern we met in Chapter 1:

spec
  -> what the user wants

status
  -> what the controller currently observes

A useful custom API should let normal operational questions be answered from the API.

Users should not have to start with controller logs simply to discover whether the requested system is ready.

Conditions

Our tiny controller deliberately reports only two status fields.

Larger APIs commonly expose structured conditions:

status:
  conditions:
    - type: Ready
      status: "True"
      reason: DeploymentAvailable

Common condition concepts include:

Ready
Available
Progressing
Degraded

Conditions are especially useful when readyReplicas: 1 is not enough to explain why the system is not ready.

Garbage collection: delete the owner

Now delete the PreviewEnvironment itself:

kubectl delete preview pr-482

Watch its children:

kubectl get deployment,service \
  -l platform.example.com/preview=pr-482 \
  -w

Because the Deployment and Service contain owner references to the custom resource, Kubernetes garbage collection can remove them when the owner disappears.

The path is:

PreviewEnvironment deleted
        |
        v
ownerReferences become invalid
        |
        v
garbage collector removes dependants

Our controller does not need explicit code saying:

on PreviewEnvironment delete:
    delete Deployment
    delete Service

for these Kubernetes-native child resources.

Recreate the preview so the rest of the chapter can continue:

kubectl apply -f preview.yaml
kubectl wait   --for=create   deployment/pr-482   --timeout=30s
kubectl rollout status deployment/pr-482

Finalizers: cleanup that garbage collection cannot perform for you

Owner references work well for Kubernetes objects.

But suppose our controller also created something outside Kubernetes:

DNS record
cloud database
SaaS tenant
external load balancer

Kubernetes garbage collection cannot delete an external database merely because an API object disappeared.

That is where finalizers fit.

A finalizer delays final deletion until responsible cleanup logic has completed.

Create a harmless demo object:

kubectl create configmap finalizer-demo \
  --from-literal=test=true

Add a finalizer:

kubectl patch configmap finalizer-demo \
  --type=merge \
  -p '{"metadata":{"finalizers":["example.com/demo"]}}'

Request deletion without waiting:

kubectl delete configmap finalizer-demo --wait=false

Inspect:

kubectl get configmap finalizer-demo -o yaml

Notice:

deletionTimestamp: ...

The object is pending deletion because its finalizer remains.

A real controller would now:

observe deletionTimestamp
        |
        v
perform external cleanup
        |
        v
remove its finalizer
        |
        v
Kubernetes completes deletion

For the lab, remove it manually:

kubectl patch configmap finalizer-demo \
  --type=json \
  -p='[
    {
      "op":"remove",
      "path":"/metadata/finalizers"
    }
  ]'

Check:

kubectl get configmap finalizer-demo

It should be gone.

This explains many resources that appear to be:

stuck Terminating

The object may be waiting for a controller to finish a cleanup contract.

Operator: controller plus domain knowledge

An Operator is not a special executable type understood by Kubernetes.

Think:

Custom API
    +
Controller
    +
Domain knowledge

Our PreviewEnvironment controller knows only a tiny domain rule:

PreviewEnvironment
    -> nginx-compatible Deployment
    -> Service on port 80

A database Operator might understand much more:

kind: DatabaseCluster
spec:
  version: "18"
  replicas: 3

and reconcile operations such as:

create members
bootstrap replication
replace failed members
take backups
restore backups
rotate credentials
perform upgrades

It is still the same reconciliation model.

The domain knowledge is richer.

Controller failure modes

Writing a controller makes several failure modes easier to understand.

Non-idempotent reconciliation

Bad:

Every reconcile creates another cloud database.

Better:

Ensure the required database exists.
Create it only when absent.
Update it only when desired state differs.

Our lab controller used server-side apply for exactly this reason:

same desired child object
        |
        v
apply repeatedly
        |
        v
convergent state

rather than:

reconcile #1 -> Deployment A
reconcile #2 -> Deployment B
reconcile #3 -> Deployment C

Update loops

A controller can trigger itself:

controller updates object
        |
        v
watch event
        |
        v
reconcile
        |
        v
controller writes identical state again
        |
        +------ loop

Only write when state actually differs.

Our status code checks whether readyReplicas or url changed before issuing another status update.

Fighting controllers

Two controllers should not continually attempt to own the same field or external resource in incompatible ways.

Otherwise:

controller A writes X
      |
controller B writes Y
      |
controller A writes X
      |
      +---- forever

Server-side apply field ownership can help make those conflicts explicit, but it does not make contradictory intent disappear.

Status that lies

Do not report Ready=True merely because a child object was created.

Report what the API contract says readiness means.

For our preview API, readyReplicas comes from the child Deployment's observed status rather than simply copying spec.replicas.

That difference is the entire point of status.

A controller is useful when it turns desired state into truthful, convergent behaviour.


36. Clean Up the Custom API [DEV]

We deliberately kept the controller running so Chapters 34 and 35 could build on the same system.

Clean it up in dependency order.

Delete the Custom Resource

kubectl delete preview pr-482 --ignore-not-found

Its owned Deployment and Service should disappear through garbage collection.

Check:

kubectl get deployment,service \
  -l platform.example.com/preview=pr-482

Delete the controller workload and RBAC

kubectl delete -f controller-deployment.yaml --ignore-not-found
kubectl delete -f controller-rbac.yaml --ignore-not-found

Delete the CRD

kubectl delete crd \
  previewenvironments.platform.example.com

Deleting a CRD deletes the custom resources stored through that API.

Only do this casually in a disposable lab environment.

Remove the local lab files

If you created the Go controller under /tmp:

rm -rf /tmp/preview-controller

Remove manifests from the current directory if present:

rm -f \
  preview.yaml \
  preview-crd.yaml \
  controller-rbac.yaml \
  controller-deployment.yaml

We have now traversed the full extension lifecycle:

CRD
  -> API type exists

Custom Resource
  -> desired state exists

Controller
  -> behaviour exists

ownerReferences
  -> child relationships exist

status
  -> observed state is reported

finalizers
  -> external cleanup can become part of deletion

That is the same Kubernetes control model, extended with our own API and code.


Part XII - CKAD Speed Layer

37. CKAD Command Patterns [CKAD]

Understanding matters first.

Once the mental model is solid, speed comes from having a small number of productive command patterns in muscle memory.

Context and namespace

kubectl config get-contexts
kubectl config current-context
kubectl config set-context --current --namespace=cookbook

Discovery

kubectl api-resources
kubectl api-versions
kubectl explain pod
kubectl explain deployment.spec.template.spec.containers

Pods

Create:

kubectl run test \
  --image=busybox:1.36 \
  --restart=Never \
  --command -- sleep 3600

Generate YAML:

kubectl run test \
  --image=busybox:1.36 \
  --restart=Never \
  --dry-run=client \
  -o yaml

Logs and exec:

kubectl logs test
kubectl logs test --previous
kubectl exec -it test -- sh

Deployments

Create:

kubectl create deployment api \
  --image=nginx:1.27-alpine \
  --replicas=3

Generate YAML:

kubectl create deployment api \
  --image=nginx:1.27-alpine \
  --replicas=3 \
  --dry-run=client \
  -o yaml

Scale:

kubectl scale deployment api --replicas=5

Image:

kubectl set image deployment/api \
  nginx=nginx:1.28-alpine

Rollouts:

kubectl rollout status deployment/api
kubectl rollout history deployment/api
kubectl rollout undo deployment/api

Services

Expose a Deployment:

kubectl expose deployment api \
  --port=80 \
  --target-port=80

Generate Service YAML without creating:

kubectl expose deployment api \
  --port=80 \
  --target-port=80 \
  --dry-run=client \
  -o yaml

Endpoints:

kubectl get endpointslices \
  -l kubernetes.io/service-name=api

ConfigMaps and Secrets

kubectl create configmap app-config \
  --from-literal=MODE=dev

kubectl create secret generic app-secret \
  --from-literal=password=example

Generate instead of create:

kubectl create configmap app-config \
  --from-literal=MODE=dev \
  --dry-run=client \
  -o yaml

kubectl create secret generic app-secret \
  --from-literal=password=example \
  --dry-run=client \
  -o yaml

Jobs and CronJobs

kubectl create job hello \
  --image=busybox:1.36 \
  -- echo hello

kubectl create cronjob clock \
  --image=busybox:1.36 \
  --schedule='*/5 * * * *' \
  -- date

Generate YAML by adding:

--dry-run=client -o yaml

Editing and patching

kubectl edit deployment api

Merge patch:

kubectl patch deployment api \
  --type=merge \
  -p '{"spec":{"replicas":2}}'

Strategic merge is useful for Kubernetes built-in list structures such as named containers:

kubectl patch deployment api \
  --type=strategic \
  -p 'spec:
        template:
          spec:
            containers:
              - name: nginx
                image: nginx:1.28-alpine'

Output

YAML:

kubectl get pod test -o yaml

JSONPath:

kubectl get pod test \
  -o jsonpath='{.status.podIP}{"\n"}'

Custom columns:

kubectl get pods \
  -o custom-columns='NAME:.metadata.name,NODE:.spec.nodeName,IP:.status.podIP'

Sort:

kubectl get pods \
  --sort-by=.metadata.creationTimestamp

Labels:

kubectl get pods --show-labels
kubectl get pods -l app=api

Waiting

kubectl wait \
  --for=condition=Ready \
  pod/test \
  --timeout=60s

Deployment rollout:

kubectl rollout status deployment/api --timeout=60s

Debugging

kubectl get pods
kubectl describe pod <pod>
kubectl logs <pod>
kubectl logs <pod> --previous
kubectl get events --sort-by=.metadata.creationTimestamp

Network path:

kubectl get svc
kubectl get pods --show-labels
kubectl get endpointslices
kubectl exec <client> -- nslookup <service>

Authorization:

kubectl auth can-i get pods

Temporary testing Pods

DNS/network shell:

kubectl run tmp \
  --image=busybox:1.36 \
  --restart=Never \
  -it --rm -- sh

If the exam environment has a different preferred utility image, use whatever image is already available and appropriate to the task.

Useful shell habit

For commands you will repeat, capture names rather than retyping hashes:

POD=$(kubectl get pod -l app=api \
  -o jsonpath='{.items[0].metadata.name}')

echo "$POD"

The goal is not command golf.

The goal is to spend exam time solving the Kubernetes problem rather than manually reconstructing YAML you could have generated.


38. Final Lab Cleanup

The easiest cleanup is deleting the namespace because nearly everything in this cookbook is namespaced:

kubectl delete namespace cookbook

If you changed your current context to default to cookbook, clear or change that namespace afterwards.

For example:

kubectl config set-context \
  --current \
  --namespace=default

Check:

kubectl config view --minify \
  -o jsonpath='{..namespace}{"\n"}'

Appendix A - The Mental Models Worth Remembering

If you remember nothing else, remember these.

Kubernetes is a control loop

spec
 |
 v
controller
 |
 v
actual state
 |
 v
status
 |
 +---- observe and reconcile ----+

Controller-managed Pods are replaceable

Deployment
    |
    v
ReplicaSet
    |
    v
Pods

Delete a Pod and desired state remains.

The controller creates another.

A Service is stable because Pods are not

DNS name
  |
  v
Service
  |
  v
EndpointSlices
  |
  v
ready Pods

Labels are relationships, not decoration

selector
  |
  v
matching labels

Services, controllers and policies all build behaviour on this idea.

Readiness and liveness answer different questions

readiness -> should traffic reach me?
liveness  -> should kubelet restart me?
startup   -> have I successfully started yet?

Configuration should not require rebuilding the image

ConfigMap -> non-secret configuration
Secret    -> sensitive configuration API
Downward API -> runtime Kubernetes metadata

Data lifetime is a separate decision from Pod lifetime

emptyDir -> Pod-scoped
PVC      -> persistent storage claim

Debug relationships rather than symptoms

Workload:

Deployment -> ReplicaSet -> Pod -> Container -> Process

Network:

DNS -> Service -> selector -> EndpointSlice -> ready Pod -> port -> process

Configuration:

Pod spec -> referenced object -> key -> mount/env -> application

Security boundaries answer different questions

authentication
    -> who are you?

authorization / RBAC
    -> may you make this API request?

admission
    -> is this write acceptable?

scheduling controls
    -> where may this workload run?

NetworkPolicy
    -> which network flows are allowed?

infrastructure IAM / SSH
    -> who may administer the actual machines?

Namespaces help provide scope, but are not an automatic security boundary by themselves.

For stronger tenancy, separate tenant Kubernetes APIs/control planes can add another boundary; see multitenancy-appendix.md.

Kubernetes extension uses the same model

CRD
  -> teach the API a type

Custom Resource
  -> declare desired state for that type

Controller
  -> reconcile it

Operator
  -> controller + domain-specific operational knowledge

Appendix B - Current CKAD Domain Map

This cookbook is intentionally organised for understanding rather than mirroring the exam outline chapter-for-chapter.

The current CKAD domains map roughly as follows:

Application Design and Build - 20%

Covered by:

  • container image definitions and builds
  • Pods and container command/args
  • workload selection
  • multi-container Pods
  • init containers and sidecars
  • Jobs and CronJobs
  • ephemeral and persistent volumes
  • StatefulSets and DaemonSets

Application Deployment - 20%

Covered by:

  • Deployments and ReplicaSets
  • scaling
  • rolling updates and rollback
  • blue-green and canary strategies
  • Helm
  • Kustomize

Application Observability and Maintenance - 15%

Covered by:

  • get, describe, logs and events
  • readiness, liveness and startup probes
  • rollout state
  • API deprecation and discovery
  • systematic debugging
  • ephemeral debug containers

Application Environment, Configuration and Security - 25%

Covered by:

  • requests and limits
  • quotas and limits policy concepts
  • ConfigMaps and Secrets
  • Downward API
  • ServiceAccounts
  • RBAC and authorization
  • admission concepts
  • SecurityContext and capabilities
  • CRDs and Operators

Services and Networking - 20%

Covered by:

  • Services
  • labels and selectors
  • EndpointSlices
  • DNS
  • network troubleshooting
  • NetworkPolicy
  • Ingress

Supplemental deep dive:

  • Gateway API concepts and hands-on Cilium Gateway API in cillium-gateay-appendix.md

Appendix C - Official References

Prefer current official documentation when a field, API version or exam detail is uncertain.

The habit this cookbook is trying to teach is simple:

Do not memorise Kubernetes as a pile of resource definitions. Learn the control loops and relationships, then let the API tell you the exact syntax.

Supplemental - Run Kubernetes Locally with kind [DEV]

The rest of our cookbook assumes you already have access to a Kubernetes cluster.

If you do not, kind is probably the quickest way to get one.

kind means Kubernetes IN Docker.

It runs Kubernetes nodes as containers on your machine:

Your machine
    |
    v
Docker
    |
    +-- kind-control-plane
            |
            +-- kube-apiserver
            +-- scheduler
            +-- controller-manager
            +-- kubelet
            +-- containerd

This is not how a production Kubernetes cluster is normally deployed.

It is, however, an excellent disposable Kubernetes lab.


Prerequisites

You need:

Docker
kubectl
kind

Check Docker:

docker version

Check kubectl:

kubectl version --client

If those work, install kind.


Install kind

macOS

With Homebrew:

brew install kind

Linux

For x86-64:

curl -Lo ./kind \
  https://kind.sigs.k8s.io/dl/v0.33.0/kind-linux-amd64

chmod +x ./kind

sudo mv ./kind /usr/local/bin/kind

For ARM64:

curl -Lo ./kind \
  https://kind.sigs.k8s.io/dl/v0.33.0/kind-linux-arm64

chmod +x ./kind

sudo mv ./kind /usr/local/bin/kind

Windows

With winget:

winget install Kubernetes.kind

Check it:

kind version

Create the Cluster

The cookbook targets Kubernetes 1.35.

Create a cluster called ckad using Kubernetes 1.35:

kind create cluster \
  --name ckad \
  --image kindest/node:v1.35.8@sha256:07b2536e30b803ed61d1677a79df6115f798ce64c80f9e22f6ed45afd09323c0 \
  --wait 2m

kind will:

create a container
        |
        v
turn it into a Kubernetes node
        |
        v
bootstrap Kubernetes
        |
        v
write kubeconfig
        |
        v
switch kubectl to the new cluster

Check:

kind get clusters

Expected:

ckad

Check your Kubernetes context:

kubectl config current-context

Expected:

kind-ckad

Observe the Cluster

Ask Kubernetes what nodes exist:

kubectl get nodes

You should see something similar to:

NAME                 STATUS   ROLES           AGE   VERSION
ckad-control-plane   Ready    control-plane   1m    v1.35.x

More detail:

kubectl get nodes -o wide

Check the control plane:

kubectl cluster-info

And see the system workloads Kubernetes created:

kubectl get pods -n kube-system

At this point you have a real Kubernetes API to experiment with.


Look Underneath

Remember that the Kubernetes node is actually a container.

Ask Docker:

docker ps

You should see:

ckad-control-plane

So the relationship is:

kubectl
   |
   v
Kubernetes API
   |
   v
ckad-control-plane
   |
   v
Docker container

From this point onwards, mostly forget about Docker.

Interact with the cluster through the Kubernetes API.


Start the Cookbook

Your cluster now exists.

Continue with Section 0 - Lab Setup.

Create the cookbook namespace:

kubectl create namespace cookbook

Make it the default namespace:

kubectl config set-context \
  --current \
  --namespace=cookbook

Check:

kubectl config view --minify \
  -o jsonpath='{..namespace}{"\n"}'

Expected:

cookbook

You are now ready for the rest of the cookbook.


One kind-specific trick worth knowing

Later you may build your own container image locally.

For example:

docker build -t my-app:v1 .

That image exists on your machine, but it does not automatically exist inside the kind node.

Load it:

kind load docker-image my-app:v1 --name ckad

Then Kubernetes can use it:

kubectl run my-app \
  --image=my-app:v1 \
  --image-pull-policy=IfNotPresent

This:

docker build
     |
     v
Host image
     |
     | kind load docker-image
     v
kind node
     |
     v
Pod

is useful when testing applications without pushing every image to a registry.

Avoid using the latest tag for this workflow. Kubernetes normally treats :latest as imagePullPolicy: Always, which may cause it to try pulling the image from a registry instead of using the image you loaded locally.


Reset the Lab

One of the best features of kind is that the entire cluster is disposable.

Destroy it:

kind delete cluster --name ckad

Check:

kind get clusters

Then recreate it whenever you want:

kind create cluster \
  --name ckad \
  --image kindest/node:v1.35.8@sha256:07b2536e30b803ed61d1677a79df6115f798ce64c80f9e22f6ed45afd09323c0 \
  --wait 2m

A completely broken lab is therefore not a disaster.

Sometimes deleting it and starting again is exactly the point.

Appendix C - Build Kubernetes on Linux with kubeadm

A runnable companion to the Kubernetes application developer cookbook. It is intentionally CKA / platform-engineering territory.

Target stack:

Kubernetes 1.35
containerd
kubeadm
kubelet
Cilium 1.20.1
kube-proxy replacement

The lab assumes Ubuntu/Debian-style Linux machines.

For the best experience use two disposable Linux VMs:

k8s-control
  2+ CPU
  2+ GiB RAM

k8s-worker
  2+ GiB RAM

Both machines must be able to reach each other directly.

The point is not merely to make Kubernetes work.

The point is to watch it not work yet, understand why, and then add the missing pieces.


C.1 What kind Was Hiding

In the local quickstart we used:

kind create cluster

and Kubernetes appeared.

That was real Kubernetes.

kind itself uses kubeadm to bootstrap its nodes.

This appendix performs the interesting pieces ourselves:

Linux
  |
  v
containerd
  |
  v
kubelet
  |
  v
kubeadm
  |
  v
control plane
  |
  v
Cilium
  |
  v
working cluster

By the end you should understand what sits underneath:

Deployment
Service
Pod

and why a Kubernetes cluster can exist while still being:

NotReady

C.2 The Cluster We Are Building

Our final topology:

                         k8s-control
                              |
                 +------------+-------------+
                 |            |             |
                 v            v             v
              API server   scheduler    controller
                 |
                 v
                etcd
                 |
                 |
        Kubernetes API :6443
                 |
          +------+------+
          |             |
          v             v
     k8s-control    k8s-worker
          |             |
       kubelet        kubelet
          |             |
      containerd     containerd
          |             |
          +------+------+
                 |
               Cilium
                 |
                 v
             Pod network

We will deliberately not install kube-proxy.

Cilium will later provide:

CNI
+
Pod networking
+
NetworkPolicy enforcement
+
Kubernetes Service load balancing
+
kube-proxy replacement

C.3 Know the Pieces

Before installing anything, keep these roles separate.

kubectl

Client for the Kubernetes API.

kubectl
   |
   v
kube-apiserver

kubelet

Node agent.

It watches for Pods assigned to its node and asks the container runtime to run them.

API server
    |
    v
 kubelet
    |
    v
containerd

containerd

Container runtime.

The kubelet communicates with it through CRI:

kubelet
   |
   | CRI
   v
containerd
   |
   v
container

kubeadm

Bootstrap and lifecycle tool.

kubeadm
   |
   v
creates/configures Kubernetes

It is not the daemon continuously running the cluster.

Cilium

Networking implementation.

Kubernetes networking APIs
          |
          v
        Cilium
          |
          v
Linux / eBPF dataplane

C.4 Prepare Both Machines

Run this section on:

k8s-control
k8s-worker

Check hostname:

hostname

Set unique names if necessary.

Control plane:

sudo hostnamectl set-hostname k8s-control

Worker:

sudo hostnamectl set-hostname k8s-worker

Log out and back in if your shell prompt does not immediately reflect the change.

Check addresses:

ip -4 addr

Make sure each machine can reach the other:

ping -c 2 <other-node-ip>

C.5 Disable Swap

Check:

swapon --show

For this lab disable swap:

sudo swapoff -a

Check again:

swapon --show

It should return nothing.

For a persistent lab also disable the swap entry in:

/etc/fstab

Modern Kubernetes can be configured to use swap, but the default kubelet behaviour used by this lab expects it to be disabled.


C.6 Enable IPv4 Forwarding

Create:

cat <<'EOF' | sudo tee /etc/sysctl.d/k8s.conf
net.ipv4.ip_forward = 1
EOF

Apply:

sudo sysctl --system

Check:

sysctl net.ipv4.ip_forward

Expected:

net.ipv4.ip_forward = 1

This is our first reminder that Kubernetes networking eventually becomes ordinary Linux networking.


C.7 Install containerd

Install on both nodes:

sudo apt-get update
sudo apt-get install -y containerd

Check:

containerd --version

Generate a default config:

sudo mkdir -p /etc/containerd

containerd config default \
  | sudo tee /etc/containerd/config.toml >/dev/null

Kubernetes and the runtime should use the same cgroup model.

Configure containerd to use the systemd cgroup driver:

sudo sed -i \
  's/SystemdCgroup = false/SystemdCgroup = true/' \
  /etc/containerd/config.toml

Verify:

grep -n SystemdCgroup /etc/containerd/config.toml

Expected:

SystemdCgroup = true

Check that CRI has not been disabled:

grep disabled_plugins /etc/containerd/config.toml || true

If cri appears in disabled_plugins, remove it.

Restart:

sudo systemctl restart containerd
sudo systemctl enable containerd

Inspect:

systemctl status containerd --no-pager

We now have:

Linux
  |
  v
containerd

but no Kubernetes.


C.8 Install kubeadm, kubelet and kubectl

Run on both nodes.

Install repository prerequisites:

sudo apt-get update

sudo apt-get install -y \
  apt-transport-https \
  ca-certificates \
  curl \
  gpg

Create the keyring directory:

sudo mkdir -p -m 755 /etc/apt/keyrings

Add the Kubernetes 1.35 repository key:

curl -fsSL \
  https://pkgs.k8s.io/core:/stable:/v1.35/deb/Release.key \
  | sudo gpg --dearmor \
  -o /etc/apt/keyrings/kubernetes-apt-keyring.gpg

Add the repository:

echo \
  'deb [signed-by=/etc/apt/keyrings/kubernetes-apt-keyring.gpg] https://pkgs.k8s.io/core:/stable:/v1.35/deb/ /' \
  | sudo tee /etc/apt/sources.list.d/kubernetes.list

Install:

sudo apt-get update

sudo apt-get install -y \
  kubelet \
  kubeadm \
  kubectl

Hold the packages so an ordinary system upgrade does not unexpectedly move the cluster to a different Kubernetes version:

sudo apt-mark hold \
  kubelet \
  kubeadm \
  kubectl

Check:

kubeadm version
kubectl version --client
kubelet --version

Enable kubelet:

sudo systemctl enable --now kubelet

Inspect:

systemctl status kubelet --no-pager

It may be unhealthy.

That is fine.

We have installed the node agent but have not told it what cluster it belongs to.

Look at its logs:

journalctl -u kubelet -n 30 --no-pager

Do not fix random errors yet.

We are missing the cluster.


C.9 Initialise the Control Plane

Run this section only on:

k8s-control

Find the IP address the worker can use to reach the control plane:

ip -4 addr

For example:

192.168.1.20

Set it:

export CONTROL_PLANE_IP=192.168.1.20

Confirm:

echo "$CONTROL_PLANE_IP"

Check the Kubernetes version installed:

kubeadm version -o short

Save it:

export K8S_VERSION="$(kubeadm version -o short)"

Now initialise Kubernetes:

sudo kubeadm init \
  --kubernetes-version="$K8S_VERSION" \
  --apiserver-advertise-address="$CONTROL_PLANE_IP" \
  --skip-phases=addon/kube-proxy

Stop and notice:

--skip-phases=addon/kube-proxy

A conventional kubeadm cluster normally installs kube-proxy.

We are deliberately leaving it out.

Eventually:

Service
   |
   v
Cilium
   |
   v
eBPF
   |
   v
Pod

will replace the traditional:

Service
   |
   v
kube-proxy
   |
   v
iptables / IPVS
   |
   v
Pod

Keep the kubeadm join ... command printed at the end.

We will use it shortly.


C.10 Configure kubectl

kubeadm init created administrator credentials:

/etc/kubernetes/admin.conf

Copy them:

mkdir -p "$HOME/.kube"

sudo cp \
  /etc/kubernetes/admin.conf \
  "$HOME/.kube/config"

sudo chown \
  "$(id -u):$(id -g)" \
  "$HOME/.kube/config"

Check:

kubectl cluster-info

Then:

kubectl get nodes

You should have something similar to:

NAME          STATUS     ROLES           AGE   VERSION
k8s-control   NotReady   control-plane   ...   v1.35.x

Excellent.

Kubernetes exists.

It is not yet a working cluster.


C.11 What kubeadm Created

Inspect system Pods:

kubectl get pods \
  -n kube-system \
  -o wide

Look for:

etcd
kube-apiserver
kube-controller-manager
kube-scheduler
coredns

Now inspect the host:

sudo ls -l /etc/kubernetes/manifests

You should find:

etcd.yaml
kube-apiserver.yaml
kube-controller-manager.yaml
kube-scheduler.yaml

These are static Pods.

/etc/kubernetes/manifests
          |
          v
       kubelet
          |
          v
    static Pods
          |
    +-----+-----+----------+-----------+
    |           |          |           |
    v           v          v           v
 API server   etcd     scheduler   controller

No Deployment created them.

No ReplicaSet owns them.

The kubelet watches that directory.

Inspect the API server:

kubectl -n kube-system get pods \
  -l component=kube-apiserver

Then:

kubectl -n kube-system describe pod \
  "$(kubectl -n kube-system get pod \
      -l component=kube-apiserver \
      -o jsonpath='{.items[0].metadata.name}')"

The API contains a mirror Pod representing the static Pod managed by the node's kubelet.


C.12 Why Is the Node NotReady?

Ask Kubernetes:

kubectl describe node k8s-control

Pay attention to:

Conditions
Events

Check CoreDNS:

kubectl get pods \
  -n kube-system \
  -l k8s-app=kube-dns

It may still be:

Pending

or otherwise unavailable.

Why?

We currently have:

API server       ✓
etcd             ✓
scheduler        ✓
controller       ✓
kubelet          ✓
containerd       ✓

Pod network      ✗

kubeadm bootstraps Kubernetes.

It does not choose the cluster networking implementation for you.


C.13 Prove kube-proxy Is Missing

Check:

kubectl get daemonsets \
  -n kube-system

Then explicitly:

kubectl get daemonset kube-proxy \
  -n kube-system

Expected:

Error from server (NotFound)

That was intentional.

We are going to ask Cilium to provide the Kubernetes Service dataplane instead.


C.14 Join the Worker

On:

k8s-worker

run the join command printed by kubeadm init.

It will look similar to:

sudo kubeadm join 192.168.1.20:6443 \
  --token <token> \
  --discovery-token-ca-cert-hash sha256:<hash>

If you lost it, regenerate one on the control plane:

kubeadm token create \
  --print-join-command

Then run the returned command with sudo on the worker.

Back on the control plane:

kubectl get nodes

You should now see:

NAME          STATUS     ROLES           VERSION
k8s-control   NotReady   control-plane   v1.35.x
k8s-worker    NotReady   <none>          v1.35.x

This is a useful state.

We have two Kubernetes nodes.

Neither has a working Pod network.


C.15 Install the Cilium CLI

Run this on the control plane.

CILIUM_CLI_VERSION="$(
  curl -s \
  https://raw.githubusercontent.com/cilium/cilium-cli/main/stable.txt
)"

CLI_ARCH=amd64

if [ "$(uname -m)" = "aarch64" ]; then
  CLI_ARCH=arm64
fi

Download:

curl -L --fail --remote-name-all \
  "https://github.com/cilium/cilium-cli/releases/download/${CILIUM_CLI_VERSION}/cilium-linux-${CLI_ARCH}.tar.gz"{,.sha256sum}

Verify:

sha256sum --check \
  "cilium-linux-${CLI_ARCH}.tar.gz.sha256sum"

Install:

sudo tar xzvfC \
  "cilium-linux-${CLI_ARCH}.tar.gz" \
  /usr/local/bin

Cleanup:

rm "cilium-linux-${CLI_ARCH}.tar.gz"{,.sha256sum}

Check:

cilium version

C.16 Install Cilium as the CNI and kube-proxy Replacement

Make sure the control-plane address is still available:

export CONTROL_PLANE_IP="$(
  kubectl get node k8s-control \
    -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'
)"

Check:

echo "$CONTROL_PLANE_IP"

Install Cilium 1.20.1:

cilium install 1.20.1 \
  --set kubeProxyReplacement=true \
  --set k8sServiceHost="$CONTROL_PLANE_IP" \
  --set k8sServicePort=6443

Why do we explicitly tell Cilium where the API server lives?

Normally Pods can reach the API server through the Kubernetes Service.

But:

Kubernetes Service
       |
       v
Service dataplane

is exactly what we have not implemented yet.

There is no kube-proxy.

So Cilium needs the real API endpoint while it bootstraps the dataplane that will eventually implement Kubernetes Services.

Watch:

kubectl get pods \
  -n kube-system \
  -w

Eventually you should see Cilium agents and the operator running.

Stop with:

Ctrl-C

C.17 Watch the Cluster Become Ready

Check:

kubectl get nodes

We want:

NAME          STATUS   ROLES           VERSION
k8s-control   Ready    control-plane   v1.35.x
k8s-worker    Ready    <none>          v1.35.x

What changed?

Kubernetes exists
       |
       v
Nodes NotReady
       |
       | install Cilium
       v
CNI available
       |
       v
Pod networking available
       |
       v
Nodes Ready

Check Cilium:

cilium status --wait

Run the connectivity test:

cilium connectivity test

The test creates workloads and verifies the dataplane rather than merely checking whether Pods exist.


C.18 Prove Cilium Replaced kube-proxy

Check again:

kubectl get daemonset kube-proxy \
  -n kube-system

It should still not exist.

Now ask Cilium:

kubectl -n kube-system exec ds/cilium -- \
  cilium-dbg status \
  | grep KubeProxyReplacement

Expected:

KubeProxyReplacement:   True

For more detail:

kubectl -n kube-system exec ds/cilium -- \
  cilium-dbg status --verbose

Our cluster now looks more like:

Kubernetes API
      |
      v
   Service
      |
      v
Cilium agent
      |
      v
 eBPF maps
      |
      v
   backend Pod

C.19 Build a Workload

Return to familiar Kubernetes APIs.

Create:

kubectl create deployment web \
  --image=nginx:1.27-alpine \
  --replicas=2

Expose:

kubectl expose deployment web \
  --port=80

Observe:

kubectl get pods \
  -o wide

Then:

kubectl get service web

Create a client:

kubectl run client \
  --image=curlimages/curl \
  --restart=Never \
  --command -- \
  sleep 3600

Wait:

kubectl wait \
  --for=condition=Ready \
  pod/client \
  --timeout=90s

Call the Service:

kubectl exec client -- \
  curl -s http://web

You should receive the NGINX page.

There is still no kube-proxy.

Inspect Cilium's Service table:

kubectl -n kube-system exec ds/cilium -- \
  cilium-dbg service list

Look for the ClusterIP belonging to:

web

The object path is:

Service
   |
   v
Kubernetes API
   |
   v
Cilium observes it
   |
   v
eBPF service state
   |
   v
backend Pods

The reconciliation model from the CKAD cookbook now reaches all the way into the network dataplane.


C.20 Observe the kubelet

On the worker:

systemctl status kubelet --no-pager

Then:

journalctl -u kubelet \
  --since "10 minutes ago" \
  --no-pager

Remember:

API server
    ^
    |
 kubelet
    |
    v
containerd
    |
    v
   Pods

The kubelet:

registers the Node
watches assigned Pods
starts containers
mounts volumes
runs probes
reports status

It is the bridge between Kubernetes desired state and a Linux machine.


C.21 Break It - Stop the kubelet

On:

k8s-worker

stop the kubelet:

sudo systemctl stop kubelet

Back on the control plane:

kubectl get nodes -w

The node eventually stops reporting healthy status.

Inspect:

kubectl describe node k8s-worker

Now restore it:

sudo systemctl start kubelet

Watch again:

kubectl get nodes -w

Important:

containerd may still have containers

but

kubelet is no longer reconciling the node

Stopping kubelet does not mean every process on the machine instantly disappears.


C.22 Break It - Stop containerd

On the worker:

sudo systemctl stop containerd

Watch the kubelet:

journalctl -u kubelet -f

The dependency is:

kubelet
   |
   | CRI
   v
containerd

No runtime means the kubelet cannot manage containers correctly.

Restore it:

sudo systemctl start containerd
sudo systemctl restart kubelet

Then:

kubectl get nodes
kubectl get pods -A

C.23 Break It - Delete a Cilium Pod

Check:

kubectl get daemonset cilium \
  -n kube-system

List agents:

kubectl get pods \
  -n kube-system \
  -l k8s-app=cilium \
  -o wide

Delete one:

kubectl delete pod \
  -n kube-system \
  <cilium-pod-name>

Watch:

kubectl get pods \
  -n kube-system \
  -l k8s-app=cilium \
  -w

The DaemonSet controller notices the mismatch and replaces it.

Delete agent
    |
    v
Desired != Actual
    |
    v
DaemonSet controller
    |
    v
Replacement agent

Same reconciliation model.

Lower in the stack.


C.24 Break It - Remove the Scheduler

This is intentionally destructive.

Do it only in this disposable lab.

On k8s-control:

sudo mv \
  /etc/kubernetes/manifests/kube-scheduler.yaml \
  /tmp/kube-scheduler.yaml

Watch:

kubectl get pods \
  -n kube-system \
  -l component=kube-scheduler

Now create:

kubectl run no-scheduler \
  --image=nginx:1.27-alpine

Inspect:

kubectl get pod no-scheduler \
  -o wide

It should remain:

Pending

Why?

kubectl
   |
   v
API server
   |
   v
Pod object exists
   |
   X
scheduler
   |
   X
spec.nodeName

The API server can accept the object while another critical control-plane component is unavailable.

Restore:

sudo mv \
  /tmp/kube-scheduler.yaml \
  /etc/kubernetes/manifests/kube-scheduler.yaml

Watch:

kubectl get pod no-scheduler -w

Eventually:

Pending -> Running

Cleanup:

kubectl delete pod no-scheduler

This is a useful distinction:

API availability

is not the same as:

complete cluster reconciliation

C.25 Inspect a Scheduler Decision

Create:

kubectl run scheduled \
  --image=nginx:1.27-alpine

Query:

kubectl get pod scheduled \
  -o jsonpath='{.spec.nodeName}{"\n"}'

The path was:

Pod submitted
     |
     v
API server
     |
     v
scheduler
     |
     | chooses node
     v
spec.nodeName
     |
     v
kubelet on chosen node
     |
     v
containerd
     |
     v
container

Cleanup:

kubectl delete pod scheduled

C.26 Where etcd Fits

Inspect:

kubectl get pod \
  -n kube-system \
  -l component=etcd \
  -o wide

Normal API clients do not talk directly to etcd.

Think:

kubectl
   |
   v
API server
   |
   v
etcd

Controllers also communicate through the API server.

The API server is the front door to cluster state.

That is why:

Kubernetes is an API-driven control system

is more useful than:

Kubernetes runs containers

C.27 Single-Node Lab Option

If you only have one Linux machine, most of this appendix still works.

The control-plane node normally has a taint preventing ordinary workloads from being scheduled there.

Check:

kubectl describe node k8s-control \
  | grep -A5 Taints

For a disposable single-node lab:

kubectl taint nodes k8s-control \
  node-role.kubernetes.io/control-plane-

Now normal workloads may be scheduled there.

This is useful for learning.

It is not our preferred production topology.


C.28 Drain and Uncordon a Node

Before maintenance:

kubectl drain k8s-worker \
  --ignore-daemonsets \
  --delete-emptydir-data

Check:

kubectl get nodes

The worker should show:

SchedulingDisabled

Inspect workload placement:

kubectl get pods -o wide

Return it to service:

kubectl uncordon k8s-worker

Check:

kubectl get nodes

This is different from simply switching the server off.

You declared operational intent to Kubernetes first.


C.29 Add Another Worker

Generate a new join command:

kubeadm token create \
  --print-join-command

Run it with sudo on another prepared worker.

Then:

kubectl get nodes

Cilium runs as a DaemonSet, so an agent will be scheduled onto the new node automatically.

Again:

new Node
   |
   v
DaemonSet desired state changes
   |
   v
Cilium agent appears

C.30 Reset a Worker

Drain it first:

kubectl drain k8s-worker \
  --ignore-daemonsets \
  --delete-emptydir-data \
  --force

On the worker:

sudo kubeadm reset -f

Remove CNI configuration left on disk:

sudo rm -rf /etc/cni/net.d

Back on the control plane:

kubectl delete node k8s-worker

Check:

kubectl get nodes

kubeadm reset is a best-effort reset.

It does not promise to return the operating system to its exact pre-Kubernetes state.

That is one reason disposable VMs make excellent learning environments.


C.31 Reset the Whole Cluster

On each worker:

sudo kubeadm reset -f
sudo rm -rf /etc/cni/net.d

On the control plane:

sudo kubeadm reset -f
sudo rm -rf /etc/cni/net.d

Remove local kubectl credentials:

rm -rf "$HOME/.kube"

For a disposable lab, destroying and recreating the VMs is cleaner still.

That is not cheating.

Reproducible infrastructure is preferable to mysterious state.


C.32 What kubeadm Did - and Did Not Do

kubeadm helped bootstrap:

certificates
kubeconfigs
control-plane static Pods
etcd
bootstrap tokens
kubelet configuration
RBAC bootstrap
cluster configuration

It did not choose:

our container runtime
our production topology
our CNI
our storage implementation
our Gateway implementation
our observability platform
our application workloads

kubeadm creates the Kubernetes foundation.

It does not create an entire application platform.


C.33 The Whole Bootstrap Sequence

Linux
  |
  v
containerd
  |
  | CRI
  v
kubelet
  |
  ^
  |
kubeadm
  |
  v
control plane
  |
  +--> API server
  +--> etcd
  +--> scheduler
  +--> controller-manager
  |
  v
Kubernetes exists
  |
  | but
  v
Nodes NotReady
  |
  | install
  v
Cilium
  |
  +--> CNI
  +--> eBPF dataplane
  +--> kube-proxy replacement
  |
  v
Nodes Ready
  |
  v
Pods / Services / Deployments

Compare that to:

kind create cluster

kind was not fake Kubernetes.

It was automating most of this journey for us.


C.34 The Model to Remember

The CKAD cookbook started with:

Desired state
     |
     v
Kubernetes API
     |
     v
Controller
     |
     v
Actual state

We can now expand it:

                      desired state
                            |
                            v
                      Kubernetes API
                            |
               +------------+------------+
               |            |            |
               v            v            v
           scheduler    controllers    Cilium
               |            |            |
               +------------+------------+
                            |
                            v
                         kubelet
                            |
                            v
                        containerd
                            |
                            v
                      Linux kernel
                            |
                            v
                      actual state

Same Kubernetes model.

We simply followed it further down.


C.35 Next - Cilium and Gateway API

Our cluster now has:

Kubernetes 1.35
Cilium 1.20.1
kube-proxy replacement

That is exactly the substrate we want for:

appendix-d-cilium-gateway-api.md

Next we will take:

Service
NetworkPolicy
GatewayClass
Gateway
HTTPRoute

and follow them through:

Kubernetes API
      |
      v
Cilium
      |
      v
eBPF / Envoy
      |
      v
actual traffic

That is where the application-facing Kubernetes API starts turning into a modern platform dataplane.


Reference Versions

This appendix was written against:

Kubernetes: 1.35
Cilium:     1.20.1

The Kubernetes 1.35 packages come from the versioned pkgs.k8s.io repository.

Cilium 1.20.1 officially supports Kubernetes 1.35.

Cilium's kube-proxy replacement documentation explicitly supports bootstrapping kubeadm with:

--skip-phases=addon/kube-proxy

Always check the matching upstream documentation before using these instructions for a newer Kubernetes or Cilium release.

Appendix - From RBAC to Multi-Tenant Kubernetes [DEV] [DEEP DIVE]

This companion starts where the main cookbook's RBAC chapter stops.

The question is no longer only:

What may this identity do inside Kubernetes?

It is now:

What boundary should exist between tenants in the first place?

This is platform-engineering material, not CKAD material.

The goal is not to memorise product-specific commands. We will use ordinary RBAC, vCluster and Kamaji as hands-on experiments to make API, control-plane and worker isolation concrete.

By the end we will have built three different models:

one shared API
    + namespace RBAC

separate tenant API
    + shared workers

separate tenant API
    + hosted control plane
    + independently joined workers

The vCluster lab assumes you are continuing from the cookbook's existing kind-ckad cluster. The Kamaji lab deliberately creates a second kind cluster so that experimentation does not disturb the CKAD environment.


1. Keep the Boundaries Separate

A useful tenancy model has several layers.

Layer 1 - identity and API authorization
----------------------------------------
authentication
RBAC
admission

"Can Alice perform this API operation?"


Layer 2 - API / control-plane isolation
---------------------------------------
shared kube-apiserver
or
separate tenant kube-apiservers

"Does Alice even share a Kubernetes API with Bob?"


Layer 3 - workload isolation
----------------------------
namespaces
Pod security
scheduler policy
taints / affinity
separate worker pools
runtime sandboxing

"Can their workloads interfere on compute?"


Layer 4 - network, storage and infrastructure isolation
-------------------------------------------------------
NetworkPolicy
CNI / VPC / VLAN / VRF
CSI and storage policy
VM boundaries
bare-metal allocation
cloud IAM

"What underlying infrastructure do tenants share?"


Layer 5 - provider management plane
-----------------------------------
management cluster
cluster lifecycle controllers
provisioning systems
hardware / cloud APIs

"Who is allowed to change the infrastructure itself?"

No single layer replaces all the others.

For example:

separate API servers
      !=
separate kernels
NoSchedule taint
      !=
authorization boundary
Kubernetes cluster-admin
      !=
SSH root on the machine

We are going to prove those distinctions rather than only state them.


2. Lab Zero: One Shared Cluster + Namespace RBAC

Before adding virtual or hosted control planes, establish the baseline.

Make sure we are on the cookbook cluster:

kubectl config use-context kind-ckad

Save the provider/host context. We will use this later because vCluster changes our current context for us:

export HOST_CONTEXT=$(kubectl config current-context)
echo "$HOST_CONTEXT"

Expected:

kind-ckad

Create a namespace for Alice:

kubectl create namespace shared-alice

Create a small developer Role:

cat <<'EOF' | kubectl apply -f -
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: developer
  namespace: shared-alice
rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log", "services"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["apps"]
    resources: ["deployments"]
    verbs: ["get", "list", "watch", "create", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: alice-developer
  namespace: shared-alice
subjects:
  - kind: User
    name: alice
    apiGroup: rbac.authorization.k8s.io
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: developer
EOF

Because our kind administrator may impersonate users, we can ask the API server what Alice would be allowed to do:

kubectl auth can-i get pods \
  --as=alice \
  -n shared-alice

Expected:

yes

She may create a Deployment in her namespace:

kubectl auth can-i create deployments.apps \
  --as=alice \
  -n shared-alice

Expected:

yes

But she cannot create arbitrary namespaces:

kubectl auth can-i create namespaces \
  --as=alice

Expected:

no

Nor delete Nodes:

kubectl auth can-i delete nodes \
  --as=alice

Expected:

no

The model is:

Alice
  |
  v
same kube-apiserver as everyone else
  |
  v
RBAC
  |
  +-- shared-alice namespace    some access
  +-- cluster-scoped objects    mostly no access
  +-- other tenant namespaces   no access unless granted

This is a legitimate tenancy model for many internal platforms.

But Alice still talks to the same Kubernetes API as everybody else.

That is the limitation we will explore next.


3. Why Give a Tenant Another Kubernetes API?

Suppose Alice needs broad Kubernetes freedom.

Inside a normal shared cluster, this would be dangerous:

Alice
  |
  v
cluster-admin
  |
  v
shared cluster

cluster-admin is intentionally enormous.

A platform provider often wants something different:

Alice
  |
  v
Alice's Kubernetes API
  |
  | cluster-admin is okay HERE
  v
Alice's tenant cluster


provider
  |
  v
provider Kubernetes API
  |
  | Alice does NOT get these credentials
  v
provider infrastructure

This changes the question from:

How carefully can I restrict Alice inside my cluster?

into:

Why should Alice be an administrator of my cluster at all?

That distinction is central to virtual clusters, hosted control planes and many managed Kubernetes systems.


4. Lab: Give Alice a vCluster

vCluster gives a tenant its own Kubernetes API while allowing the platform operator to host the tenant control plane on existing infrastructure.

We will begin with its shared-node model because we can run the whole experiment on our existing kind cluster.

4.1 Install the vCluster CLI

On macOS with Homebrew:

brew install loft-sh/tap/vcluster

Verify it:

vcluster --version

Make sure the host context is still our cookbook cluster:

kubectl config use-context "$HOST_CONTEXT"

4.2 Create Alice's tenant cluster

Create a tenant cluster called alice inside the provider namespace tenant-alice:

vcluster create alice \
  --namespace tenant-alice

The vCluster CLI automatically connects you to the new tenant cluster when creation completes.

Check the current context:

kubectl config current-context

Then ask the API server what namespaces exist:

kubectl get namespaces

You should see an ordinary Kubernetes-looking namespace view such as:

default
kube-node-lease
kube-public
kube-system

Alice is not looking at the provider cluster's namespace list.

She is talking to another Kubernetes API.

Conceptually:

                         kind-ckad
                    provider Kubernetes
                           API
                            |
                            v
                     tenant-alice
                            |
                      vCluster CP
                    + API server
                    + controllers
                    + datastore
                    + syncer
                            |
                            v
                          Alice

5. Alice Can Be an Administrator of Her Cluster

By default, the kubeconfig generated by vcluster connect uses tenant administrator credentials.

Prove what our current credential can do:

kubectl auth can-i create namespaces

Expected:

yes

Try another cluster-scoped operation:

kubectl auth can-i create clusterrolebindings.rbac.authorization.k8s.io

Expected:

yes

Make the permission set explicit:

kubectl auth can-i '*' '*'

With tenant administrator credentials this should report:

yes

Now compare that with Alice against the provider API:

kubectl --context "$HOST_CONTEXT" \
  auth can-i delete nodes \
  --as=alice

Expected:

no

This is the important model:

              TENANT API

Alice's tenant credential
        |
        v
  tenant cluster-admin
        |
        v
      YES


             PROVIDER API

Alice
  |
  v
provider RBAC
  |
  +-- delete Nodes?                NO
  +-- create ClusterRoleBinding?   NO

cluster-admin is not a magical global property attached to a human being.

It is authorization against a particular Kubernetes API.

In this lab we use vCluster-generated credentials and Kubernetes impersonation to make the boundary obvious. A production platform might authenticate the same human through OIDC or another identity provider on both APIs and assign different authorization in each one.


6. Issue a Less Powerful Tenant Credential

A tenant does not need to give every user administrator access either.

vCluster can generate a kubeconfig backed by a ServiceAccount and bind that ServiceAccount to a tenant-local ClusterRole.

Create/connect using a read-only tenant identity:

vcluster connect alice \
  --namespace tenant-alice \
  --service-account kube-system/alice-viewer \
  --cluster-role view

Now test it:

kubectl auth can-i get pods --all-namespaces

Expected:

yes

But:

kubectl auth can-i create namespaces

Expected:

no

And:

kubectl auth can-i create deployments.apps

Expected:

no

So there are now two authorization domains in play:

provider API
    |
    +-- provider decides who may manage vCluster infrastructure

Alice tenant API
    |
    +-- tenant decides who is admin, viewer, developer, etc.

Reconnect with the default tenant administrator credential for the next exercise:

vcluster connect alice \
  --namespace tenant-alice

7. Create a Workload Inside the Tenant

Create another namespace from inside Alice's cluster:

kubectl create namespace apps

Deploy nginx:

kubectl create deployment web \
  --image=nginx:1.27-alpine \
  -n apps

Watch it:

kubectl get pods \
  -n apps \
  -o wide

From Alice's perspective this is just Kubernetes:

Alice
  |
  v
POST Deployment to tenant API
  |
  v
Deployment controller
  |
  v
ReplicaSet
  |
  v
Pod

But shared-node vCluster has another layer underneath.

Let's look behind the curtain.


8. Look at the Same Workload From the Provider Side

Disconnect from the tenant:

vcluster disconnect

If necessary, explicitly restore the host context:

kubectl config use-context "$HOST_CONTEXT"

Inspect the namespace where the vCluster lives:

kubectl get pods \
  -n tenant-alice \
  -o wide

You should see the tenant control-plane Pod and translated workloads.

A workload created inside the tenant might appear with a rewritten host-side name similar to:

web-xxxxxxxxxx-yyyyy-x-apps-x-alice

The exact generated name is not important.

The important relationship is:

TENANT VIEW
-----------
namespace: apps
pod:       web-xxxxx

        |
        | sync / translation
        v

PROVIDER VIEW
-------------
namespace: tenant-alice
pod:       rewritten host-side name

In shared-node mode, the vCluster syncer translates workload resources onto the provider cluster so the provider scheduler and kubelets can run them.

The tenant does not need credentials for that provider API.


9. Follow One Pod Through the vCluster Boundary

Reconnect:

vcluster connect alice \
  --namespace tenant-alice

Look at the tenant Pod:

kubectl get pods \
  -n apps \
  -o wide

Switch back to the host:

vcluster disconnect
kubectl config use-context "$HOST_CONTEXT"

Then:

kubectl get pods \
  -n tenant-alice \
  -o wide

You have just observed the complete path:

kubectl
   |
   v
Alice tenant kube-apiserver
   |
   v
tenant Pod object
   |
   v
vCluster syncer
   |
   v
provider kube-apiserver
   |
   v
translated Pod
   |
   v
provider scheduler
   |
   v
provider node / kubelet

When the provider-side Pod changes status, vCluster synchronizes that observation back to the tenant API.

So even here we are still using the same control-loop model from Chapter 1:

desired tenant object
       |
       v
translation / reconciliation
       |
       v
provider object
       |
       v
actual workload
       |
       v
status flows back

10. Shared API Isolation Is Not Worker Isolation

We have proven that Alice has a separate Kubernetes API.

But in this default shared-node model, her nginx workload ultimately runs on the same provider worker infrastructure as other workloads.

Tenant A API         Tenant B API
     |                    |
     v                    v
 translated Pods     translated Pods
       \                  /
        \                /
         v              v
          provider nodes
                |
                v
           shared kernel

So:

separate tenant API       yes
separate tenant RBAC      yes
separate host namespace   yes
separate physical node    not necessarily
separate kernel           no, not in shared-node mode

Current vCluster guidance treats shared nodes as appropriate for trusted tenants such as internal development, CI and testing.

It explicitly does not treat this model as the worker security boundary for untrusted external tenants with arbitrary Kubernetes workload access.

That is not a defect in RBAC.

It is a different layer of the architecture.


11. A Small Shared-Node Hardening Recipe

Even for trusted tenants, the provider should not assume the separate API is sufficient on its own.

For example, vCluster can create host-side NetworkPolicy around the tenant workload namespace:

policies:
  networkPolicy:
    enabled: true

A more opinionated baseline can look like:

policies:
  podSecurityStandard: restricted
  resourceQuota:
    enabled: true
  limitRange:
    enabled: true
  networkPolicy:
    enabled: true
    workload:
      publicEgress:
        enabled: false
sync:
  toHost:
    pods:
      useSecretsForSATokens: true

Save that as vcluster-hardening.yaml, then apply it as an upgrade:

vcluster create alice \
  --namespace tenant-alice \
  --upgrade \
  --connect=false \
  -f vcluster-hardening.yaml

But do not confuse configuration with enforcement.

NetworkPolicy object exists
       !=
CNI actually enforces NetworkPolicy

The host CNI must implement the policy.

Likewise:

Pod Security Standard
       !=
separate kernel
resource quota
       !=
strong workload isolation

Hardening improves the shared-node model. It does not change its fundamental trust boundary.

restricted Pod Security may also break workloads that assume root privileges. Treat that as useful feedback about the workload rather than blindly weakening the platform baseline.


12. What Would vCluster Private Nodes Change?

vCluster also supports a model where tenant workers are not shared with the provider worker pool.

This requires vCluster Platform and separate Linux worker machines, so it is not part of our simple kind-ckad lab.

The configuration begins with something like:

privateNodes:
  enabled: true
  vpn:
    enabled: true
networking:
  podCIDR: 10.64.0.0/16
  serviceCIDR: 10.128.0.0/16

A private-node tenant cluster is created with:

vcluster create alice-private \
  --namespace tenant-alice-private \
  --values vcluster-private.yaml

An interesting thing happens immediately:

kubectl get nodes

Expected before joining any workers:

No resources found.

That is not a broken cluster.

control plane exists
       !=
worker nodes exist

Once private workers are joined:

Alice tenant API
      |
      v
Alice private workers

Bob tenant API
      |
      v
Bob private workers

The provider-hosted control plane and tenant compute have become separate choices.


13. Clean Up the vCluster Lab

Before moving to Kamaji, delete Alice's vCluster:

kubectl config use-context "$HOST_CONTEXT"
vcluster delete alice \
  --namespace tenant-alice

Remove the baseline namespace too:

kubectl delete namespace shared-alice

Our original CKAD cluster remains intact.


14. Kamaji: Host the Tenant Control Plane, Not Its Workers

Kamaji approaches the broad problem differently.

Instead of every tenant owning dedicated control-plane VMs, Kamaji runs upstream Kubernetes control-plane components as workloads inside a provider-operated Management Cluster.

                      Management Cluster

                        Kamaji operator
                              |
             +----------------+----------------+
             |                                 |
             v                                 v
      Tenant A control plane            Tenant B control plane
      kube-apiserver                    kube-apiserver
      controller-manager                controller-manager
      scheduler                         scheduler
             |                                 |
             v                                 v
      tenant A API endpoint              tenant B API endpoint
             |                                 |
             v                                 v
      tenant A workers                   tenant B workers

This creates a clean separation:

control-plane lifecycle
        |
        v
provider management cluster

worker lifecycle
        |
        v
VMs / bare metal / Cluster API / other provisioning

Let's build one.


15. Lab: Create a Kamaji Management Cluster on kind

Kamaji's official kind walkthrough is intended for development and learning only.

We will create a separate kind cluster named kamaji.

Prerequisites:

docker
kind
kubectl
helm
jq

Create the management cluster:

kind create cluster --name kamaji

Verify:

kubectl config current-context

Expected:

kind-kamaji

Save the context:

export KAMAJI_CONTEXT=$(kubectl config current-context)

16. Install cert-manager

Kamaji uses admission webhooks and relies on cert-manager for their TLS certificates.

Add the repository:

helm repo add jetstack https://charts.jetstack.io
helm repo update

Install cert-manager:

helm upgrade --install cert-manager jetstack/cert-manager \
  --namespace cert-manager \
  --create-namespace \
  --version v1.18.3 \
  --set crds.enabled=true

Watch it become ready:

kubectl get pods \
  -n cert-manager \
  -w

Press Ctrl-C once the Pods are Ready.

The pinned version above follows the current Kamaji kind walkthrough when this appendix was written. If upstream moves on, prefer the dependency versions in the current official guide.


17. Install MetalLB for Tenant API Endpoints

A Kamaji Tenant Control Plane needs an API endpoint.

In this kind lab, the official walkthrough uses MetalLB to provide LoadBalancer addresses on the Docker kind network.

Install MetalLB:

kubectl apply -f \
  https://raw.githubusercontent.com/metallb/metallb/v0.15.3/config/manifests/metallb-native.yaml

Wait for its controller:

kubectl wait \
  --namespace metallb-system \
  --for=condition=Available \
  deployment/controller \
  --timeout=120s

Get the IPv4 gateway of the Docker kind network:

GW_IP=$(docker network inspect kind \
  | jq -r '.[0].IPAM.Config[] | select(.Gateway | test("^[0-9]+\\.[0-9]+\\.[0-9]+\\.[0-9]+$")) | .Gateway')

echo "$GW_IP"

Derive the first two octets:

NET_IP=$(echo "$GW_IP" \
  | sed -E 's|^([0-9]+\.[0-9]+)\..*$|\1|g')

echo "$NET_IP"

Create an address pool near the top of that Docker subnet:

cat <<EOF | sed -E "s|172.19|${NET_IP}|g" | kubectl apply -f -
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
  name: kind-ip-pool
  namespace: metallb-system
spec:
  addresses:
    - 172.19.255.200-172.19.255.250
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
  name: kind-l2
  namespace: metallb-system
EOF

This is lab networking, not a production LoadBalancer design.


18. Install Kamaji

Add the Clastix chart repository:

helm repo add clastix https://clastix.github.io/charts
helm repo update

Install Kamaji:

helm upgrade --install kamaji clastix/kamaji \
  --namespace kamaji-system \
  --create-namespace \
  --set 'resources=null' \
  --version 0.0.0+latest

Watch the installation:

kubectl get pods \
  -n kamaji-system \
  -w

Then verify that Kamaji extended the Kubernetes API:

kubectl get crds \
  | grep -i kamaji

This should already look familiar:

install operator
      |
      v
new CRD appears
      |
      v
new declarative API available

19. Create a TenantControlPlane

Apply Kamaji's current sample TenantControlPlane:

kubectl apply -f \
  https://raw.githubusercontent.com/clastix/kamaji/master/config/samples/kamaji_v1alpha1_tenantcontrolplane.yaml

Watch it reconcile:

kubectl get tcp -w

Eventually you should see a state similar to:

NAME      VERSION   STATUS   CONTROL-PLANE ENDPOINT   KUBECONFIG
k8s-133   ...       Ready    ...:6443                 k8s-133-admin-kubeconfig

Stop the watch with Ctrl-C.

Inspect what Kamaji created in the management cluster:

kubectl get tcp,deploy,pods,svc

The conceptual chain is:

TenantControlPlane
       |
       v
Kamaji controller
       |
       v
Deployment / Service / certificates / datastore state
       |
       v
running tenant kube-apiserver
       + controller-manager
       + scheduler

This is the same spec -> controller -> status model we began the entire cookbook with.

Kamaji is simply using it to create Kubernetes control planes.


20. Retrieve the Tenant kubeconfig

Kamaji writes the tenant administrator kubeconfig into a Secret named after the TenantControlPlane.

For the sample tenant:

kubectl get secret k8s-133-admin-kubeconfig

Extract it:

kubectl get secret k8s-133-admin-kubeconfig \
  -o jsonpath='{.data.admin\.conf}' \
  | base64 -d \
  > /tmp/kamaji-tenant.conf

Inspect its target without changing your main kubeconfig:

kubectl --kubeconfig=/tmp/kamaji-tenant.conf \
  config view --minify

We now have separate credentials for separate APIs:

~/.kube/config
    -> provider / management clusters

/tmp/kamaji-tenant.conf
    -> tenant Kubernetes API

21. Talk Directly to the Tenant API

Try:

kubectl --kubeconfig=/tmp/kamaji-tenant.conf \
  cluster-info

Then:

kubectl --kubeconfig=/tmp/kamaji-tenant.conf \
  get namespaces

Now the important command:

kubectl --kubeconfig=/tmp/kamaji-tenant.conf \
  get nodes

Expected:

No resources found.

This is not an error.

We successfully have:

kube-apiserver          yes
controller-manager      yes
scheduler               yes
Kubernetes API          yes
worker node             no

That gives us a clean mental model:

control plane exists
        !=
compute exists

The tenant has a Kubernetes cluster control plane before it has somewhere to run application Pods.


22. macOS / Docker Desktop Note

On native Linux, the MetalLB address on the Docker kind network is often directly reachable from the host.

On macOS with Docker Desktop, the Docker bridge network may not be routed directly into macOS.

That means this can happen:

TenantControlPlane     Ready
LoadBalancer IP        allocated
kubeconfig             valid

but

kubectl from macOS     cannot route to that Docker-network IP

Do not interpret that as Kamaji reconciliation failing.

First confirm from the management side:

kubectl --context "$KAMAJI_CONTEXT" get tcp

and:

kubectl --context "$KAMAJI_CONTEXT" get svc

The official Kamaji kind guide calls out this Docker bridge/macOS case and suggests running tenant API checks from an environment that can reach the kind Docker network, including the kind control-plane container when appropriate.

The networking lesson is useful in its own right:

API exists
   !=
my current machine has a route to that API

Production Kamaji environments expose the API deliberately with normal LoadBalancer, DNS, Gateway or equivalent networking rather than relying on Docker Desktop bridge routing.


23. What Would Joining a Worker Look Like?

Kamaji deliberately does not create tenant worker machines for you.

Workers might come from:

cloud VMs
bare metal
Cluster API
an infrastructure platform
manual provisioning

Once a Linux machine has the required container runtime, kubelet and kubeadm components installed, the tenant control plane can generate an ordinary kubeadm join command.

Conceptually:

kubeadm --kubeconfig=/tmp/kamaji-tenant.conf \
  token create \
  --print-join-command

That produces the familiar shape:

kubeadm join <tenant-api>:6443 \
  --token ... \
  --discovery-token-ca-cert-hash ...

Run that command on the intended worker machine.

Then:

kubectl --kubeconfig=/tmp/kamaji-tenant.conf \
  get nodes

would move from:

No resources found.

towards something like:

NAME        STATUS   ROLES    AGE
worker-01   Ready    <none>   30s

The architecture is therefore:

Kamaji management cluster
       |
       +-- tenant kube-apiserver
       +-- tenant controller-manager
       +-- tenant scheduler

                 |
                 | Kubernetes API
                 v

        independently provisioned
             worker nodes

Kamaji can also integrate with Cluster API so that worker lifecycle becomes declarative rather than a manual kubeadm join exercise.


24. The Provider Does Not Need to Give Away Its Management Cluster

Return to the original access-control question.

A tenant can receive:

/tmp/kamaji-tenant.conf

which grants access to:

Tenant A kube-apiserver

without receiving credentials for:

kind-kamaji management API

and without receiving:

SSH access to the management hosts

Those are different trust boundaries:

                     TENANT
                        |
                        v
                 Tenant API endpoint
                        |
                        v
               tenant Kubernetes RBAC

================================================
                  provider boundary
================================================

               Management Cluster API
                        |
               Kamaji / controllers
                        |
                host infrastructure

The tenant may be cluster-admin above the line.

That does not imply they are administrator below it.


25. vCluster and Kamaji: Compare What We Actually Built

We have now used both models instead of only drawing them.

Question Shared namespace + RBAC vCluster shared nodes Kamaji
Tenant has separate kube-apiserver No Yes Yes
Tenant-local RBAC No separate API Yes Yes
Tenant can have broad admin without provider-cluster admin Not as shared cluster-admin Yes Yes
Provider hosts tenant control plane Shared provider CP Yes Yes
Tenant workload becomes a Pod on provider cluster Directly Yes, translated/synced No - worker joins tenant API
Workers can be shared Yes Yes Not the core model
Dedicated workers possible By platform design Yes, private nodes Yes
Worker lifecycle bundled into basic control-plane creation N/A Depends on worker model No
Provider management credentials required by tenant Same shared API No No

The important difference is not which product has the longest feature list.

It is where each architecture places the boundary.


26. Revisit the Control-Plane Node Question

The original question was roughly:

How do I let users operate Kubernetes without letting them control the control-plane nodes?

There were actually four questions hiding inside it.

26.1 May Alice modify the Node API object?

That is Kubernetes authorization:

kubectl auth can-i delete nodes \
  --as=alice

RBAC answers this.

26.2 May Alice's Pod run on a particular node?

That is scheduling and admission policy:

nodeSelector
affinity
taints / tolerations
admission

A control-plane NoSchedule taint is useful placement policy.

It is not a hard authorization boundary if Alice is allowed to submit arbitrary tolerations.

26.3 May Alice log into the machine?

That is infrastructure access:

cloud IAM
SSH keys
VPN / firewall
bastions
OS users

Kubernetes RBAC does not remove Alice's SSH key.

26.4 Does Alice need to see the provider API at all?

That is the architectural question we explored here:

shared cluster
    |
    +-- Alice and provider use same API

versus

tenant control plane
    |
    +-- Alice uses tenant API
    +-- provider API remains provider-only

This fourth option can dramatically reduce how much permission engineering has to happen inside the provider cluster.


27. Strong Tenancy Is Still a Stack

Even a separate tenant kube-apiserver does not solve every problem.

For an untrusted external tenant, a design may need something closer to:

Tenant identity
    |
    v
Tenant API server
    |
    +-- tenant-local RBAC
    +-- admission / policy
    |
    v
Dedicated or strongly isolated compute
    |
    +-- runtime security
    +-- device isolation
    |
    v
Tenant network boundary
    |
    +-- NetworkPolicy
    +-- VPC / VLAN / routing policy
    |
    v
Tenant storage boundary
    |
    +-- CSI / volume policy
    +-- encryption / credentials
    |
    v
Provider management plane
    |
    X tenant credentials do not cross this boundary

Do not collapse those controls into one mental bucket.

RBAC
  != NetworkPolicy
  != scheduler placement
  != node isolation
  != VM isolation
  != storage isolation
  != host IAM

28. A Practical Platform Decision Sequence

When designing a platform, ask these questions in order.

1. Are the tenants mutually trusted?
        |
        +-- yes -> shared cluster + namespace/RBAC may be sufficient
        |
        +-- no  -> continue

2. Do tenants need broad Kubernetes administration?
        |
        +-- yes -> consider separate tenant API/control-plane boundaries

3. Can tenants execute arbitrary workloads?
        |
        +-- yes -> decide what worker/kernel isolation is required

4. Can workloads communicate across tenants?
        |
        +-- no -> enforce network boundaries

5. Can storage or devices be shared safely?
        |
        +-- design CSI/device/IOMMU/etc. boundaries accordingly

6. Can tenant credentials reach the provider management plane?
        |
        +-- they generally should not

Then pick technology.

Do not begin with:

"We should use vCluster."

Begin with:

"What boundary are we trying to create?"

29. Connect This Back to the Main Cookbook

Nearly everything in this appendix is built from concepts we already learned.

Authentication / RBAC
    -> Chapter 23

SecurityContext / Pod security
    -> Chapter 24

NetworkPolicy
    -> Chapter 25

Taints / scheduling eligibility
    -> scheduling chapters

CRDs
    -> Chapter 33

Controllers
    -> Chapter 34

spec / status / reconciliation
    -> Chapter 1 and everywhere else

The platform gets more sophisticated.

The primitive stays familiar:

API object
   |
   v
desired state
   |
   v
controller
   |
   v
lower-level infrastructure
   |
   v
status

A Deployment reconciles Pods.

A vCluster syncer reconciles tenant resources into provider resources.

Kamaji reconciles a TenantControlPlane into a running Kubernetes control plane.

Cluster API can reconcile a cluster specification into worker machines.

Once the control-loop model is clear, these systems stop looking magical.


30. Cleanup the Kamaji Lab

When finished, return to the CKAD context:

kubectl config use-context kind-ckad

Delete the entire Kamaji learning environment:

kind delete cluster --name kamaji

Remove the temporary tenant kubeconfig:

rm -f /tmp/kamaji-tenant.conf

Your original cookbook cluster remains available.


31. What You Should Remember

If you retain only a few things from this appendix, make them these:

RBAC answers:
"What can this identity do against this Kubernetes API?"
A separate tenant API answers:
"Why should this tenant be an administrator of my provider API at all?"
Separate API servers do not automatically mean separate workers or kernels.
cluster-admin is relative to a cluster/API boundary.
control plane exists != worker nodes exist

and:

a strong tenant boundary is normally a stack of controls,
not one Kubernetes object or one product.

32. Official References

These projects evolve quickly. The commands and architecture above were aligned with their current documentation when this appendix was written; use current upstream docs when making a real platform decision.

vCluster

Kamaji

Kubernetes

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment