You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Application
|
v
Kubernetes API
|
+--> Deployment controller
|
+--> Service objects
|
+--> NetworkPolicy
|
v
Cilium
|
+--> CNI
|
+--> NetworkPolicy enforcement
|
+--> kube-proxy replacement
|
v
eBPF dataplane
We are going to add L7 routing:
Client
|
v
Gateway
|
v
Cilium Gateway controller
|
v
Envoy
|
v
HTTPRoute
|
v
Service
|
v
Pod
Gateway API separates concerns more cleanly than the original Ingress model.
The important resource chain is:
GatewayClass
|
v
Gateway
|
v
HTTPRoute
|
v
Service
|
v
Pod
D.2 Why Gateway API
Ingress gives us roughly:
Ingress
|
v
Controller
|
v
Service
Gateway API makes roles explicit.
Platform owns
----------------
GatewayClass
Gateway
Application team owns
---------------------
HTTPRoute
Service
Deployment
Think:
GatewayClass
-> Which implementation handles this?
Gateway
-> Where and how does traffic enter?
HTTPRoute
-> Where should HTTP requests go?
Service
-> Which backend Pods receive them?
That separation becomes particularly useful in shared clusters.
You should see that the route cannot completely resolve its references.
Query:
kubectl get httproute api \
-o jsonpath='{range .status.parents[*].conditions[*]}{.type}={.status}{" reason="}{.reason}{" message="}{.message}{"\n"}{end}'
The important difference from the simple canary in the CKAD cookbook is:
Old approach
------------
9 Pods stable
1 Pod canary
Service selects all 10
Gateway API approach
--------------------
HTTPRoute
|
+-- weight 90 -> Service v1
|
+-- weight 10 -> Service v2
We normally leave the Gateway API CRDs and Cilium installation in place because they are cluster-level platform components.
For a fully disposable lab, reset the cluster using Appendix C instead.
D.31 The Full Request Path
We can now describe a request from outside the cluster.
curl
|
| Host: api.example.test
v
Linux node :8080
|
v
Cilium / Envoy
|
| listener selected
v
Gateway/public
|
| route selected
v
HTTPRoute/api
|
| backendRef
v
Service/api-v1
|
| endpoints
v
EndpointSlice
|
v
Pod IP
|
v
Cilium dataplane
|
v
container
At the same time:
Kubernetes API
|
+--> stores Gateway
|
+--> stores HTTPRoute
|
+--> stores Service
|
+--> stores EndpointSlice
|
v
controllers observe
|
v
lower-level state changes
That is Kubernetes reconciliation expressed as networking.
D.32 The Architecture We Now Understand
At the beginning of the cookbook:
Service -> Pods
was enough.
Now we can expand it:
Kubernetes API
|
+----------------+----------------+
| | |
v v v
Gateway HTTPRoute Service
| | |
+--------+-------+ |
| |
v v
Cilium controller EndpointSlice
| |
v |
Envoy |
| |
+-----------+------------+
|
v
Cilium agent
|
v
eBPF
|
v
Linux kernel
|
v
Pod
And with observability:
packet / request
|
v
Cilium dataplane
|
+----> enforcement
|
+----> Hubble
|
v
observable flow
D.33 Portability vs Platform Power
The core cookbook deliberately favoured APIs such as:
Service
NetworkPolicy
Ingress
because they are portable Kubernetes APIs.
This appendix used:
Gateway API
which is portable across conformant Gateway implementations.
Portable API
|
+--> easier platform portability
|
+--> common Kubernetes mental model
Implementation-specific API
|
+--> richer platform capability
|
+--> tighter coupling to the implementation
The important thing is to know which side of the boundary you are on.
D.34 The Model to Remember
For networking:
Application intent
|
v
Kubernetes API
|
v
Controller / network implementation
|
v
Envoy / eBPF / Linux
|
v
actual traffic
For debugging:
API status
|
v
references
|
v
Services / endpoints
|
v
implementation status
|
v
observed flows
|
v
dataplane
Do not jump straight to the bottom.
Prove each layer.
D.35 Where You Are Now
The learning path has become:
kind
|
v
"Give me Kubernetes"
|
v
CKAD cookbook
|
v
"Teach me Kubernetes APIs"
|
v
kubeadm
|
v
"Show me where Kubernetes comes from"
|
v
Cilium
|
v
"Show me how networking is implemented"
|
v
Gateway API
|
v
"Give platform and application teams a modern routing API"
|
v
Hubble / eBPF
|
v
"Show me what the dataplane is actually doing"
At this point the cluster should feel much less magical.
The Kubernetes cookbook teaches you how Kubernetes reconciles desired state into running workloads.
This handbook takes the next step:
How does application code safely become desired state in a real cluster?
The goal is not to memorise Argo CD commands.
The goal is to understand a modern delivery system well enough that you can reason about it, debug it, secure it, and eventually build a platform around it.
We will build this path:
developer
|
| git push / pull request
v
application repository
|
| CI tests, builds, scans
v
OCI registry
|
| immutable image + digest
v
promotion pull request
|
v
GitOps repository
|
| reviewed desired-state change
v
Argo CD
|
| reconciliation
v
Kubernetes API
|
v
Argo Rollouts
|
| canary / blue-green / analysis / promotion
v
running application
By the end, we will have replaced the classic pipeline:
CI job
|
| kubectl apply
| cluster-admin kubeconfig
v
production
with:
CI
|
+-- builds an artifact
+-- signs / attests it
+-- proposes a desired-state change
|
v
Git pull request
|
| human / policy review
v
Git
|
| pulled by controller
v
cluster
Sections are marked:
[DEV] - application developer knowledge
[OPS] - operating delivery systems
[PLATFORM] - platform engineering and multi-team design
[DEEP DIVE] - concepts worth understanding beyond the immediate lab
The recurring teaching loop is the same as the main Kubernetes cookbook:
problem
|
v
mental model
|
v
small experiment
|
v
observe the controllers
|
v
change one thing
|
v
observe the consequence
|
v
break an assumption
|
v
explain why
Part I - Git Becomes Desired State
0. Lab Setup [DEV] [OPS]
This handbook assumes you completed enough of the Kubernetes cookbook to be comfortable with:
Deployments and Services
spec versus status
rollouts and ReplicaSets
RBAC
Kustomize
container images
kind
The examples assume the existing lab cluster:
kubectl config use-context kind-ckad
Check it:
kubectl get nodes
We will install platform components into their own namespaces rather than the cookbook namespace.
Useful local tools:
git
kubectl
docker
kind
For the complete hosted Git workflow, a GitHub account is convenient.
Optional but useful on macOS:
brew install gh argocd
Argo Rollouts will be installed later.
Two repositories, not one
We will eventually use two repositories:
web-app
|
+-- application source
+-- Dockerfile
+-- tests
+-- CI workflow
platform-gitops
|
+-- Kubernetes desired state
+-- environment overlays
+-- image digest promoted to each environment
+-- Argo CD Application definitions
That separation is intentional.
The application repository answers:
What software should we build?
The GitOps repository answers:
What exact version of that software should this environment run?
Argo should now remove resources that were previously managed but no longer exist in desired state.
Check:
kubectl get configmap web-banner \
-n web-dev
Expected:
NotFound
Pruning is powerful.
In production, think carefully before allowing automatic pruning of high-impact resources such as namespaces, storage resources, or shared infrastructure.
Argo also supports requiring explicit confirmation before certain prune/delete operations.
Part II - Build Once, Promote an Immutable Artifact
7. The Application Repository [DEV]
So far our GitOps repository deploys public nginx.
Now we will build our own application.
Create a separate repository:
mkdir -p ~/gitops-lab/web-app
cd~/gitops-lab/web-app
git init -b main
A newly-published GHCR package is private by default.
Your GitHub Actions runner can push it because the workflow has package permission.
Your kind node has no such credential.
If you merge the promotion without dealing with that boundary, you will rediscover:
ImagePullBackOff
For the shortest public lab, open the package settings in GitHub and change the container package visibility to Public. Public GHCR container packages can be pulled anonymously.
For a private-registry lab, leave it private and create a read credential. For example, using a token that can read the package:
Now Pods using that ServiceAccount can present the registry credential when pulling.
For production, do not manually sprinkle long-lived developer tokens through namespaces. Use a deliberate registry-auth pattern such as scoped robot/service credentials, a secret controller, or a node credential provider where appropriate.
This boundary is the same lesson as the kind image sidequest from the main cookbook:
image exists somewhere
!=
node is authorised and able to obtain it
The action output from docker/build-push-action includes the registry digest.
That digest is the value we actually want to promote.
Did CI pass?
Is the image signed?
Did vulnerability policy pass?
Does the digest correspond to the intended source commit?
Is this environment allowed to consume it?
Merge the PR.
Then watch Argo rather than running kubectl apply:
kubectl get application web-dev \
-n argocd \
-w
Watch Kubernetes underneath it:
kubectl get deployment,rs,pods \
-n web-dev \
-w
You should recognise the same rollout machinery from the Kubernetes cookbook.
Git changed.
Argo changed the Deployment desired state.
The Deployment controller changed ReplicaSets and Pods.
The hierarchy is:
Git commit
|
v
Argo Application
|
v
Deployment.spec.template
|
v
ReplicaSet
|
v
Pods
Prove which image is running
Ask Kubernetes:
kubectl get deployment web \
-n web-dev \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
You should see the digest-pinned image reference.
Now Git history and cluster state can be connected directly.
14. Why the Pull Request Is Part of the Control Plane [PLATFORM]
It is tempting to think of the PR as bureaucracy around the "real deployment".
In this model, the PR is part of the deployment control plane.
The pull request is where we can perform controls before desired state changes:
You no longer know whether production contains what was tested.
Promotion PRs
A typical flow becomes:
application merge
|
v
CI creates immutable image
|
v
PR: digest -> dev
|
v
dev verification
|
v
PR: same digest -> staging
|
v
staging verification
|
v
PR: same digest -> production
A platform can automate the creation of those PRs while still preserving review gates.
16. Separate Application Change from Environment Promotion [PLATFORM]
This is one of the most useful organisational consequences of the two-repository model.
Do not casually permit application projects to deploy into the argocd namespace itself.
An application able to modify Argo's own configuration or credentials can potentially turn application deployment permission into platform administration.
This is the same privilege-escalation reasoning we used in the Kubernetes RBAC chapter.
18. Argo's Kubernetes Permissions Still Matter [OPS] [PLATFORM]
GitOps does not abolish RBAC.
It moves the actor.
With imperative deployment:
developer / CI identity
|
v
Kubernetes RBAC
With Argo:
developer
|
v
Git permissions
|
v
Argo controller identity
|
v
Kubernetes RBAC
Ask which ServiceAccounts exist:
kubectl get serviceaccounts \
-n argocd
Inspect the controller identity:
kubectl get pod \
-n argocd \
-l app.kubernetes.io/name=argocd-application-controller \
-o jsonpath='{.items[0].spec.serviceAccountName}{"\n"}'
External Secrets Operator
Git stores a reference
secret value lives in Vault / cloud secret manager
SOPS
encrypted secret material lives in Git
authorised controller decrypts it
Sealed Secrets
encrypted object lives in Git
cluster controller decrypts it
The core design question is:
Can Git contain enough desired state to reference the secret without exposing the secret itself?
A production deployment might contain:
Git:
"web needs database credential named web-db"
secret manager:
actual credential bytes
controller:
materialises Kubernetes Secret
This is another reconciliation loop.
Modern platforms are often a composition of controllers, each owning one slice of desired state.
Part IV - Progressive Delivery
20. Why a Successful Git Sync Is Not the Same as a Safe Release [DEV] [OPS]
Argo CD answers:
Does the cluster match Git?
That does not necessarily answer:
Is this new version safe for 100% of users?
A normal Deployment may replace instances gradually, but it does not inherently evaluate business or reliability signals before continuing.
Progressive delivery adds another control loop:
new desired image
|
v
small amount of exposure
|
v
observe health
|
+---+---+
| |
good bad
| |
v v
more abort
traffic
We will use Argo Rollouts for this lab.
Argo CD and Argo Rollouts solve different problems:
Argo CD
Git -> Kubernetes desired state
Argo Rollouts
desired release -> controlled transition between versions
kubectl rollout status \
deployment/argo-rollouts \
-n argo-rollouts \
--timeout=300s
On macOS, install the kubectl plugin:
brew install argoproj/tap/kubectl-argo-rollouts
Check:
kubectl argo rollouts version
Find the new APIs:
kubectl api-resources \
| grep argoproj
You should now see resources including:
Rollout
AnalysisTemplate
AnalysisRun
Experiment
Again, a product feature has become Kubernetes API objects plus controllers.
22. Migrate a Deployment to Rollouts Without Controllers Fighting [DEV] [OPS]
A tempting migration is:
change kind: Deployment
to Rollout
For a throwaway manifest that can be enough.
For a running workload, it hides an important ownership problem.
If a Deployment and Rollout temporarily exist with overlapping selectors, both controllers can be active at the same time.
And if Argo CD keeps enforcing a Deployment replica count while Argo Rollouts is trying to scale that Deployment down during migration, two reconcilers can fight over the same field.
That gives us a better lab.
We will migrate using a Rollout workloadRef.
The existing Deployment remains the source of the Pod template.
The Rollout becomes responsible for progressive delivery.
Conceptually:
Git / Kustomize
|
v
Deployment Pod template
|
| workloadRef
v
Rollout
|
+-- manages progressive ReplicaSets
|
+-- scales old Deployment down
First decide who owns Deployment.spec.replicas
Our Argo Application currently has self-healing enabled.
Git says the original Deployment has:
replicas: 2
But during migration, Argo Rollouts needs to scale that Deployment down.
If both controllers insist on owning the same field, we can create a reconciliation tug-of-war:
Git / Argo CD: replicas must be 2
Argo Rollouts: replicas must be 0
The fix is not to disable reconciliation globally.
The fix is to define field ownership deliberately.
Tell Argo CD to ignore the replica field of this specific Deployment, and to respect that exclusion during sync:
python3 - <<'PY2'from pathlib import Pathp = Path("apps/web/base/kustomization.yaml")s = p.read_text()if " - rollout.yaml\n" not in s: s = s.replace("resources:\n", "resources:\n - rollout.yaml\n")p.write_text(s)PY2
Do not remove the Deployment.
The Rollout is using its Pod template through workloadRef.
Render the desired state:
kubectl kustomize environments/dev/web
You should see both:
Deployment/web
Rollout/web
That is intentional during this migration pattern.
Commit and push:
git add .
git commit -m "migrate web delivery to Argo Rollouts"
git push
Watch the two controllers:
kubectl get deployment,rollout,rs,pods \
-n web-dev \
-w
Inspect the Rollout:
kubectl argo rollouts get rollout web \
-n web-dev
The Rollout should establish its own stable ReplicaSet while progressively scaling down the old Deployment-managed workload.
Inspect the original Deployment:
kubectl get deployment web \
-n web-dev
Its replica count can now be changed by the Rollouts controller without Argo CD immediately restoring the Git value.
Where does the image live now?
Because we used workloadRef, the Pod template still lives on the Deployment:
kubectl get deployment web \
-n web-dev \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
Our existing Kustomize image override therefore continues to work exactly where it did before.
When a promotion PR changes the Deployment Pod template image, Argo Rollouts observes the referenced workload change and performs the canary transition.
This gives us a clean division of ownership:
Git / Argo CD
owns Deployment Pod template
owns Rollout policy
Argo Rollouts
owns progressive ReplicaSets
owns migration scaling
Kubernetes
owns Pod execution and status
The migration itself has taught us something broader than Argo Rollouts:
Multiple controllers can cooperate safely only when we understand which fields and resources each controller owns.
28. Canary Replica Ratios vs Real Traffic Shaping [OPS] [PLATFORM]
Our basic canary uses replica counts.
With five replicas:
1 canary + 4 stable ~= 20%
That is useful but crude.
It cannot express concepts such as:
1% traffic to canary
only users with header X
mirror requests without using canary responses
90/10 traffic while keeping equal replica counts
For that, Argo Rollouts integrates with traffic-management systems.
This is where our Gateway API / Cilium work connects directly.
Conceptually:
Rollout desired weight
|
v
Argo Rollouts
|
v
Gateway API route weights
|
v
Cilium
|
v
real network traffic
Modern Argo Rollouts can integrate with Gateway API through its traffic-router plugin system.
That allows the progressive-delivery controller to update Gateway API routing state rather than merely scaling stable/canary replica counts.
For learners following the Cilium Gateway API companion, this is the natural next exercise after mastering the basic Rollout.
The important architectural connection is:
Git
|
v
Argo CD
|
v
Rollout
|
v
Argo Rollouts
|
+--> ReplicaSets
|
+--> Gateway API route weight
|
v
Cilium
Each controller owns a different concern.
Another controller-ownership trap
A traffic router introduces the same field-ownership issue we saw during Deployment-to-Rollout migration.
If Argo Rollouts dynamically changes route weights while Argo CD insists that the weights must always equal the static values committed in Git, the controllers can report permanent drift or overwrite one another.
The usual GitOps pattern is to keep the route object in Git while explicitly ignoring the fields that the progressive-delivery controller is expected to mutate.
It is defining ownership precisely enough for multiple reconcilers to cooperate.
29. Blue-Green: Build the Replacement Before You Switch Traffic [DEV] [OPS]
Canary is not the only progressive-delivery strategy.
A canary asks:
Can we expose a small amount of real traffic to the new version and increase it gradually?
Blue-green asks a different question:
Can we build the entire replacement, test it separately, and switch production traffic only when we are happy?
The mental model is:
+--------------------+
| stable ReplicaSet |
| v3 |
+---------+----------+
^
|
active Service
|
production
+--------------------+
| preview ReplicaSet |
| v4 |
+---------+----------+
^
|
preview Service
|
tests / humans
Both versions exist.
But only one receives production traffic.
When we promote:
before
------
web Service ----------> v3
web-preview Service --> v4
promotion
---------
web Service ----------> v4
web-preview Service --> v4
v3 remains briefly
available for rollback
This gives blue-green a very useful property:
new version is deployed
!=
new version is serving production traffic
That distinction is worth experiencing directly.
Canary vs blue-green
The two strategies solve slightly different release problems.
RollingUpdate
replace instances gradually
simplest operational model
Canary
expose some real production traffic
measure behaviour
increase exposure gradually
Blue-green
build a complete replacement
validate it away from production
switch traffic when ready
Blue-green is often easier to reason about because there is a hard traffic boundary between the active and preview versions.
It does have a cost.
For at least part of the release we may run two versions simultaneously:
active capacity
+
preview capacity
For an expensive workload that may matter.
Argo Rollouts lets us reduce preview capacity with previewReplicaCount, but the new version must be scaled to the full desired replica count before it becomes active.
When would you choose each strategy?
Blue-green is attractive when:
we can validate a version before production traffic
cutover should happen quickly
old and new versions cannot safely share live traffic
rollback speed matters
extra temporary capacity is acceptable
Canary is attractive when:
real-user behaviour is part of validation
we have useful production metrics
we want gradual blast-radius expansion
our application tolerates multiple versions serving simultaneously
we have a traffic router for precise percentages
Neither is universally better.
A platform can support both.
The release should choose the strategy that matches the application's risk model.
30. Convert Our Rollout from Canary to Blue-Green [DEV] [OPS] [PLATFORM]
We already have:
Git
|
v
Argo CD
|
v
Rollout/web
|
v
ReplicaSets
We are going to keep that delivery chain.
We will change only the progressive-delivery strategy.
The blue-green controller needs two Services:
web
active production traffic
web-preview
pre-production validation traffic
Argo Rollouts will dynamically add a ReplicaSet hash to those Service selectors so that each Service points at exactly the intended version.
That creates an important ownership question.
Who owns the Service selector?
Our Service currently lives in Git:
spec:
selector:
app: web
For blue-green, Argo Rollouts will turn the live selector into something conceptually like:
gh pr create \
--title "Use blue-green delivery for web" \
--body "Add a preview Service and move web from canary steps to an Argo Rollouts blue-green strategy."
Review it.
Then merge it:
gh pr merge \
--merge \
--delete-branch
Return to main:
git checkout main
git pull
Watch Argo CD reconcile:
kubectl get application web-dev \
-n argocd \
-w
In another terminal inspect the Rollout:
kubectl argo rollouts get rollout web \
-n web-dev
And the Services:
kubectl get service web web-preview \
-n web-dev
At steady state, both Services may initially point at the current stable ReplicaSet.
The interesting behaviour appears on the next release.
31. Ship a Blue-Green Release: Preview First, Production Later [DEV] [OPS]
Now we will ship a version whose lifecycle is deliberately visible.
The application change still begins in the application repository.
That is important.
We do not edit the GitOps repository by hand to invent a release.
The application CI builds an immutable artifact and proposes its digest for promotion.
kubectl get service web web-preview \
-n web-dev \
-o custom-columns='SERVICE:.metadata.name,HASH:.spec.selector.rollouts-pod-template-hash,APP:.spec.selector.app'
You should see different hashes while the new version is awaiting promotion:
SERVICE HASH APP
web <old-hash> web
web-preview <new-hash> web
This is the blue-green boundary made concrete.
Argo Rollouts has changed routing without changing either Service's identity.
image built successfully
|
image was pulled successfully
|
Pods became Ready
|
preview endpoint works
|
automated gate passed
|
production still serves old version
This is a much stronger decision point than:
CI job went green -> deploy everything
Promote the preview
When satisfied, promote it:
kubectl argo rollouts promote web \
-n web-dev
Watch:
kubectl argo rollouts get rollout web \
-n web-dev \
--watch
During promotion, Rollouts will ensure the new ReplicaSet reaches the full desired size before switching production traffic.
Inspect the Services again:
kubectl get service web web-preview \
-n web-dev \
-o custom-columns='SERVICE:.metadata.name,HASH:.spec.selector.rollouts-pod-template-hash'
Both should now point at the new ReplicaSet hash.
Re-run the production request:
curl -s http://127.0.0.1:8082 \
| grep version
Now production should say:
version: blue-green-v4
The important operation was not replacing the Service.
The Service stayed stable:
web.web-dev.svc.cluster.local
Argo Rollouts changed which ReplicaSet that stable identity selected.
Observe the old version before it disappears
Immediately after promotion:
kubectl get rs \
-n web-dev \
-l app=web
The old active ReplicaSet should remain scaled for roughly our configured delay:
scaleDownDelaySeconds: 60
Watch it:
kubectl get rs \
-n web-dev \
-l app=web \
-w
Eventually the old ReplicaSet scales down.
This delay is intentionally different from keeping old ReplicaSet metadata around.
Kubernetes may retain the old ReplicaSet object for revision history even after its replica count reaches zero.
The distinction is:
revision retained
!=
old application still consuming full compute
32. Break Blue-Green Safely: Fail Before Production Sees It [DEV] [OPS]
while allowing workloads to choose an appropriate rollout policy.
Rollback means different things at different moments
Before blue-green promotion:
new version rejected
production never moved
There is no traffic rollback to perform.
We only need to repair desired state in Git.
Immediately after blue-green promotion:
new version active
old ReplicaSet may still be scaled up
Operational rollback can be extremely fast.
After old capacity has been scaled down:
Git revert
|
v
old digest becomes desired
|
v
Rollouts restores the previous revision
With a rollback window, recent revisions can be fast-tracked.
For canary:
abort
|
v
stable ReplicaSet regains exposure
but Git must still be corrected if it describes the rejected image.
The recurring lesson is:
Operational safety actions and source-of-truth changes are related, but they are not the same operation.
The platform-engineering view
At this point our delivery system contains several reconcilers:
Git
|
v
Argo CD
|
+------------------------------+
| |
v v
Rollout policy Services
| ^
v |
Argo Rollouts ------------------+
|
v
ReplicaSets
|
v
Pods
For canary with a traffic router we add another loop:
Argo Rollouts
|
v
Gateway API route weights
|
v
Cilium
For blue-green, Rollouts controls active/preview Service selectors instead.
This is why platform engineering is fundamentally about more than installing controllers.
You need to know:
which controller owns which resource
which controller owns which field
what the source of truth is
which mutations are temporary operational state
which mutations must go back through Git
That is the difference between a platform composed of controllers and a pile of controllers fighting each other.
Part V - Platform Engineering
34. CI Should Have Registry and Git Permissions - Not Cluster Admin [PLATFORM]
Our final developer pipeline has approximately these permissions:
Another platform may use one repository per environment or business unit.
The useful questions are:
Who may approve this path?
What blast radius does changing this directory have?
Can one repository compromise every cluster?
How is an artifact promoted?
How do we audit the change?
Repository topology is not merely aesthetic.
It is part of your security and operating model.
37. ApplicationSet: When One Application Becomes Hundreds [OPS] [PLATFORM]
Creating one Argo Application by hand is fine.
Creating 500 nearly-identical Application objects manually is not.
Argo CD provides ApplicationSet to generate Applications from data.
Conceptually:
list of clusters / directories / tenants
|
v
ApplicationSet
|
v
many Applications
|
v
many reconciliations
Do not use ApplicationSet merely because it exists.
Use it when the repetition itself is data-driven.
Typical platform examples include:
one app per cluster
one app per tenant
one app per region
one environment directory per service
This is the GitOps version of moving from hand-created Pods to controllers.
38. App of Apps, ApplicationSet and Platform Bootstrapping [DEEP DIVE]
Eventually you may want GitOps to manage the platform components that enable GitOps.
That sounds circular because it is.
A bootstrap process often looks like:
create cluster
|
v
install minimal Argo CD
|
v
point Argo at platform bootstrap repository
|
v
Argo installs / manages
- policies
- ingress / Gateway API
- observability
- operators
- team Applications
Once bootstrapped, most ongoing platform state can flow through Git.
You still need an answer for:
Who creates the cluster?
Who installs the first Argo instance?
Who provides its Git credentials?
Who upgrades Argo itself?
GitOps moves the bootstrap boundary.
It does not eliminate it.
This is where tools such as Cluster API, hosted control planes, Terraform/OpenTofu, MAAS/NiCO, vCluster, or Kamaji may sit beneath the application platform.
A useful full-stack model is:
infrastructure desired state
|
v
clusters / nodes / networks
|
v
GitOps bootstrap
|
v
platform services
|
v
application GitOps
|
v
workloads
39. Supply-Chain Policy: Build Evidence Must Reach Deployment Policy [PLATFORM]
We built and signed an artifact.
That is only half the story.
A mature platform connects build evidence to deployment policy.
For example:
CI produces:
image digest
signature
provenance
SBOM
vulnerability result
promotion policy requires:
trusted builder identity
no forbidden vulnerabilities
approved source repository
immutable digest
admission policy verifies:
image is allowed to start
This closes a gap that Git review alone cannot.
A perfectly reviewed Git commit could still point to:
unsigned malicious image
if nothing verifies the artifact.
Likewise, a perfectly signed artifact can still be accidentally promoted to the wrong environment if Git permissions are too broad.
The controls reinforce one another.
40. Harbor as a Platform Registry [PLATFORM]
Harbor becomes particularly useful when an organisation wants the registry itself to enforce platform rules.
A production pattern might be:
Harbor project: payments
CI robot:
push
read
runtime identity:
pull only
platform admin:
configure retention
immutability
scanning
trust policy
Useful policy decisions:
Make release tags immutable
Allow:
sha-abc123
v1.4.2
but prevent those tags from being overwritten once published.
This does not replace digest pinning.
It makes human-readable references less surprising.
Generate SBOMs
Harbor can integrate SBOM generation with its scanner.
The SBOM gives visibility into the packages inside an artifact.
Enforce signatures
Harbor can associate Cosign/Notation signatures with OCI artifacts and can enforce content trust on a project.
The desired path becomes:
unsigned image
X
cannot be consumed
signed trusted image
|
v
runtime may pull it
Use robot accounts
Automations should not log in using a platform administrator's personal username and password.
Create separate identities for:
CI publisher
replication
scanner
runtime pull
with the smallest permissions each needs.
This mirrors Kubernetes ServiceAccounts and RBAC.
Machine identity should be explicit.
41. Observability for the Delivery System [OPS] [PLATFORM]
The deployment system is production software too.
Monitor it.
Useful questions include:
Is Argo able to read Git?
Are Applications OutOfSync?
Are Applications Degraded?
How long do syncs take?
Are Rollouts paused or aborted?
Are AnalysisRuns failing?
Can nodes pull images?
Did registry scanning fail?
Are promotion PRs stuck?
A useful event chain for one release is:
source commit
|
v
CI run
|
v
image digest
|
v
promotion PR
|
v
Git merge commit
|
v
Argo sync revision
|
v
Rollout revision
|
v
ReplicaSet
|
v
Pods
Preserving those identifiers makes incident response dramatically easier.
Good platform metadata might include:
source git SHA
image digest
GitOps commit SHA
service name
environment
owner
release timestamp
42. Failure Drills [OPS]
A delivery system is not understood until you have watched it fail.
It is one extremely useful reconciliation boundary inside the platform.
45. From Kubernetes Learner to Platform Engineer [DEV] [OPS] [PLATFORM]
The progression through these cookbooks now looks approximately like this:
LEVEL 1 - Kubernetes user
Pod
Deployment
Service
ConfigMap
Secret
PVC
probes
|
v
LEVEL 2 - Kubernetes developer
rollouts
resources
scheduling
RBAC
Kustomize
Helm
Gateway API
|
v
LEVEL 3 - Kubernetes internals
spec / status
controllers
CRDs
ownerReferences
finalizers
custom Go reconciler
|
v
LEVEL 4 - delivery engineer
OCI images
immutable digests
CI
GitOps
Argo CD
promotion PRs
progressive delivery
rollback
|
v
LEVEL 5 - platform engineer
AppProjects
multi-tenancy
hosted control planes
Gateway API / Cilium
registry policy
artifact trust
secret systems
fleet management
ApplicationSets
observability
policy
|
v
LEVEL 6 - platform designer
Who owns desired state?
Where are trust boundaries?
Which controller owns each resource?
How does software move between environments?
How is privilege constrained?
How do we prove what is running?
How does the platform fail safely?
The tools will change.
Those questions age much more slowly.
46. Production Checklist [PLATFORM]
Before calling a GitOps delivery platform production-ready, be able to answer all of these.
Artifact
[ ] Is every release associated with an immutable digest?
[ ] Is the same digest promoted through environments?
[ ] Are release tags immutable where practical?
[ ] Are artifacts vulnerability scanned?
[ ] Is an SBOM available?
[ ] Are artifacts signed / attested?
[ ] Is signature/provenance verification enforced somewhere?
CI
[ ] Does CI avoid general Kubernetes credentials?
[ ] Are registry credentials scoped?
[ ] Is GitOps write access scoped to the required repository/path?
[ ] Is automation represented by a machine identity, not a person's token?
[ ] Can the build identity be audited?
Git
[ ] Are production branches protected?
[ ] Are CODEOWNERS / reviewers appropriate?
[ ] Is direct push disabled where required?
[ ] Can every production digest be traced to a reviewed change?
[ ] Is rollback performed through source-of-truth changes?
Argo CD
[ ] Are AppProjects explicit rather than relying on default?
[ ] Are source repositories restricted?
[ ] Are destination namespaces/clusters restricted?
[ ] Is Argo's own Kubernetes RBAC understood?
[ ] Is access to the argocd namespace tightly controlled?
[ ] Are prune and self-heal policies deliberate?
[ ] Are Argo upgrades and backups planned?
Progressive delivery
[ ] Is rollout strategy appropriate for the service?
[ ] Are canary signals meaningful?
[ ] Can a release be aborted quickly?
[ ] Is there a documented Git rollback path?
[ ] Are automated analyses observable and auditable?
[ ] Is traffic shaping coarse replica ratio or real routed traffic by design?
Secrets
[ ] Are plaintext production secrets kept out of ordinary Git history?
[ ] Is secret rotation supported?
[ ] Can workloads access only the secrets they need?
Operations
[ ] Can you identify source commit -> image digest -> GitOps commit -> running Pods?
[ ] Are reconciliation failures alerted?
[ ] Are OutOfSync and Degraded applications visible?
[ ] Are stuck/aborted Rollouts visible?
[ ] Have failure drills actually been run?
If several answers are "we assume so", keep building.
47. Cleanup [DEV]
Remove the Argo-managed application first:
kubectl delete application web-dev \
-n argocd
Depending on finalizer/prune configuration, inspect whether its managed resources remain:
kubectl get all \
-n web-dev
Delete the lab namespace if required:
kubectl delete namespace web-dev
Remove Argo Rollouts:
kubectl delete namespace argo-rollouts
Remove Argo CD:
kubectl delete namespace argocd
The CRDs are cluster-scoped and may remain after deleting namespaces.
Inspect:
kubectl get crd \
| grep argoproj
For a disposable kind lab, the cleanest full reset is often simply deleting and recreating the cluster.
Do not blindly use that advice on a cluster containing anything you care about.
Appendix A - The Whole Reconciliation Stack
A useful final mental model:
SOURCE CODE
|
| developer changes application behaviour
v
APPLICATION GIT
|
| CI reconciles source into an artifact
v
OCI IMAGE DIGEST
|
| promotion automation proposes desired deployment
v
GITOPS GIT
|
| Argo CD reconciles Git into Kubernetes resources
v
ROLLOUT / DEPLOYMENT SPEC
|
| workload controller reconciles replicas
v
REPLICASETS / PODS
|
| kubelet reconciles Pod specs on nodes
v
CONTAINERS
Bad CI/CD often gives one automation system enormous authority:
source
|
v
CI
|
| build
| mutate production
| own cluster credentials
v
cluster
A better delivery platform separates concerns:
source
|
v
CI
|
| produces immutable evidence
v
artifact
|
v
promotion PR
|
| changes reviewed desired state
v
Git
|
v
Argo CD
|
| reconciles cluster intent
v
Argo Rollouts
|
| controls exposure
v
runtime
At each boundary we can ask:
What is the desired state?
Who may change it?
Which controller reconciles it?
What identity does that controller use?
How do we observe failure?
How do we return to a known-good state?
If you can answer those questions, you are no longer merely deploying to Kubernetes.
A runnable Kubernetes cookbook for application developers and CKAD candidates.
The goal is not to memorise YAML.
The goal is to build a mental model of Kubernetes, use the API deliberately, observe what the control plane did, break things on purpose, and work out why they broke.
Target: Kubernetes 1.35 and the current CKAD curriculum.
Sections are marked:
[CKAD] - directly relevant to CKAD
[DEV] - practical application developer knowledge
[DEEP DIVE] - controllers, operators and platform engineering
This cookbook is intentionally cumulative. We will keep reusing the same resources so that later concepts explain earlier behaviour rather than appearing as unrelated YAML fragments.
The recurring teaching loop is:
problem
|
v
mental model
|
v
small experiment
|
v
observe Kubernetes
|
v
change one thing
|
v
observe the consequence
|
v
break an assumption
|
v
explain why
When a section does not need every step, we will not force it into a template.
Part I - Learn the Control Loop
0. Lab Setup [CKAD]
Use one namespace for the whole cookbook so that commands stay short and cleanup is easy.
kubectl create namespace cookbook
Make it the default namespace for the current context:
Despite the name, all does not mean every Kubernetes API resource.
For actual API discovery:
kubectl api-resources
A note about the lab
Most examples work on any ordinary Kubernetes cluster.
A few exercises depend on optional cluster capabilities:
PersistentVolumeClaims need a storage provisioner or an existing PersistentVolume.
NetworkPolicy enforcement needs a CNI that implements NetworkPolicy.
Ingress routing needs an Ingress controller.
CRD exercises need permission to create cluster-scoped CustomResourceDefinitions.
When one of those assumptions matters, the recipe will say so.
1. The Kubernetes Control Loop [CKAD] [DEV]
Before learning all of Kubernetes' resource types, learn the one idea that connects nearly all of them.
Kubernetes is not primarily a system for running commands on servers.
It is a system for declaring what you want and continuously trying to make reality match it.
Conceptually:
you declare what you want
|
v
Kubernetes API
|
v
controller
/ | \
observe compare act
\ | /
reality
|
+-------- repeat
This repeating process is called reconciliation.
Let's make it happen rather than just defining it.
Start with one application
Create an nginx Deployment with one replica:
kubectl create deployment api \
--image=nginx:1.27-alpine \
--replicas=1
Look at the Deployment:
kubectl get deployment api
And the Pod it caused Kubernetes to create:
kubectl get pods -l app=api
You asked Kubernetes for one copy of nginx.
A Pod now exists.
The important part is how Kubernetes represents that request.
spec is what you asked for
Ask the API how many replicas the Deployment should have:
kubectl get deployment api \
-o jsonpath='{.spec.replicas}{"\n"}'
Expected:
1
That value lives under something like:
spec:
replicas: 1
spec describes desired state.
It is your instruction to Kubernetes:
I want one replica of this application.
Now ask Kubernetes how many replicas are actually ready:
kubectl get deployment api \
-o jsonpath='{.status.readyReplicas}{"\n"}'
Once nginx is ready, you should see:
1
That value comes from status:
status:
readyReplicas: 1
A useful first approximation is:
spec -> what you want
status -> what Kubernetes currently sees
You normally change spec.
Kubernetes and its controllers populate status.
Change what you want
Suppose one replica is no longer enough.
Tell Kubernetes you want three:
kubectl scale deployment api --replicas=3
This command did not directly start two containers.
It changed the desired state stored in the Kubernetes API.
Check:
kubectl get deployment api \
-o jsonpath='{.spec.replicas}{"\n"}'
Now:
3
Watch the Pods:
kubectl get pods -l app=api -w
Two more Pods should appear.
Press Ctrl-C once all three are running.
Now compare desired state with observed state:
kubectl get deployment api \
-o jsonpath='{.spec.replicas}{" desired, "}{.status.readyReplicas}{" ready\n"}'
Eventually:
3 desired, 3 ready
For a short time the system may have looked like:
spec.replicas: 3
status.readyReplicas: 1
Desired state and observed state disagreed.
Controllers noticed the difference and acted.
Eventually:
spec.replicas: 3
status.readyReplicas: 3
Reality caught up with intent.
That is reconciliation.
Change reality instead
Now do the opposite.
Instead of changing what we want, interfere with what actually exists.
Delete one Pod:
kubectl delete pod \
$(kubectl get pods \ -l app=api \ -o jsonpath='{.items[0].metadata.name}')
Immediately watch:
kubectl get pods -l app=api -w
The Pod you deleted disappears.
Then another Pod appears.
This distinction matters:
The old Pod did not restart. Kubernetes created a replacement.
Why?
Deleting a Pod changed actual state, but it did not change the Deployment's desired state.
The Deployment still says, in effect:
spec:
replicas: 3
Kubernetes therefore sees:
desired replicas: 3
actual replicas: 2
and reconciles the difference:
desired: 3
|
v
observe: 2
|
v
difference: -1
|
v
create another Pod
|
v
actual: 3
This is why Kubernetes applications should normally be designed around replaceable workloads rather than individual machines or individual containers.
There is more than one controller involved
We have simplified the story slightly.
A Deployment does not directly create Pods.
The relationship is approximately:
Deployment
|
v
ReplicaSet
|
+-- Pod
+-- Pod
+-- Pod
The Deployment controller manages ReplicaSets.
The ReplicaSet controller makes sure the correct number of matching Pods exists.
See the hierarchy:
kubectl get deployment api
kubectl get replicasets
kubectl get pods -l app=api
Inspect Pod ownership:
kubectl get pods \
-l app=api \
-o custom-columns='POD:.metadata.name,OWNER:.metadata.ownerReferences[0].name'
Do not worry about memorising ReplicaSets yet.
For now, remember the pattern:
declare
|
v
observe
|
v
compare
|
v
reconcile
|
+------ repeat
That pattern is Kubernetes.
How do we know the controller saw our change?
Many controller-managed objects expose another useful pair of fields.
kubectl get deployment api \
-o jsonpath='{.metadata.generation}{" desired generation, "}{.status.observedGeneration}{" observed generation\n"}'
metadata.generation changes when the desired configuration changes.
status.observedGeneration tells us which generation the controller has observed.
You do not need to memorise those fields yet.
The useful idea is that Kubernetes APIs can expose both:
what was requested
and:
how far the controller has progressed towards it
That becomes very useful when debugging.
Useful detour: talk to the application
We know nginx is running.
Prove it with a local port-forward:
kubectl port-forward deployment/api 8080:80
In another terminal:
curl http://127.0.0.1:8080
You should receive nginx's HTML response.
kubectl port-forward is extremely useful while developing and debugging because it creates a temporary tunnel from your machine to a workload in the cluster.
It is not application networking.
your laptop
|
localhost:8080
|
kubectl
|
API server
|
v
Pod
When the command stops, the tunnel stops.
Later we will create a Service, which solves a different problem: giving replaceable Pods a stable network identity inside Kubernetes.
What to keep from this chapter
Do not memorise the commands yet.
Remember the model:
spec
|
| what should exist
v
controller
|
| continuously reconciles
v
actual state
|
| reported back as
v
status
Changing spec changes your intent.
Changing reality without changing spec causes Kubernetes to try to repair reality.
Almost everything else in this cookbook builds on that idea.
2. API Objects and Changing Desired State [CKAD] [DEV]
apiVersion -> which API schema?
kind -> what type of object?
metadata -> which object?
spec -> what do I want?
status -> what does Kubernetes observe?
status is normally omitted from manifests you write.
Controllers populate it after the object exists.
Ask Kubernetes for the object we already created
kubectl get deployment api -o yaml
There is a lot of output.
Do not try to read it all.
Find these areas:
metadata:
spec:
status:
The object in the API is richer than the small amount of configuration we initially supplied because Kubernetes defaults fields and controllers add status.
Generate YAML instead of writing it from memory
Generate a Deployment manifest without creating anything:
This is usually faster and safer than guessing YAML.
Query individual fields
Deployment name:
kubectl get deployment api \
-o jsonpath='{.metadata.name}{"\n"}'
Desired replicas:
kubectl get deployment api \
-o jsonpath='{.spec.replicas}{"\n"}'
Ready replicas:
kubectl get deployment api \
-o jsonpath='{.status.readyReplicas}{"\n"}'
Container image:
kubectl get deployment api \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
Pod IPs:
kubectl get pods \
-l app=api \
-o custom-columns='NAME:.metadata.name,IP:.status.podIP'
This is querying the Kubernetes API.
It is not executing commands inside a container.
Learn to switch output formats
Human summary:
kubectl get pods
More columns:
kubectl get pods -o wide
Full object:
kubectl get pod <pod-name> -o yaml
JSON:
kubectl get pod <pod-name> -o json
Selected fields:
kubectl get pods \
-o custom-columns='NAME:.metadata.name,PHASE:.status.phase,NODE:.spec.nodeName'
A strong Kubernetes workflow is often:
get
|
v
describe
|
v
query exact fields
|
v
logs/events when needed
Part II - Pods and Workload Controllers
4. Pods: The Unit Kubernetes Schedules [CKAD]
A Pod is the smallest workload Kubernetes schedules onto a node.
A Pod contains one or more containers that share:
an IP address
a network namespace
localhost
optional volumes
scheduling fate
The most important practical property is this:
Pods are replaceable.
Do not treat a Pod name or Pod IP as a durable application endpoint.
The container image is an input to Kubernetes
Kubernetes schedules and runs container images; it does not normally build your application image for you.
A minimal image definition might look like:
FROM nginx:1.27-alpine
COPY index.html /usr/share/nginx/html/index.html
Build it with an OCI-compatible image tool such as Docker or Podman:
docker build -t example/web:v1 .
or:
podman build -t example/web:v1 .
The cluster then needs to be able to pull or otherwise access that image. In a real workflow that usually means pushing it to a registry the cluster can reach.
Dockerfile / Containerfile
|
v
build an OCI image
|
v
registry / cluster image store
|
v
Pod spec references image
|
v
kubelet asks container runtime to run it
Most exercises use public images so we can concentrate on Kubernetes itself. In Chapter 9 we will deliberately build our own image and discover what “the cluster needs to be able to access that image” actually means.
Create a standalone Pod
The Deployment from Chapter 1 is controller-managed.
Create a separate Pod so we can inspect Pod behaviour without a controller replacing it:
kubectl run pod-lab \
--image=nginx:1.27-alpine \
--restart=Never \
--labels=app=pod-lab
Observe:
kubectl get pod pod-lab
kubectl get pod pod-lab -o wide
describe connects configuration to runtime events
kubectl describe pod pod-lab
Pay attention to:
Node
IP
Containers
Conditions
Events
get -o yaml gives the complete API object.
describe gives a human-oriented operational summary.
Both are useful for different reasons.
Logs are container output
kubectl logs pod-lab
NGINX may have little to show until it receives traffic.
Port-forward temporarily:
kubectl port-forward pod/pod-lab 8081:80
In another terminal:
curl http://127.0.0.1:8081
Now check logs again:
kubectl logs pod-lab
exec runs a process inside the existing container
kubectl exec pod-lab -- nginx -v
Try:
kubectl exec pod-lab -- ps
For an interactive shell:
kubectl exec -it pod-lab -- sh
Exit with:
exit
Container command and args
Kubernetes can override an image's configured entrypoint and arguments.
Pending Pod
|
v
scheduler could not bind it to a node
|
v
inspect scheduling events
Cleanup:
kubectl delete pod limited unschedulable
rm -f limited.yaml unschedulable.yaml
ResourceQuota and LimitRange
Requests and limits are workload configuration.
Namespaces can also have policy around resource consumption.
Inspect any existing quota:
kubectl get resourcequota
kubectl get limitrange
A ResourceQuota can cap aggregate namespace consumption.
A LimitRange can constrain or default per-container or per-Pod values.
You do not need to confuse these with requests and limits:
request / limit -> what this workload asks for
LimitRange -> policy/defaults around individual workloads
ResourceQuota -> aggregate namespace budget
Scheduling controls placement, not API permission
Later we will meet node selectors, affinity, taints and tolerations in broader platform contexts.
Keep one distinction in mind now:
scheduler controls
-> where a Pod is eligible to run
RBAC
-> what API operations an identity may perform
A NoSchedule taint on a control-plane node can keep ordinary workloads away from that node. It does not decide who may edit Node objects through the API, and it does not control who may SSH into the machine.
We will bring those boundaries together in Chapter 23.
7. Deployments and ReplicaSets [CKAD]
We already used a Deployment before formally studying it because it gave us the smallest useful reconciliation experiment.
Now we can unpack the machinery.
A Deployment is appropriate for replaceable application replicas where individual Pod identity does not matter.
The ownership chain is:
Deployment
|
v
ReplicaSet
|
+-- Pod
+-- Pod
+-- Pod
Inspect the Deployment
kubectl get deployment api
ReplicaSet:
kubectl get replicasets -l app=api
Pods:
kubectl get pods -l app=api
Follow ownership through the API
Pod -> ReplicaSet:
kubectl get pods \
-l app=api \
-o custom-columns='POD:.metadata.name,OWNER:.metadata.ownerReferences[0].name'
ReplicaSet -> Deployment:
kubectl get rs \
-l app=api \
-o custom-columns='RS:.metadata.name,OWNER:.metadata.ownerReferences[0].name'
The ownership chain is not hidden state.
It is represented in the API.
Scaling is reconciliation, not cloning
Scale to five:
kubectl scale deployment api --replicas=5
Watch:
kubectl get pods -l app=api -w
Then inspect:
kubectl get deployment api \
-o jsonpath='{.spec.replicas}{" desired, "}{.status.readyReplicas}{" ready\n"}'
Scale back:
kubectl scale deployment api --replicas=3
Scaling does not create a new Deployment revision because it does not change the Pod template.
The Pod template is the important boundary
Inspect:
kubectl get deployment api \
-o jsonpath='{.spec.template}{"\n"}'
A Deployment says, approximately:
I want N Pods that look like this template.
Changing replicas changes how many Pods.
Changing .spec.template changes what the Pods should be, which leads directly to rolling updates.
8. Rolling Updates, Revisions and Rollback [CKAD]
Suppose we want a new application version.
Replacing every Pod at once would create unnecessary downtime.
A Deployment can instead progressively move from one ReplicaSet to another.
Check the current image
kubectl get deployment api \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
Change the Pod template
kubectl set image deployment/api \
nginx=nginx:1.28-alpine
Watch rollout progress:
kubectl rollout status deployment/api
Inspect ReplicaSets:
kubectl get rs -l app=api
You should now see an older ReplicaSet and a newer one.
Why did Kubernetes create a new ReplicaSet this time but not when we scaled?
Because the Pod template changed.
change .spec.replicas
|
v
same Pod template
|
v
same Deployment revision
change .spec.template
|
v
new desired Pod definition
|
v
new ReplicaSet / revision
Rollout history
kubectl rollout history deployment/api
Roll back
kubectl rollout undo deployment/api
Watch:
kubectl rollout status deployment/api
Check the image again:
kubectl get deployment api \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
RollingUpdate strategy
Inspect the strategy:
kubectl get deployment api \
-o jsonpath='{.spec.strategy}{"\n"}'
kubectl set image deployment/api \
nginx=example/web:v2
Watch the rollout:
kubectl rollout status deployment/api
The local development loop is now explicit:
edit source
|
v
build image
|
v
make image available to nodes
|
v
change Deployment spec
|
v
controllers reconcile
|
v
new Pods become ready
For kind:
docker build -t example/web:v2 .
kind load docker-image example/web:v2 --name ckad
kubectl set image deployment/api nginx=example/web:v2
kubectl rollout status deployment/api
On a real multi-node or remote cluster, the middle step is normally a registry rather than kind load.
Why use a new image tag?
During local development it is tempting to rebuild the same tag repeatedly:
example/web:v1
with different contents.
Prefer immutable or at least unique version tags while learning this workflow:
example/web:v1
example/web:v2
example/web:v3
Then the desired image reference tells you which build Kubernetes was asked to run.
Production systems often go further and deploy immutable image digests.
The important principle is the same:
know exactly which image desired state refers to
What the two failures taught us
Both experiments produced an image pull failure.
But the causes were different.
Failure 1
requested image
|
v
image does not exist
|
v
pull fails
The fix was to correct desired state:
kubectl rollout undo
Failure 2
requested image
|
v
image exists on developer machine
|
v
image absent from Kubernetes node
|
v
registry cannot provide it
|
v
pull fails
The desired state was valid.
The fix was to make the requested artifact available to the node:
kind load docker-image
That distinction is much more useful than memorising:
ImagePullBackOff = image problem
A better mental model is:
status tells you where the system is stuck
|
v
Events and object relationships tell you why
And underneath both examples is still the same Kubernetes control loop:
desired Pod template
|
v
controller creates replacement workload
|
v
node attempts to realise it
|
+---- cannot obtain image ----> status + Events expose failure
|
v
condition fixed
|
v
reconciliation continues
Part III - Stable Networking for Replaceable Pods
10. Services: Stable Identity for Replaceable Pods [CKAD]
Our Deployment Pods are intentionally disposable.
That creates a networking problem.
A Pod can disappear and its replacement can have a different IP address.
Clients therefore need something more stable than a Pod IP.
A Service provides a stable virtual endpoint over a selected set of backends.
Conceptually:
Client
|
v
Service
|
v
EndpointSlice
|
+-- Pod
+-- Pod
+-- Pod
Expose the Deployment
kubectl expose deployment api \
--name=api \
--port=80 \
--target-port=80
Inspect:
kubectl get service api
kubectl describe service api
The important distinction is:
port -> port clients use on the Service
targetPort -> port the application receives traffic on
How does the Service find Pods?
Query its selector:
kubectl get service api \
-o jsonpath='{.spec.selector}{"\n"}'
You should see something equivalent to:
map[app:api]
Check matching Pods:
kubectl get pods -l app=api --show-labels
The Service does not contain a permanent list of those Pod IPs.
It declares a selector.
Another controller continuously maintains the backend endpoint data.
There is our reconciliation model again.
11. EndpointSlices and DNS [CKAD] [DEV]
A common simplified drawing is:
Service -> Pods
Useful, but incomplete.
For Services with selectors, Kubernetes normally maintains EndpointSlices containing matching backend endpoints.
Inspect EndpointSlices
kubectl get endpointslices
Filter for our Service:
kubectl get endpointslices \
-l kubernetes.io/service-name=api
Show addresses:
kubectl get endpointslices \
-l kubernetes.io/service-name=api \
-o jsonpath='{.items[*].endpoints[*].addresses[*]}{"\n"}'
Compare with Pod IPs:
kubectl get pods \
-l app=api \
-o custom-columns='NAME:.metadata.name,IP:.status.podIP'
The addresses should correspond.
The relationship is closer to:
Service selector
|
v
matching Pods
|
v
EndpointSlice controller
|
v
EndpointSlices
|
v
networking implementation
Create a client Pod
We need something inside the cluster to test cluster networking.
Service selector = app=broken
Pods = app=api
No match
|
v
No ready backend endpoints
Fix desired state
kubectl patch service api \
--type=merge \
-p '{"spec":{"selector":{"app":"api"}}}'
Verify:
kubectl exec client -- wget -qO- http://api
A useful Service debugging order
When DNS resolves but traffic fails, walk the path rather than trying random commands:
Service
|
| selector and ports correct?
v
Pod labels
|
| selector matches?
v
EndpointSlice
|
| ready addresses present?
v
Pod readiness
|
| endpoint considered ready?
v
application port
|
| process actually listening?
v
application
We will revisit this after adding readiness probes.
Part IV - Configuration and Health
13. ConfigMaps: Configuration Without Rebuilding the Image [CKAD]
Our current nginx-based image has application content baked into it.
Suppose that content or configuration needs to vary between environments.
Rebuilding an image for every small configuration change is often the wrong abstraction.
A ConfigMap stores non-secret configuration in the Kubernetes API.
Create configuration
kubectl create configmap api-content \
--from-literal=index.html='hello from kubernetes'
Inspect:
kubectl get configmap api-content -o yaml
Query only the value:
kubectl get configmap api-content \
-o jsonpath='{.data.index\.html}{"\n"}'
ConfigMap-backed volumes are eventually updated by the kubelet.
After a short delay, retry:
kubectl exec client -- wget -qO- http://api
You should eventually see:
configuration changed
A subtle but important caveat:
A ConfigMap mounted using subPath does not receive the normal projected-volume updates.
That is one reason this example mounts the ConfigMap as the directory rather than mounting one key with subPath.
ConfigMap as environment variables
ConfigMaps can also populate environment variables.
The difference matters operationally:
ConfigMap volume
-> projected into filesystem
-> updates can appear later
ConfigMap environment variable
-> value captured when container starts
-> Pod must be recreated to see a new value
14. Secrets: Sensitive Configuration [CKAD]
Some configuration is sensitive enough that it should not be mixed casually with ordinary ConfigMaps.
init container starts
|
v
prepares shared data
|
v
init container completes
|
v
application container starts
Inspect the separate status lists:
kubectl get pod init-demo \
-o jsonpath='{range .status.initContainerStatuses[*]}init:{.name}={.state.terminated.reason}{"\n"}{end}{range .status.containerStatuses[*]}app:{.name}={.state.running.startedAt}{"\n"}{end}'
The init container terminated successfully before nginx began its normal lifetime.
Cleanup:
kubectl delete pod init-demo
rm -f init-demo.yaml
Native sidecar: start in init ordering, then stay alive
Native sidecars are restartable init containers.
They use:
restartPolicy: Always
Unlike an ordinary init container, the sidecar does not need to finish before the application can keep running.
You should see the application messages even though the log-forwarder was declared under initContainers.
Inspect its state:
kubectl get pod sidecar-demo \
-o jsonpath='{range .status.initContainerStatuses[*]}{.name}{" running="}{.state.running.startedAt}{" restarts="}{.restartCount}{"\n"}{end}'
The slightly surprising result is:
spec.initContainers
|
+-- ordinary init container -> eventually terminates
|
+-- restartPolicy: Always -> remains running as a sidecar
That is why native sidecars can participate in init ordering while still living for the Pod lifetime.
Good uses include:
log forwarding
local proxying
configuration synchronisation
security helpers
Use a sidecar when the supporting functionality genuinely belongs to the same Pod lifecycle.
Do not group unrelated services into one Pod merely because Kubernetes allows multiple containers.
Cleanup:
kubectl delete pod sidecar-demo
rm -f sidecar-demo.yaml
Part VI - Finite, Stateful and Node-Scoped Workloads
19. Jobs and CronJobs [CKAD]
A Deployment represents work that should keep running.
Part VII - Identity, Authorization and Runtime Security
23. Who Is Allowed to Do What? Users, ServiceAccounts, RBAC and Admission [CKAD] [DEV]
Until now we have mostly used kubectl as a highly privileged lab administrator.
That is useful for learning, but it hides an important production question:
Who should be allowed to do what?
There are several separate decisions in the Kubernetes API request path.
request
|
v
authentication
|
| who are you?
v
authorization
|
| may that identity perform this action?
v
admission
|
| for writes: is this object acceptable, or should it be mutated?
v
Kubernetes API state
Keeping those stages separate avoids a lot of security confusion.
Humans and workloads use different kinds of identity
Kubernetes commonly deals with two broad identity types:
human / external client workload inside Kubernetes
| |
v v
User / Group ServiceAccount
| |
+---------------+----------------+
|
v
RBAC
A ServiceAccount is a Kubernetes API object.
A normal human User is not.
Kubernetes does not provide a User resource that you create with:
kubectl create user alice
Instead, human authentication normally comes from something outside the Kubernetes object model, such as:
After authentication, the API server has identity information such as:
username: alice
groups:
- developers
RBAC then decides what that identity may do.
Lab: give Alice namespace-scoped developer access
We do not need to configure a real identity provider just to learn authorization.
kubectl can ask the API server to evaluate a request as another identity using impersonation.
--as=alice does not create Alice. It asks the API server to evaluate the request as the username alice. Your current identity must itself be allowed to impersonate users. The administrator credentials used by our kind lab normally are.
Alice
|
v
Kubernetes API
|
+-- read Pods in cookbook yes
+-- change Deployments in cookbook yes
+-- read Secrets in cookbook no
+-- change Deployments in default no
+-- delete Nodes no
You can ask for a broader view of the permissions Kubernetes calculates:
A Deployment controller can then create ReplicaSets and Pods on her behalf.
Alice
|
| create Deployment allowed
v
Deployment
|
v
Deployment controller
|
v
ReplicaSet
|
v
Pods
Authorization is evaluated against the API request Alice makes.
You therefore need to reason about what a permitted object can cause controllers to do, not merely about the object's name.
This becomes particularly important with powerful workload features such as privileged containers, host mounts and scheduling controls.
ServiceAccount: workload identity
Now do the same exercise for an application rather than a human.
Create a ServiceAccount:
kubectl create serviceaccount api-sa
Inspect it:
kubectl get serviceaccount api-sa -o yaml
Modern Kubernetes normally gives Pods short-lived projected ServiceAccount credentials rather than relying on automatically created permanent token Secrets.
Request a temporary token when the cluster allows it:
kubectl create token api-sa
Assign the ServiceAccount to our Deployment:
kubectl set serviceaccount deployment/api api-sa
Wait:
kubectl rollout status deployment/api
Check:
kubectl get deployment api \
-o jsonpath='{.spec.template.spec.serviceAccountName}{"\n"}'
Human and workload authorization now look almost identical after authentication:
User/alice --------------------+
|
ServiceAccount/api-sa ---------+
|
v
RBAC
|
allowed verbs
on resources
in a scope
Role, ClusterRole, RoleBinding and ClusterRoleBinding
RBAC has four main API objects.
Role
-> permission rules defined for one namespace
ClusterRole
-> reusable permission rules
-> can also describe cluster-scoped resources
RoleBinding
-> grants a Role or ClusterRole inside one namespace
ClusterRoleBinding
-> grants a ClusterRole across the cluster
A useful relationship is:
subject
|
| User / Group / ServiceAccount
v
binding
|
v
role containing rules
|
v
verbs + resources
For example:
alice
|
v
RoleBinding/cookbook
|
v
Role/developer
|
+-- get/list/watch Pods
+-- create/update/patch Deployments
Be especially careful with ClusterRoleBinding.
This:
Alice can administer one namespace
and this:
Alice can administer the whole cluster
can differ by only the binding used.
A namespace is a scope, not an automatic security boundary
Namespaces are extremely useful administrative boundaries.
But merely placing two teams in different namespaces does not automatically isolate them.
namespace alone
!=
permission boundary
You normally combine namespaces with controls such as:
The exact isolation you need depends on whether the tenants trust one another.
API permission, workload placement and machine access are different controls
This distinction matters particularly around control-plane nodes.
There are at least three independent questions:
1. API authorization
"Can Alice delete or modify this Node object?"
2. workload placement
"Can this Pod be scheduled onto this node?"
3. machine access
"Can Alice SSH into or otherwise administer the actual host?"
They are enforced by different layers:
control-plane machine
|
+-----------------------+-----------------------+
| | |
v v v
Kubernetes API scheduler target Linux / VM / metal
| | |
RBAC taints / tolerations IAM / SSH / firewall
affinity / selectors OS permissions
For example, control-plane nodes are commonly tainted so ordinary workloads do not schedule there:
node-role.kubernetes.io/control-plane:NoSchedule
That is a scheduling control.
It is not the same thing as denying API access to Node objects, and it is not the same thing as denying SSH access to the machine.
It is also not, by itself, a strong tenant security boundary: a workload that is allowed to specify a matching toleration can become eligible for the node.
Likewise:
kubectl auth can-i delete nodes --as=alice
answers an API authorization question.
It tells you nothing about whether Alice has infrastructure credentials for the underlying VM or bare-metal host.
What about the kubelet itself? [DEV] [DEEP DIVE]
Nodes are API clients too.
A kubelet commonly authenticates with an identity resembling:
system:node:worker-01
Kubernetes has a special-purpose Node authorizer that can constrain kubelet API access based on the Pods assigned to that node.
The NodeRestriction admission plugin adds additional restrictions around what kubelets may modify.
You do not need to configure these for CKAD, but knowing that node identity has its own authorization path prevents the misleading idea that every Kubernetes permission problem is just a RoleBinding.
RBAC is not the same thing as multi-tenancy [DEV] [DEEP DIVE]
RBAC can give multiple teams restricted access to one Kubernetes API:
Alice ----+
|
Bob ------+--> one kube-apiserver
| |
| RBAC
| |
+--> namespace-scoped views
That can be entirely appropriate for trusted teams.
But stronger tenancy may instead give each tenant its own Kubernetes API/control-plane boundary:
Alice ---> tenant A API
Bob -----> tenant B API
|
v
provider infrastructure
Projects such as vCluster and Kamaji operate in this design space, although they implement it differently.
RBAC still exists inside each tenant cluster. The difference is that the tenant boundary no longer depends only on permissions inside one shared API server.
The companion multitenancy-appendix.md continues this model and compares shared-cluster RBAC, vCluster and Kamaji without turning the CKAD path into a platform-engineering course.
Admission control
Authorization is not necessarily the final decision for a write request.
After authentication and authorization, admission can validate or mutate an incoming object before it is persisted.
Examples include:
Pod security requirements
quotas
policy rules
injected defaults
validating or mutating webhooks
This explains a useful failure class:
I am authenticated
|
v
I am authorized
|
v
write request still rejected
|
v
check admission or policy error
One subtle distinction: normal read operations such as get, list and watch do not pass through admission control in the same way write requests do.
The permission model to keep
When something is denied, ask which boundary you are actually debugging:
Who am I?
-> authentication
May I make this API request?
-> authorization / RBAC
Is this write acceptable?
-> admission
May this workload land on that node?
-> scheduling controls
May this workload talk to another workload?
-> NetworkPolicy / network controls
May this person administer the actual machine?
-> infrastructure IAM / SSH / OS controls
Those controls cooperate, but they are not substitutes for one another.
Cleanup the lab RBAC objects, but keep the ServiceAccount because the Deployment currently uses it:
runAsNonRoot
-> refuse a root runtime identity
runAsUser
-> choose a numeric runtime UID
allowPrivilegeEscalation: false
-> process cannot gain more privileges than its parent
capabilities.drop
-> remove specific Linux capabilities
readOnlyRootFilesystem
-> make the container root filesystem read-only
seccompProfile: RuntimeDefault
-> apply the runtime's default syscall filter
The broader lesson is the same as elsewhere in Kubernetes:
security intent in spec
|
v
runtime tries to satisfy it
|
+-- possible -> container runs
|
+-- impossible -> visible failure
Cleanup:
kubectl delete pod secure-demo
rm -f secure-demo.yaml
Part VIII - Network Policy and External HTTP Routing
25. NetworkPolicy: Which Pods May Talk? [CKAD]
A Service answers:
Where should traffic go?
A NetworkPolicy answers a different question:
Which network flows should be allowed?
NetworkPolicy enforcement depends on the cluster's CNI plugin.
If your CNI does not enforce NetworkPolicy, the objects can exist without changing packet flow.
kubectl get ingress api
kubectl describe ingress api
Inspect status directly:
kubectl get ingress api \
-o jsonpath='{.status.loadBalancer.ingress}{"\n"}'
If there is no controller, that status will normally remain empty and no HTTP listener will magically appear.
That is useful behaviour to understand:
Ingress stored in API yes
routing implementation no
external traffic path no
If your cluster already has a maintained Ingress controller, set the matching class:
spec:
ingressClassName: <class-name>
and follow that controller's documented exposure method.
The abstract path is still:
HTTP request
|
v
Ingress controller
|
| host/path rule
v
Service
|
v
ready backend Pods
Ingress does not replace a Service.
It routes to one.
Why we do not install an Ingress controller just for this lab
The Kubernetes project now recommends Gateway API instead of Ingress.
Ingress remains a stable API and is not being removed, but the API is frozen and no longer gaining features.
Also avoid old tutorials that tell you to install ingress-nginx: that project was retired in March 2026 and no longer receives fixes or security updates.
So the main cookbook keeps Ingress because it is still important Kubernetes and CKAD knowledge, but we do not introduce a legacy controller merely to make this one exercise route traffic.
Instead, the networking deep dive takes the modern path.
Continue with cillium-gateay-appendix.md, where we actually install and exercise Cilium's Gateway API implementation:
GatewayClass
|
v
Gateway
|
v
HTTPRoute
|
v
Service
|
v
Pod
That appendix sends real traffic, breaks backend references, inspects status, exercises header routing, weighted backends, cross-namespace ReferenceGrant, and follows the implementation down through Envoy, Cilium and eBPF.
The important progression is:
Ingress
-> simple, stable, frozen HTTP routing API
Gateway API
-> role-oriented, extensible service-networking APIs
For a broader tour of the Gateway API resource model and its routing features, Roman Glushko's deep dive is also excellent:
Server-side validation is especially useful when you want current cluster schema and admission behaviour to participate.
A reliable workflow is:
old example found online
|
v
kubectl api-resources / api-versions
|
v
kubectl explain
|
v
server-side dry run
|
v
apply
31. A Systematic Debugging Workflow [CKAD] [DEV]
Random commands make Kubernetes feel mysterious.
Most failures become easier when you follow the relationship between objects.
Workload failure
Follow ownership downward:
Deployment
|
v
ReplicaSet
|
v
Pod
|
v
Container
|
v
Process
Commands:
kubectl get deployment api
kubectl get rs -l app=api
kubectl get pods -l app=api
kubectl describe pod <pod>
kubectl logs <pod>
Previous crashed container instance:
kubectl logs <pod> --previous
Events:
kubectl get events \
--sort-by=.metadata.creationTimestamp
Questions to ask in order:
Did the controller create the expected child resource?
Did the Pod schedule?
Did the image pull?
Did the container start?
Did the process stay alive?
Did readiness succeed?
Networking failure
Follow the request path:
DNS
|
v
Service
|
v
selector
|
v
EndpointSlice
|
v
ready Pod
|
v
container port
|
v
process
Useful commands:
kubectl exec client -- nslookup api
kubectl get service api -o yaml
kubectl get pods -l app=api --show-labels
kubectl get endpointslices \
-l kubernetes.io/service-name=api
kubectl describe pod <pod>
Then test the application directly from inside the cluster when possible.
Configuration failure
Follow references:
Pod spec
|
+-- ConfigMap name/key
|
+-- Secret name/key
|
+-- volume name
|
+-- mount path
Useful commands:
kubectl describe pod <pod>
kubectl get configmap <name> -o yaml
kubectl get secret <name> -o yaml
A misspelled Secret key can prevent a container from starting even though the Secret object itself exists.
Authorization failure
Ask Kubernetes instead of guessing:
kubectl auth can-i get pods
kubectl auth can-i create deployments
For another identity, when impersonation is permitted:
kubectl auth can-i list pods \
--as=system:serviceaccount:cookbook:api-sa
The general debugging habit is:
Follow the API relationships until you find the first place where observed state stops matching your expectation.
32. Debugging Minimal Containers with Ephemeral Containers [DEV]
Production images may intentionally not contain:
bash
curl
dig
tcpdump
ps
That is often desirable.
A production image does not need a complete incident-response toolbox merely to make debugging convenient.
Let's hit that problem before solving it.
First try ordinary exec
Pick one API Pod:
POD=$(kubectl get pod -l app=api \ -o jsonpath='{.items[0].metadata.name}')
CONTAINER=$(kubectl get pod "$POD" \ -o jsonpath='{.spec.containers[0].name}')printf'pod=%s container=%s\n'"$POD""$CONTAINER"
Inside the debug container, call the application over the Pod's shared network namespace:
wget -qO- http://127.0.0.1
You can also inspect processes:
ps
Then exit:
exit
Inspect what Kubernetes added:
kubectl get pod "$POD" \
-o jsonpath='{range .spec.ephemeralContainers[*]}{.name}{" image="}{.image}{" target="}{.targetContainerName}{"\n"}{end}'
The original application image did not change.
The Pod now has temporary debugging tooling attached to it.
Conceptually:
minimal production container
|
| exec lacks tooling
v
debugging blocked
|
| kubectl debug
v
ephemeral container joins Pod
|
+-- same Pod network
+-- optional process targeting
+-- extra tools
Ephemeral containers are intentionally different from normal application containers:
they are added to an existing Pod for troubleshooting
they are not part of the normal workload template
they do not restart like ordinary workload containers
they cannot define normal container resources such as ports or probes
Depending on the runtime and security configuration, process visibility and debugging capabilities can differ.
The model to keep is:
production container stays minimal
+
temporary debug tooling when needed
rather than:
ship every debugging utility in every production image forever
Part XI - Extending Kubernetes
33. CRDs and Custom Resources [CKAD] [DEV]
Kubernetes' built-in API contains types such as:
Pod
Deployment
Service
Secret
Job
But platform teams often want domain-specific APIs.
A production controller normally uses watches, informers, a work queue, retries and often leader election.
We are deliberately starting with something smaller:
every 2 seconds
|
v
list PreviewEnvironments
|
v
for each one
|
+-- apply desired Deployment
+-- apply desired Service
+-- read observed Deployment status
+-- update PreviewEnvironment status
Polling is not the architecture we would choose for a serious controller.
It is useful here because the reconciliation logic stays visible.
The controller uses server-side apply for its child resources. Re-running the same desired definition therefore converges instead of creating another Deployment every loop.
Create the Go project
You need Go installed for this deep dive.
Check:
go version
Create a workspace:
mkdir -p /tmp/preview-controller
cd /tmp/preview-controller
go mod init example.com/preview-controller
go get \
k8s.io/apimachinery@v0.35.0 \
k8s.io/client-go@v0.35.0
client-go uses matching v0.X.Y versions for Kubernetes v1.X.Y releases, so v0.35.0 aligns with our Kubernetes 1.35 target.
PreviewEnvironment.spec
|
v
our Go controller
|
+-- server-side apply Deployment
|
+-- server-side apply Service
|
v
Deployment.status
|
v
PreviewEnvironment.status
Change desired state
Change the custom resource rather than the generated Deployment:
The controller sees the new desired state and changes the Deployment.
Create drift on purpose
Now fight the controller.
Scale its child Deployment directly:
kubectl scale deployment pr-482 --replicas=1
Check immediately:
kubectl get deployment pr-482
Then check again a few seconds later:
sleep 3
kubectl get deployment pr-482
It should return to three replicas.
Why?
PreviewEnvironment.spec.replicas = 3
|
v
controller observes child replicas = 1
|
v
server-side apply desired Deployment
|
v
child replicas = 3
This is Chapter 1's control loop implemented by code we wrote ourselves.
Delete a child
Delete the generated Deployment:
kubectl delete deployment pr-482
Watch the labelled Deployment set rather than asking for the temporarily missing object by name:
kubectl get deployment -l platform.example.com/preview=pr-482 -w
Within a reconciliation cycle, it should reappear.
Again, the controller is not issuing an imperative "restart" command.
It is repeatedly asserting:
A Deployment named pr-482 should exist with this spec.
Press Ctrl-C in the terminal running go run . before the next step.
Run the controller as a Kubernetes workload
Running from your laptop proved the control loop.
Now make the controller obey the same platform rules as every other workload.
Create a Dockerfile:
FROM golang:1.25-alpine AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY main.go ./
RUN CGO_ENABLED=0 GOOS=linux go build -o /controller .
FROM scratch
COPY --from=build /controller /controller
USER 65532:65532
ENTRYPOINT ["/controller"]
Build it:
docker build -t example/preview-controller:v1 .
Our lab is kind, so reuse the image-distribution lesson from Chapter 9:
kubectl apply -f controller-deployment.yaml
kubectl rollout status deployment/preview-controller
Follow its logs:
kubectl logs deployment/preview-controller -f
The controller now uses:
a ServiceAccount from the RBAC chapter
namespace discovery from the Downward API chapter
a hardened SecurityContext from Chapter 24
a locally built image loaded into kind as in Chapter 9
a custom API from Chapter 33
server-side apply to express child desired state
This is why the earlier chapters matter.
They compose.
Leave the controller, CRD and PreviewEnvironment/pr-482 running.
Chapter 35 will inspect the ownership and status relationships we just created.
What we deliberately left out
Our controller polls every two seconds because that keeps the first implementation understandable.
A production controller normally evolves toward:
watch / informer
|
v
work queue
|
v
reconcile(key)
|
+-- retry with backoff
+-- status / conditions
+-- metrics
+-- leader election when replicated
Libraries such as client-go and controller-runtime provide those building blocks.
The essential idea does not change:
Reconciliation should be idempotent. Running it repeatedly should converge the system toward the same desired state, not create more side effects every time.
35. Ownership, Finalizers, Status and Operators [DEV] [DEEP DIVE]
Once you understand controllers, several Kubernetes mechanisms become easier to place.
They are not arbitrary advanced features.
They help controllers express responsibility, cleanup and observed state.
We now have our own controller running, so we can inspect these mechanisms on something we built rather than only discussing them in the abstract.
Ownership: which object is responsible for this child?
First inspect the Deployment created for our preview:
kubectl get deployment pr-482 \
-o jsonpath='{.metadata.ownerReferences}{"\n"}'
Then the Service:
kubectl get service pr-482 \
-o jsonpath='{.metadata.ownerReferences}{"\n"}'
Both should point at:
PreviewEnvironment/pr-482
The relationship is:
PreviewEnvironment/pr-482
|
+-- Deployment/pr-482
| |
| v
| ReplicaSet
| |
| v
| Pods
|
+-- Service/pr-482
Compare that with the built-in chain we saw earlier:
Deployment
|
v
ReplicaSet
|
v
Pod
Our controller is using the same Kubernetes ownership mechanism as built-in controllers.
Labels and ownership are not the same thing:
label
-> which objects match this query/selector?
ownerReference
-> which object controls this dependent resource?
Controller responsibility: delete a child
Delete the generated Service:
kubectl delete service pr-482
Watch the labelled Service set:
kubectl get service -l platform.example.com/preview=pr-482 -w
Within a reconciliation cycle, the Service should reappear.
That happened because our controller still sees:
PreviewEnvironment/pr-482 exists
|
v
Service/pr-482 should exist
Ownership did not recreate the Service.
Reconciliation did.
That distinction matters.
Owner references describe relationships and enable garbage collection; controllers are the actors that continuously restore desired state.
Conditions are especially useful when readyReplicas: 1 is not enough to explain why the system is not ready.
Garbage collection: delete the owner
Now delete the PreviewEnvironment itself:
kubectl delete preview pr-482
Watch its children:
kubectl get deployment,service \
-l platform.example.com/preview=pr-482 \
-w
Because the Deployment and Service contain owner references to the custom resource, Kubernetes garbage collection can remove them when the owner disappears.
The path is:
PreviewEnvironment deleted
|
v
ownerReferences become invalid
|
v
garbage collector removes dependants
Our controller does not need explicit code saying:
on PreviewEnvironment delete:
delete Deployment
delete Service
for these Kubernetes-native child resources.
Recreate the preview so the rest of the chapter can continue:
We have now traversed the full extension lifecycle:
CRD
-> API type exists
Custom Resource
-> desired state exists
Controller
-> behaviour exists
ownerReferences
-> child relationships exist
status
-> observed state is reported
finalizers
-> external cleanup can become part of deletion
That is the same Kubernetes control model, extended with our own API and code.
Part XII - CKAD Speed Layer
37. CKAD Command Patterns [CKAD]
Understanding matters first.
Once the mental model is solid, speed comes from having a small number of productive command patterns in muscle memory.
authentication
-> who are you?
authorization / RBAC
-> may you make this API request?
admission
-> is this write acceptable?
scheduling controls
-> where may this workload run?
NetworkPolicy
-> which network flows are allowed?
infrastructure IAM / SSH
-> who may administer the actual machines?
Namespaces help provide scope, but are not an automatic security boundary by themselves.
For stronger tenancy, separate tenant Kubernetes APIs/control planes can add another boundary; see multitenancy-appendix.md.
Kubernetes extension uses the same model
CRD
-> teach the API a type
Custom Resource
-> declare desired state for that type
Controller
-> reconcile it
Operator
-> controller + domain-specific operational knowledge
Appendix B - Current CKAD Domain Map
This cookbook is intentionally organised for understanding rather than mirroring the exam outline chapter-for-chapter.
The current CKAD domains map roughly as follows:
Application Design and Build - 20%
Covered by:
container image definitions and builds
Pods and container command/args
workload selection
multi-container Pods
init containers and sidecars
Jobs and CronJobs
ephemeral and persistent volumes
StatefulSets and DaemonSets
Application Deployment - 20%
Covered by:
Deployments and ReplicaSets
scaling
rolling updates and rollback
blue-green and canary strategies
Helm
Kustomize
Application Observability and Maintenance - 15%
Covered by:
get, describe, logs and events
readiness, liveness and startup probes
rollout state
API deprecation and discovery
systematic debugging
ephemeral debug containers
Application Environment, Configuration and Security - 25%
Covered by:
requests and limits
quotas and limits policy concepts
ConfigMaps and Secrets
Downward API
ServiceAccounts
RBAC and authorization
admission concepts
SecurityContext and capabilities
CRDs and Operators
Services and Networking - 20%
Covered by:
Services
labels and selectors
EndpointSlices
DNS
network troubleshooting
NetworkPolicy
Ingress
Supplemental deep dive:
Gateway API concepts and hands-on Cilium Gateway API in cillium-gateay-appendix.md
Appendix C - Official References
Prefer current official documentation when a field, API version or exam detail is uncertain.
Later you may build your own container image locally.
For example:
docker build -t my-app:v1 .
That image exists on your machine, but it does not automatically exist inside the kind node.
Load it:
kind load docker-image my-app:v1 --name ckad
Then Kubernetes can use it:
kubectl run my-app \
--image=my-app:v1 \
--image-pull-policy=IfNotPresent
This:
docker build
|
v
Host image
|
| kind load docker-image
v
kind node
|
v
Pod
is useful when testing applications without pushing every image to a registry.
Avoid using the latest tag for this workflow. Kubernetes normally treats :latest as imagePullPolicy: Always, which may cause it to try pulling the image from a registry instead of using the image you loaded locally.
Reset the Lab
One of the best features of kind is that the entire cluster is disposable.
The lab assumes Ubuntu/Debian-style Linux machines.
For the best experience use two disposable Linux VMs:
k8s-control
2+ CPU
2+ GiB RAM
k8s-worker
2+ GiB RAM
Both machines must be able to reach each other directly.
The point is not merely to make Kubernetes work.
The point is to watch it not work yet, understand why, and then add the missing pieces.
C.1 What kind Was Hiding
In the local quickstart we used:
kind create cluster
and Kubernetes appeared.
That was real Kubernetes.
kind itself uses kubeadm to bootstrap its nodes.
This appendix performs the interesting pieces ourselves:
Linux
|
v
containerd
|
v
kubelet
|
v
kubeadm
|
v
control plane
|
v
Cilium
|
v
working cluster
By the end you should understand what sits underneath:
Deployment
Service
Pod
and why a Kubernetes cluster can exist while still being:
NotReady
C.2 The Cluster We Are Building
Our final topology:
k8s-control
|
+------------+-------------+
| | |
v v v
API server scheduler controller
|
v
etcd
|
|
Kubernetes API :6443
|
+------+------+
| |
v v
k8s-control k8s-worker
| |
kubelet kubelet
| |
containerd containerd
| |
+------+------+
|
Cilium
|
v
Pod network
We will deliberately not install kube-proxy.
Cilium will later provide:
CNI
+
Pod networking
+
NetworkPolicy enforcement
+
Kubernetes Service load balancing
+
kube-proxy replacement
C.3 Know the Pieces
Before installing anything, keep these roles separate.
kubectl
Client for the Kubernetes API.
kubectl
|
v
kube-apiserver
kubelet
Node agent.
It watches for Pods assigned to its node and asks the container runtime to run them.
API server
|
v
kubelet
|
v
containerd
containerd
Container runtime.
The kubelet communicates with it through CRI:
kubelet
|
| CRI
v
containerd
|
v
container
kubeadm
Bootstrap and lifecycle tool.
kubeadm
|
v
creates/configures Kubernetes
It is not the daemon continuously running the cluster.
Cilium
Networking implementation.
Kubernetes networking APIs
|
v
Cilium
|
v
Linux / eBPF dataplane
C.4 Prepare Both Machines
Run this section on:
k8s-control
k8s-worker
Check hostname:
hostname
Set unique names if necessary.
Control plane:
sudo hostnamectl set-hostname k8s-control
Worker:
sudo hostnamectl set-hostname k8s-worker
Log out and back in if your shell prompt does not immediately reflect the change.
Check addresses:
ip -4 addr
Make sure each machine can reach the other:
ping -c 2 <other-node-ip>
C.5 Disable Swap
Check:
swapon --show
For this lab disable swap:
sudo swapoff -a
Check again:
swapon --show
It should return nothing.
For a persistent lab also disable the swap entry in:
/etc/fstab
Modern Kubernetes can be configured to use swap, but the default kubelet behaviour used by this lab expects it to be disabled.
C.6 Enable IPv4 Forwarding
Create:
cat <<'EOF' | sudo tee /etc/sysctl.d/k8s.confnet.ipv4.ip_forward = 1EOF
Apply:
sudo sysctl --system
Check:
sysctl net.ipv4.ip_forward
Expected:
net.ipv4.ip_forward = 1
This is our first reminder that Kubernetes networking eventually becomes ordinary Linux networking.
It does not create an entire application platform.
C.33 The Whole Bootstrap Sequence
Linux
|
v
containerd
|
| CRI
v
kubelet
|
^
|
kubeadm
|
v
control plane
|
+--> API server
+--> etcd
+--> scheduler
+--> controller-manager
|
v
Kubernetes exists
|
| but
v
Nodes NotReady
|
| install
v
Cilium
|
+--> CNI
+--> eBPF dataplane
+--> kube-proxy replacement
|
v
Nodes Ready
|
v
Pods / Services / Deployments
Compare that to:
kind create cluster
kind was not fake Kubernetes.
It was automating most of this journey for us.
C.34 The Model to Remember
The CKAD cookbook started with:
Desired state
|
v
Kubernetes API
|
v
Controller
|
v
Actual state
We can now expand it:
desired state
|
v
Kubernetes API
|
+------------+------------+
| | |
v v v
scheduler controllers Cilium
| | |
+------------+------------+
|
v
kubelet
|
v
containerd
|
v
Linux kernel
|
v
actual state
Appendix - From RBAC to Multi-Tenant Kubernetes [DEV] [DEEP DIVE]
This companion starts where the main cookbook's RBAC chapter stops.
The question is no longer only:
What may this identity do inside Kubernetes?
It is now:
What boundary should exist between tenants in the first place?
This is platform-engineering material, not CKAD material.
The goal is not to memorise product-specific commands. We will use ordinary RBAC, vCluster and Kamaji as hands-on experiments to make API, control-plane and worker isolation concrete.
By the end we will have built three different models:
one shared API
+ namespace RBAC
separate tenant API
+ shared workers
separate tenant API
+ hosted control plane
+ independently joined workers
The vCluster lab assumes you are continuing from the cookbook's existing kind-ckad cluster. The Kamaji lab deliberately creates a second kind cluster so that experimentation does not disturb the CKAD environment.
1. Keep the Boundaries Separate
A useful tenancy model has several layers.
Layer 1 - identity and API authorization
----------------------------------------
authentication
RBAC
admission
"Can Alice perform this API operation?"
Layer 2 - API / control-plane isolation
---------------------------------------
shared kube-apiserver
or
separate tenant kube-apiservers
"Does Alice even share a Kubernetes API with Bob?"
Layer 3 - workload isolation
----------------------------
namespaces
Pod security
scheduler policy
taints / affinity
separate worker pools
runtime sandboxing
"Can their workloads interfere on compute?"
Layer 4 - network, storage and infrastructure isolation
-------------------------------------------------------
NetworkPolicy
CNI / VPC / VLAN / VRF
CSI and storage policy
VM boundaries
bare-metal allocation
cloud IAM
"What underlying infrastructure do tenants share?"
Layer 5 - provider management plane
-----------------------------------
management cluster
cluster lifecycle controllers
provisioning systems
hardware / cloud APIs
"Who is allowed to change the infrastructure itself?"
No single layer replaces all the others.
For example:
separate API servers
!=
separate kernels
NoSchedule taint
!=
authorization boundary
Kubernetes cluster-admin
!=
SSH root on the machine
We are going to prove those distinctions rather than only state them.
2. Lab Zero: One Shared Cluster + Namespace RBAC
Before adding virtual or hosted control planes, establish the baseline.
Make sure we are on the cookbook cluster:
kubectl config use-context kind-ckad
Save the provider/host context. We will use this later because vCluster changes our current context for us:
Alice
|
v
same kube-apiserver as everyone else
|
v
RBAC
|
+-- shared-alice namespace some access
+-- cluster-scoped objects mostly no access
+-- other tenant namespaces no access unless granted
This is a legitimate tenancy model for many internal platforms.
But Alice still talks to the same Kubernetes API as everybody else.
That is the limitation we will explore next.
3. Why Give a Tenant Another Kubernetes API?
Suppose Alice needs broad Kubernetes freedom.
Inside a normal shared cluster, this would be dangerous:
Alice
|
v
cluster-admin
|
v
shared cluster
cluster-admin is intentionally enormous.
A platform provider often wants something different:
Alice
|
v
Alice's Kubernetes API
|
| cluster-admin is okay HERE
v
Alice's tenant cluster
provider
|
v
provider Kubernetes API
|
| Alice does NOT get these credentials
v
provider infrastructure
This changes the question from:
How carefully can I restrict Alice inside my cluster?
into:
Why should Alice be an administrator of my cluster at all?
That distinction is central to virtual clusters, hosted control planes and many managed Kubernetes systems.
4. Lab: Give Alice a vCluster
vCluster gives a tenant its own Kubernetes API while allowing the platform operator to host the tenant control plane on existing infrastructure.
We will begin with its shared-node model because we can run the whole experiment on our existing kind cluster.
4.1 Install the vCluster CLI
On macOS with Homebrew:
brew install loft-sh/tap/vcluster
Verify it:
vcluster --version
Make sure the host context is still our cookbook cluster:
kubectl config use-context "$HOST_CONTEXT"
4.2 Create Alice's tenant cluster
Create a tenant cluster called alice inside the provider namespace tenant-alice:
vcluster create alice \
--namespace tenant-alice
The vCluster CLI automatically connects you to the new tenant cluster when creation completes.
Check the current context:
kubectl config current-context
Then ask the API server what namespaces exist:
kubectl get namespaces
You should see an ordinary Kubernetes-looking namespace view such as:
default
kube-node-lease
kube-public
kube-system
Alice is not looking at the provider cluster's namespace list.
She is talking to another Kubernetes API.
Conceptually:
kind-ckad
provider Kubernetes
API
|
v
tenant-alice
|
vCluster CP
+ API server
+ controllers
+ datastore
+ syncer
|
v
Alice
5. Alice Can Be an Administrator of Her Cluster
By default, the kubeconfig generated by vcluster connect uses tenant administrator credentials.
TENANT API
Alice's tenant credential
|
v
tenant cluster-admin
|
v
YES
PROVIDER API
Alice
|
v
provider RBAC
|
+-- delete Nodes? NO
+-- create ClusterRoleBinding? NO
cluster-admin is not a magical global property attached to a human being.
It is authorization against a particular Kubernetes API.
In this lab we use vCluster-generated credentials and Kubernetes impersonation to make the boundary obvious. A production platform might authenticate the same human through OIDC or another identity provider on both APIs and assign different authorization in each one.
6. Issue a Less Powerful Tenant Credential
A tenant does not need to give every user administrator access either.
vCluster can generate a kubeconfig backed by a ServiceAccount and bind that ServiceAccount to a tenant-local ClusterRole.
kubectl
|
v
Alice tenant kube-apiserver
|
v
tenant Pod object
|
v
vCluster syncer
|
v
provider kube-apiserver
|
v
translated Pod
|
v
provider scheduler
|
v
provider node / kubelet
When the provider-side Pod changes status, vCluster synchronizes that observation back to the tenant API.
So even here we are still using the same control-loop model from Chapter 1:
desired tenant object
|
v
translation / reconciliation
|
v
provider object
|
v
actual workload
|
v
status flows back
10. Shared API Isolation Is Not Worker Isolation
We have proven that Alice has a separate Kubernetes API.
But in this default shared-node model, her nginx workload ultimately runs on the same provider worker infrastructure as other workloads.
Tenant A API Tenant B API
| |
v v
translated Pods translated Pods
\ /
\ /
v v
provider nodes
|
v
shared kernel
So:
separate tenant API yes
separate tenant RBAC yes
separate host namespace yes
separate physical node not necessarily
separate kernel no, not in shared-node mode
Current vCluster guidance treats shared nodes as appropriate for trusted tenants such as internal development, CI and testing.
It explicitly does not treat this model as the worker security boundary for untrusted external tenants with arbitrary Kubernetes workload access.
That is not a defect in RBAC.
It is a different layer of the architecture.
11. A Small Shared-Node Hardening Recipe
Even for trusted tenants, the provider should not assume the separate API is sufficient on its own.
For example, vCluster can create host-side NetworkPolicy around the tenant workload namespace:
Hardening improves the shared-node model. It does not change its fundamental trust boundary.
restricted Pod Security may also break workloads that assume root privileges. Treat that as useful feedback about the workload rather than blindly weakening the platform baseline.
12. What Would vCluster Private Nodes Change?
vCluster also supports a model where tenant workers are not shared with the provider worker pool.
This requires vCluster Platform and separate Linux worker machines, so it is not part of our simple kind-ckad lab.
Alice tenant API
|
v
Alice private workers
Bob tenant API
|
v
Bob private workers
The provider-hosted control plane and tenant compute have become separate choices.
13. Clean Up the vCluster Lab
Before moving to Kamaji, delete Alice's vCluster:
kubectl config use-context "$HOST_CONTEXT"
vcluster delete alice \
--namespace tenant-alice
Remove the baseline namespace too:
kubectl delete namespace shared-alice
Our original CKAD cluster remains intact.
14. Kamaji: Host the Tenant Control Plane, Not Its Workers
Kamaji approaches the broad problem differently.
Instead of every tenant owning dedicated control-plane VMs, Kamaji runs upstream Kubernetes control-plane components as workloads inside a provider-operated Management Cluster.
Management Cluster
Kamaji operator
|
+----------------+----------------+
| |
v v
Tenant A control plane Tenant B control plane
kube-apiserver kube-apiserver
controller-manager controller-manager
scheduler scheduler
| |
v v
tenant A API endpoint tenant B API endpoint
| |
v v
tenant A workers tenant B workers
This creates a clean separation:
control-plane lifecycle
|
v
provider management cluster
worker lifecycle
|
v
VMs / bare metal / Cluster API / other provisioning
Let's build one.
15. Lab: Create a Kamaji Management Cluster on kind
Kamaji's official kind walkthrough is intended for development and learning only.
We will create a separate kind cluster named kamaji.
The pinned version above follows the current Kamaji kind walkthrough when this appendix was written. If upstream moves on, prefer the dependency versions in the current official guide.
17. Install MetalLB for Tenant API Endpoints
A Kamaji Tenant Control Plane needs an API endpoint.
In this kind lab, the official walkthrough uses MetalLB to provide LoadBalancer addresses on the Docker kind network.
NAME VERSION STATUS CONTROL-PLANE ENDPOINT KUBECONFIG
k8s-133 ... Ready ...:6443 k8s-133-admin-kubeconfig
Stop the watch with Ctrl-C.
Inspect what Kamaji created in the management cluster:
kubectl get tcp,deploy,pods,svc
The conceptual chain is:
TenantControlPlane
|
v
Kamaji controller
|
v
Deployment / Service / certificates / datastore state
|
v
running tenant kube-apiserver
+ controller-manager
+ scheduler
This is the same spec -> controller -> status model we began the entire cookbook with.
Kamaji is simply using it to create Kubernetes control planes.
20. Retrieve the Tenant kubeconfig
Kamaji writes the tenant administrator kubeconfig into a Secret named after the TenantControlPlane.
kubectl --kubeconfig=/tmp/kamaji-tenant.conf \
get namespaces
Now the important command:
kubectl --kubeconfig=/tmp/kamaji-tenant.conf \
get nodes
Expected:
No resources found.
This is not an error.
We successfully have:
kube-apiserver yes
controller-manager yes
scheduler yes
Kubernetes API yes
worker node no
That gives us a clean mental model:
control plane exists
!=
compute exists
The tenant has a Kubernetes cluster control plane before it has somewhere to run application Pods.
22. macOS / Docker Desktop Note
On native Linux, the MetalLB address on the Docker kind network is often directly reachable from the host.
On macOS with Docker Desktop, the Docker bridge network may not be routed directly into macOS.
That means this can happen:
TenantControlPlane Ready
LoadBalancer IP allocated
kubeconfig valid
but
kubectl from macOS cannot route to that Docker-network IP
Do not interpret that as Kamaji reconciliation failing.
First confirm from the management side:
kubectl --context "$KAMAJI_CONTEXT" get tcp
and:
kubectl --context "$KAMAJI_CONTEXT" get svc
The official Kamaji kind guide calls out this Docker bridge/macOS case and suggests running tenant API checks from an environment that can reach the kind Docker network, including the kind control-plane container when appropriate.
The networking lesson is useful in its own right:
API exists
!=
my current machine has a route to that API
Production Kamaji environments expose the API deliberately with normal LoadBalancer, DNS, Gateway or equivalent networking rather than relying on Docker Desktop bridge routing.
23. What Would Joining a Worker Look Like?
Kamaji deliberately does not create tenant worker machines for you.
Workers might come from:
cloud VMs
bare metal
Cluster API
an infrastructure platform
manual provisioning
Once a Linux machine has the required container runtime, kubelet and kubeadm components installed, the tenant control plane can generate an ordinary kubeadm join command.
A control-plane NoSchedule taint is useful placement policy.
It is not a hard authorization boundary if Alice is allowed to submit arbitrary tolerations.
26.3 May Alice log into the machine?
That is infrastructure access:
cloud IAM
SSH keys
VPN / firewall
bastions
OS users
Kubernetes RBAC does not remove Alice's SSH key.
26.4 Does Alice need to see the provider API at all?
That is the architectural question we explored here:
shared cluster
|
+-- Alice and provider use same API
versus
tenant control plane
|
+-- Alice uses tenant API
+-- provider API remains provider-only
This fourth option can dramatically reduce how much permission engineering has to happen inside the provider cluster.
27. Strong Tenancy Is Still a Stack
Even a separate tenant kube-apiserver does not solve every problem.
For an untrusted external tenant, a design may need something closer to:
Tenant identity
|
v
Tenant API server
|
+-- tenant-local RBAC
+-- admission / policy
|
v
Dedicated or strongly isolated compute
|
+-- runtime security
+-- device isolation
|
v
Tenant network boundary
|
+-- NetworkPolicy
+-- VPC / VLAN / routing policy
|
v
Tenant storage boundary
|
+-- CSI / volume policy
+-- encryption / credentials
|
v
Provider management plane
|
X tenant credentials do not cross this boundary
Do not collapse those controls into one mental bucket.
RBAC
!= NetworkPolicy
!= scheduler placement
!= node isolation
!= VM isolation
!= storage isolation
!= host IAM
28. A Practical Platform Decision Sequence
When designing a platform, ask these questions in order.
1. Are the tenants mutually trusted?
|
+-- yes -> shared cluster + namespace/RBAC may be sufficient
|
+-- no -> continue
2. Do tenants need broad Kubernetes administration?
|
+-- yes -> consider separate tenant API/control-plane boundaries
3. Can tenants execute arbitrary workloads?
|
+-- yes -> decide what worker/kernel isolation is required
4. Can workloads communicate across tenants?
|
+-- no -> enforce network boundaries
5. Can storage or devices be shared safely?
|
+-- design CSI/device/IOMMU/etc. boundaries accordingly
6. Can tenant credentials reach the provider management plane?
|
+-- they generally should not
Then pick technology.
Do not begin with:
"We should use vCluster."
Begin with:
"What boundary are we trying to create?"
29. Connect This Back to the Main Cookbook
Nearly everything in this appendix is built from concepts we already learned.
API object
|
v
desired state
|
v
controller
|
v
lower-level infrastructure
|
v
status
A Deployment reconciles Pods.
A vCluster syncer reconciles tenant resources into provider resources.
Kamaji reconciles a TenantControlPlane into a running Kubernetes control plane.
Cluster API can reconcile a cluster specification into worker machines.
Once the control-loop model is clear, these systems stop looking magical.
30. Cleanup the Kamaji Lab
When finished, return to the CKAD context:
kubectl config use-context kind-ckad
Delete the entire Kamaji learning environment:
kind delete cluster --name kamaji
Remove the temporary tenant kubeconfig:
rm -f /tmp/kamaji-tenant.conf
Your original cookbook cluster remains available.
31. What You Should Remember
If you retain only a few things from this appendix, make them these:
RBAC answers:
"What can this identity do against this Kubernetes API?"
A separate tenant API answers:
"Why should this tenant be an administrator of my provider API at all?"
Separate API servers do not automatically mean separate workers or kernels.
cluster-admin is relative to a cluster/API boundary.
control plane exists != worker nodes exist
and:
a strong tenant boundary is normally a stack of controls,
not one Kubernetes object or one product.
32. Official References
These projects evolve quickly. The commands and architecture above were aligned with their current documentation when this appendix was written; use current upstream docs when making a real platform decision.