A Second NodePool: Kata Containers on ARO HCP
Value Proposition
Standard OpenShift workloads run inside containers that share the host kernel. For most enterprise workloads this is acceptable. For regulated, multi-tenant, or high-risk workloads — processing untrusted code, running third-party payloads, handling data with strict isolation requirements — kernel sharing is an unacceptable risk surface.
OpenShift Sandboxed Containers addresses this by running each pod inside a lightweight KVM virtual machine managed by Kata Containers. The pod API is unchanged — no application manifest modifications are required — but the isolation boundary shifts from namespace and cgroup to a hardware-enforced VM boundary. A compromised container cannot escape to the host kernel; it can only reach the Kata guest kernel.
In ARO HCP, this capability requires a dedicated NodePool targeting a VM SKU that supports nested virtualization — the ability to run KVM inside an Azure VM. The Standard_D4s_v5 SKU used by the existing worker NodePool supports nested virtualization, so no additional VM family or quota is required. Kata nodes are kept separate from standard workloads via a NoSchedule taint, ensuring runc containers never land on Kata nodes unintentionally.
The kata-nodepool NodePool was already provisioned at the end of Module 1. This module labels and taints those nodes, installs the Sandboxed Containers Operator, and examines the three resources the operator manages. Module 3 completes the configuration by deploying KataConfig and running the first sandboxed workload.
Learning Objectives
By the end of this module you will be able to:
-
Confirm the Kata NodePool provisioned in Module 1 is ready and inspect its configuration
-
Apply node labels and taints that bind the NodePool to KataConfig in Module 3
-
Verify KVM availability on Kata nodes from the hosted cluster context
-
Deploy the Sandboxed Containers Operator via an OLM Subscription
-
Identify the three resources managed by the Sandboxed Containers Operator and explain each one’s role
-
Evaluate
installPlanApprovalsettings for production operator lifecycle management
Section 1 — Create the Kata NodePool
Why a Dedicated NodePool
Kata Containers require KVM — a hardware-enforced hypervisor that runs each pod in its own VM. Two constraints follow:
-
VM SKU: Not all Azure VM SKUs support nested virtualization. The Standard_D4s_v5 SKU used by the existing worker NodePool supports it — no separate SKU family or additional quota is needed.
-
Node isolation: Each Kata pod carries per-VM overhead. Running standard runc workloads on Kata nodes wastes that overhead. A NoSchedule taint enforces the boundary.
A node label ties this NodePool to KataConfig in Module 3. KataConfig uses a kataConfigPoolSelector to identify which nodes should receive Kata binaries. If the label on this NodePool does not exactly match that selector, KataConfig will proceed silently with zero nodes configured — the most common Sandboxed Containers deployment error.
Confirm the NodePool Is Ready
The kata-nodepool NodePool was created at the end of Module 1. Confirm it is fully provisioned before proceeding.
-
Check the NodePool provisioning state:
az aro hcp cluster nodepool show \ --cluster-name {aro_hcp_cluster_name} \ --resource-group {aro_hcp_cluster_rg} \ --name kata-nodepool \ --query "{State:properties.provisioningState, Replicas:properties.replicas, VMSize:properties.platform.vmSize}" \ --output tableprovisioningStatemust beSucceededbefore continuing.
Label and Taint the Kata Nodes
Once the nodes are Ready, apply the label and taint that Module 3 depends on.
-
Confirm the Kata nodes have joined the cluster and identify them by instance type:
oc get nodes -o custom-columns=NAME:.metadata.name,INSTANCE:.metadata.labels."node\.kubernetes\.io/instance-type"The
Standard_D4s_v5nodes are the Kata NodePool. The existing workers showStandard_D4s_v6. -
Apply the Kata node label using the instance-type as the selector — ARO HCP does not surface the AKS-style agentpool label, so instance type is the reliable differentiator:
oc label nodes -l node.kubernetes.io/instance-type=Standard_D4s_v5 \ node-role.kubernetes.io/kata=
|
The label |
-
Apply the NoSchedule taint using the same instance-type selector:
oc adm taint nodes -l node.kubernetes.io/instance-type=Standard_D4s_v5 \ kata=true:NoSchedule -
Confirm both label and taint are applied:
oc get nodes -l node-role.kubernetes.io/kata= \ -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taintsEach Kata node should show the
kata=true:NoScheduletaint.
Verify KVM Availability
KVM must be present on each Kata node for the Kata runtime to function. Confirm this before proceeding to the operator installation.
-
Capture a Kata node name:
KATA_NODE=$(oc get nodes -l node-role.kubernetes.io/kata= \ -o jsonpath='{.items[0].metadata.name}') echo $KATA_NODE -
Open a debug shell on the Kata node:
oc debug node/$KATA_NODE -
Inside the debug shell, verify the KVM device is present:
ls -la /host/dev/kvmExpected output: a character device at
/host/dev/kvm. If this path is absent, nested virtualization is not enabled on this VM SKU in this Azure region — Kata will not function and a different SKU or region is required. -
Exit the debug shell:
exit
Section 2 — Install the Sandboxed Containers Operator
The Sandboxed Containers Operator manages the lifecycle of Kata binaries and runtime configuration on the cluster. It is installed from the Red Hat operator catalog via a standard OLM Subscription.
Apply the Operator Subscription
The installation requires three resources in sequence: a Namespace, an OperatorGroup, and a Subscription.
-
Apply all three resources in one step:
oc apply -f - <<'EOF' apiVersion: v1 kind: Namespace metadata: name: openshift-sandboxed-containers-operator --- apiVersion: operators.coreos.com/v1 kind: OperatorGroup metadata: name: openshift-sandboxed-containers-operator namespace: openshift-sandboxed-containers-operator spec: targetNamespaces: - openshift-sandboxed-containers-operator --- apiVersion: operators.coreos.com/v1alpha1 kind: Subscription metadata: name: sandboxed-containers-operator namespace: openshift-sandboxed-containers-operator spec: channel: stable installPlanApproval: Automatic name: sandboxed-containers-operator source: redhat-operators sourceNamespace: openshift-marketplace EOF
Watch the CSV Reach Succeeded
-
Watch the ClusterServiceVersion (CSV) status:
oc get csv -n openshift-sandboxed-containers-operator -wWait for the
PHASEcolumn to showSucceeded. This typically takes 2–3 minutes as OLM downloads the operator bundle and starts the manager pod. PressCtrl+Conce it reachesSucceeded. -
Confirm the operator pod is running:
oc get pods -n openshift-sandboxed-containers-operatorThe operator manager pod should be in
Runningstate.
Section 3 — Operator Resources and Production Considerations
The Three Resources the Operator Manages
The Sandboxed Containers Operator installs three CRDs. Examine them:
oc get crd | grep -E 'kata|peerpod'
Each CRD represents a distinct capability:
-
kataconfigs.kataconfiguration.openshift.io— triggers installation of Kata binaries onto labeled nodes via MachineConfig. When you create a KataConfig resource in Module 3, the operator generates a MachineConfig that rolls onto every node matching thekataConfigPoolSelector. Once the MachineConfig is applied, the Kata runtime is installed and the containerd configuration is updated on each matching node. -
RuntimeClass(created by the operator, not a CRD) — registers thekatahandler with containerd. After KataConfig completes, a RuntimeClass namedkataappears in the cluster. Pods that setruntimeClassName: katain their spec use the Kata VM runtime instead of runc. No other change to the pod spec is required. -
peerpodsconfigs.confidentialcontainers.org— configures PeerPod mode. In PeerPod mode, each pod runs in a separate cloud provider VM rather than a nested VM on the worker node. This enables Kata-style isolation on VM SKUs that do not support nested virtualization, at the cost of longer pod start times. PeerPod mode is not covered in this lab.
Examine the Installed CSV
-
Confirm what CRDs the operator owns:
oc get csv -n openshift-sandboxed-containers-operator \ -o jsonpath='{.items[0].spec.customresourcedefinitions.owned[*].name}' \ | tr ' ' '\n'This lists every CRD the operator is responsible for managing — these are the resources you can create to drive operator behaviour.
Production Consideration: installPlanApproval
-
Check the Subscription’s
installPlanApprovalsetting:oc get subscription sandboxed-containers-operator \ -n openshift-sandboxed-containers-operator \ -o jsonpath='{.spec.installPlanApproval}'This returns
Automatic. WithAutomaticapproval, OLM upgrades the operator whenever a new version appears in thestablechannel without any human gate.
For production environments that require change-controlled upgrades:
-
Set
installPlanApproval: Manualin the Subscription -
Review pending upgrades:
oc get installplan -n openshift-sandboxed-containers-operator -
Approve a specific plan:
oc patch installplan <name> -n openshift-sandboxed-containers-operator --type merge -p '{"spec":{"approved":true}}'
Automatic is appropriate for lab and development environments. For production, align the approval mode with your change management process.
Summary
What You Learned
-
A dedicated NodePool is required for Kata workloads. The Standard_D4s_v5 SKU supports nested virtualization (KVM) — no additional quota or VM family is needed. The node label must exactly match the
kataConfigPoolSelectorin KataConfig — a mismatch causes silent, partial installation. The NoSchedule taint prevents runc workloads from incurring Kata overhead. -
KVM availability is verifiable before KataConfig runs.
oc debug node/andls /host/dev/kvmconfirms the hypervisor device is present — the necessary condition for Kata to function. -
The Sandboxed Containers Operator manages three resources. KataConfig installs Kata binaries via MachineConfig. RuntimeClass registers the kata handler with containerd. PeerPodConfig enables Kata-style isolation on SKUs without nested virtualization.
-
Operator upgrade policy is a production decision.
AutomaticinstallPlanApproval is convenient;Manualis appropriate for environments where operator upgrades require change management approval.
Key Takeaways for Customer Conversations
"Why do I need a separate NodePool for Kata?"
Two reasons: VM SKU and isolation. Kata requires KVM, which is only available on certain Azure VM SKUs. And each Kata pod carries per-VM overhead that runc workloads should not pay for. A dedicated NodePool with a NoSchedule taint enforces this separation without application changes.
"What if my VM SKU doesn’t support nested virtualization?"
Use PeerPod mode. Instead of a nested VM on the worker node, each pod runs in a separate Azure VM. The isolation boundary is still hardware-enforced, but the Kata guest runs in Azure rather than on the worker. This works on any SKU at the cost of longer pod start times.
"What happens if the node label doesn’t match the KataConfig selector?"
KataConfig reports success but installs Kata binaries on zero nodes. Verify the match before applying KataConfig: oc get nodes -l node-role.kubernetes.io/kata= should list your Kata NodePool workers. If it returns nothing, the label was not applied or does not match the selector.
"Is the Sandboxed Containers Operator supported on ARO HCP?"
Yes. The operator runs on worker nodes — entirely within the customer’s data plane. The Kata runtime and KVM are worker-node concerns. The hosted control plane has no involvement in Kata pod scheduling or VM management.