Where Is the Control Plane? Proving the HCP Boundary

Value Proposition

Every OpenShift administrator has an intuition about what a cluster contains: master nodes running the API server, scheduler, etcd, and controller manager, alongside worker nodes running workloads. ARO HCP breaks that intuition deliberately. The control plane exists — the API server is responding, you used it to log in — but its pods are invisible from the customer’s context. They run in Red Hat’s managed infrastructure, outside the customer’s subscription entirely.

The NodePool object defines exactly where that boundary falls. Everything above the NodePool — the hosted control plane components, the API server load balancer, etcd — is Red Hat’s responsibility. Everything at and below the NodePool — worker node configuration, workloads, namespaces — is the customer’s. Understanding this line is essential for field conversations about support escalation, incident response, and SLA scope.

The decoupling also changes the upgrade model. In standard ARO, the control plane and worker nodes are upgraded together. In ARO HCP, the control plane is upgraded independently by Red Hat SRE — transparently, without customer action. The customer manages only the worker NodePool version, and workers can remain at a different OpenShift version than the control plane during a rolling upgrade window.

This module makes the boundary concrete: you confirm empirically from the OCP layer that no control plane pods are visible, then exercise the customer’s side of the boundary directly by inspecting and scaling a NodePool through the Azure CLI and watching ARO HCP translate a single declarative API call into an Azure VM creation.

Learning Objectives

By the end of this module you will be able to:

  • Prove from the OCP layer that no control plane pods are visible in the customer’s cluster context and explain why

  • Identify the NodePool as the customer’s primary operational surface and explain the support boundary in practice

  • Inspect NodePool configuration using the az aro hcp CLI

  • Scale a worker NodePool and trace the resulting Azure VM provisioning

  • Explain the ARO HCP operational model to a customer using live evidence

Section 1 — The Control Plane Is Not in Your Cluster

Confirm a Worker-Only Node View

In a self-managed OpenShift cluster, oc get nodes lists both master and worker nodes. In ARO HCP, the customer’s cluster context shows only worker nodes — master nodes run in Red Hat’s managed infrastructure and are not part of the customer’s node inventory.

oc get nodes

Every node listed has the worker role. No master or control-plane role is present. The cluster is responding through an API server that is itself invisible to this view.

Confirm Empty Control Plane Namespaces

In self-managed OpenShift, the openshift-etcd and openshift-kube-apiserver namespaces contain running control plane pods. In ARO HCP, those pods run in Red Hat’s managed infrastructure.

  1. Attempt to list pods in the etcd namespace:

    oc get pods -n openshift-etcd

    In self-managed OpenShift, this namespace contains three etcd pods running the distributed key-value store for all cluster state. In ARO HCP, these pods run in Red Hat’s infrastructure and are not visible from the customer’s context.

  2. Confirm the API server pods are equally absent:

    oc get pods -n openshift-kube-apiserver

    This is the key evidence: the API server is responding — every oc command you have run proves it — but its implementation is outside the customer’s view. The control plane exists; you simply do not own it.

Verify Operators Report Healthy Without Visible Pods

The most counterintuitive aspect of ARO HCP: cluster operators report Available=True and Degraded=False — but the pods implementing them are invisible from the customer’s context.

oc get clusteroperator

Scan the AVAILABLE, PROGRESSING, and DEGRADED columns. Operators like kube-apiserver and kube-controller-manager report healthy — these are control plane operators whose pods run in Red Hat’s managed infrastructure. The operator objects that do appear exist as status reporters, not as pod controllers.

This is the definitive proof of the ARO HCP model: a fully functioning, healthy control plane with zero customer-managed pods.

Section 2 — NodePool: Inspect Your Operational Surface

The NodePool is the customer’s primary handle on the cluster. VM count, VM SKU, OpenShift version, and update strategy are all managed at this layer. In ARO HCP, NodePool operations go through the Azure CLI — the az aro hcp subcommand group surfaces the same configuration the HyperShift operator manages internally.

List the Cluster and NodePools

  1. List the ARO HCP cluster in your resource group:

    az aro hcp cluster list --resource-group {aro_hcp_cluster_rg} --output table
  2. List the NodePools for your cluster:

    az aro hcp cluster nodepool list --cluster-name {aro_hcp_cluster_name} --resource-group {aro_hcp_cluster_rg} --output table
  3. Capture the default NodePool name and inspect it:

    NODEPOOL_NAME=$(az aro hcp cluster nodepool list \
      --cluster-name {aro_hcp_cluster_name} \
      --resource-group {aro_hcp_cluster_rg} \
      --query "[0].name" --output tsv)
    echo "NodePool: $NODEPOOL_NAME"
    az aro hcp cluster nodepool show \
      --cluster-name {aro_hcp_cluster_name} \
      --resource-group {aro_hcp_cluster_rg} \
      --name $NODEPOOL_NAME

    Review these fields in the output:

    • properties.replicas — the desired worker node count. Changing this is how you scale the cluster.

    • properties.platform.vmSize — the Azure VM SKU for workers. The customer controls this; Red Hat does not choose it.

    • properties.version.id — the OpenShift version for this NodePool. Workers can be at a different version than the control plane during a rolling upgrade. The control plane version is managed by Red Hat SRE independently of this field.

    • properties.provisioningState — NodePool-level health. A non-Succeeded state is in customer scope to investigate.

    • properties.autoRepair — when true, ARO HCP automatically replaces unhealthy nodes without customer intervention.

    • properties.status.conditions — detailed health conditions. Check these when provisioningState is not Succeeded.

The NodePool in ARO HCP is the customer-facing equivalent of the HyperShift NodePool CRD that lives on the management cluster. The Azure CLI surfaces the same configuration through a managed API without requiring access to the management cluster’s Kubernetes API.

Section 3 — NodePool Provisioning: Adding a Kata-Dedicated NodePool

This section demonstrates the operational model in action: one Azure CLI call drives Azure VM provisioning transparently. You create a second NodePool using the Standard_D4s_v5 SKU — a VM generation with available quota that supports nested virtualization (KVM). This NodePool carries directly into Module 2 as the Kata Containers NodePool, so it is not deleted at the end of this module.

Record the Baseline

  1. Record the current VMs in the managed resource group before adding capacity:

    az vm list \
      --resource-group {aro_hcp_cluster_managed_rg} \
      --output table \
      --query "[].{Name:name, Size:hardwareProfile.vmSize}"

    Note the number and SKU of VMs. After creating the new NodePool, two additional Standard_D4s_v5 VMs will appear here.

Create the Kata NodePool

  1. Capture the OpenShift version from the existing NodePool — the new NodePool must match:

    VERSION=$(az aro hcp cluster nodepool show \
      --cluster-name {aro_hcp_cluster_name} \
      --resource-group {aro_hcp_cluster_rg} \
      --name $NODEPOOL_NAME \
      --query "properties.version.id" --output tsv)
    echo "Version: $VERSION"
  2. Create a two-node Kata NodePool using the Standard_D4s_v5 SKU, which supports nested virtualization (KVM):

    az aro hcp cluster nodepool create \
      --cluster-name {aro_hcp_cluster_name} \
      --resource-group {aro_hcp_cluster_rg} \
      --name kata-nodepool \
      --vm-size Standard_D4s_v5 \
      --replicas 2 \
      --version $VERSION \
      --channel-group stable

Inspect Node Labels While Provisioning

New VM provisioning typically takes 5–10 minutes. Use that time in terminal-2 to examine how Azure VM metadata surfaces as Kubernetes node labels on an existing worker — the OCP-layer evidence of the Azure infrastructure you inspected in Module 0.

  1. Capture an existing worker node name:

    WORKER_NODE=$(oc get nodes -o jsonpath='{.items[0].metadata.name}')
    echo $WORKER_NODE
  2. Inspect the Azure-specific node labels:

    oc describe node $WORKER_NODE | grep -E 'topology|instance-type|node\.kubernetes\.io'

    Note these labels, set automatically by the cloud controller manager when the node registered:

    • node.kubernetes.io/instance-type — the Azure VM SKU, matching properties.platform.vmSize in the NodePool spec. You will also see the legacy alias beta.kubernetes.io/instance-type with the same value.

    • topology.kubernetes.io/region — the Azure region (e.g. eastus2)

    • topology.kubernetes.io/zone — the Azure Availability Zone as a number (0, 1, or 2), not a named zone

    • topology.disk.csi.azure.com/zone — the zone used by the Azure CSI disk driver for persistent volume placement

    • k8s.ovn.org/layer2-topology-version — OVN-Kubernetes networking label; not related to Azure placement

Cross-Reference in Azure

  1. While the NodePool is provisioning, query the managed resource group for new VMs:

    az vm list \
      --resource-group {aro_hcp_cluster_managed_rg} \
      --output table \
      --query "[].{Name:name, Size:hardwareProfile.vmSize, State:provisioningState}"

    Refresh this command every 30–60 seconds. You will see two new Standard_D4s_v5 VMs appear with provisioningState: Creating, then transition to Succeeded.

You made one Azure CLI call — creating a NodePool — and Azure provisioned a VM. You never called az vm create, never allocated a NIC, never configured a disk. The ARO HCP service translated the declarative intent into every required Azure API call. This is the operational model ARO HCP delivers at every scale.

  1. Confirm the new node joins the cluster:

    oc get nodes -w

    Wait for the new node to appear in Ready status, then press Ctrl+C.

  2. Inspect the new nodes' labels and compare with the existing workers:

    NEW_NODE=$(oc get nodes --sort-by=.metadata.creationTimestamp \
      -o jsonpath='{.items[-1].metadata.name}')
    oc describe node $NEW_NODE | grep -E 'topology|instance-type|node\.kubernetes\.io'

    The new node shows Standard_D4s_v5 as the instance-type — different from the existing Standard_D4s_v6 workers. This confirms the two NodePools are backed by different VM generations, each independently configurable.

The kata-nodepool NodePool remains running. In Module 2, you will apply a node label and NoSchedule taint to these nodes and install the Sandboxed Containers Operator to enable Kata workloads on them.

Summary

What You Learned

  • The control plane is invisible from the customer’s context. oc get nodes lists only worker nodes. Control plane namespaces contain no running pods. Cluster operators report Available=True with no visible pods. The API server is responding, but its implementation runs in Red Hat’s managed infrastructure.

  • The NodePool is the customer’s operational surface. Every worker scaling action, VM SKU selection, and NodePool version update goes through the NodePool resource. Everything above it — the hosted control plane — is Red Hat’s responsibility.

  • One Azure CLI call drives VM provisioning. Creating a NodePool was the only action required. The ARO HCP service handled every Azure API call transparently — VM creation, NIC allocation, disk provisioning.

Key Takeaways for Customer Conversations

"How do I scale my cluster?"

Either update the replica count on an existing NodePool with az aro hcp cluster nodepool update --replicas N, or add a new NodePool for a different VM SKU with az aro hcp cluster nodepool create. Either way, that is the entire action — no Azure portal, no az vm create, no capacity reservation to manage separately. The ARO HCP service translates the declarative intent into every required Azure API call.

"What if I need a different VM size for my workers?"

Update the VM size on the NodePool. ARO HCP will provision new nodes with the new SKU and drain the old ones. The control plane is unaffected — it runs independently and is not sized by the customer.

"How do I know when my NodePool is unhealthy?"

Check the NodePool’s provisioning state and conditions via az aro hcp cluster nodepool show. A non-Succeeded provisioning state is in customer scope to investigate — typically a VM quota issue, a subnet routing problem, or a failing bootstrap. Control plane health, by contrast, is Red Hat’s responsibility.

"How do I upgrade my cluster?"

Control plane upgrades are handled by Red Hat SRE transparently — no customer action required and no maintenance window to schedule. Worker NodePool upgrades are the customer’s responsibility: update the OpenShift version on the NodePool and ARO HCP performs a rolling replacement of worker nodes. Because the control plane and workers upgrade independently, you can stage worker upgrades across NodePools without touching the control plane at all.

"Why can’t I see the master nodes?"

They run in a Red Hat-managed Azure service account. You pay for them through the ARO HCP service tier, but you never patch, monitor, or recover them. Your SRE team’s on-call scope ends at the worker tier.