What Makes GKE Different from Other Managed Kubernetes Services
Google Kubernetes Engine (GKE) sits alongside AWS EKS and Azure AKS as one of the major managed Kubernetes offerings in public cloud. What sets GKE apart is a simple but consequential fact: Google created the Kubernetes project itself, and GKE has accumulated more production operations experience on top of it than any other service.
With GKE, Google manages the control plane — the API server, etcd, and the scheduler. In a Regional cluster, the control plane is distributed across three zones, so the Kubernetes API keeps responding even when an entire zone goes down. GKE Autopilot takes this further by having Google manage the nodes as well, dramatically reducing the operational burden on engineering teams. The GCP-ACE exam focuses heavily on the differences between these operating models, scaling strategies, and security integration patterns.
---
GKE Standard vs GKE Autopilot: How to Choose
The first major decision any team faces when adopting GKE is whether to use GKE Standard or GKE Autopilot. Both modes expose the same Kubernetes API, but they differ fundamentally in who owns the node infrastructure.
| Feature | GKE Standard | GKE Autopilot | |---------|-------------|---------------| | Node provisioning | Operator managed | Google managed | | Node OS patching | Operator responsibility (auto-upgrade available) | Google responsibility | | Node Pool configuration | Fully customizable | Not available (use per-Pod resource requests instead) | | Billing basis | Node VM uptime | Pod CPU/memory requests | | Privileged containers | Allowed | Prohibited | | Custom DaemonSets | Allowed | Restricted (system DaemonSets only) | | Operational overhead | High | Low | | Spot nodes | Configured per Node Pool | Configured via Pod annotations |
GKE Autopilot is well-suited for teams with limited Kubernetes operations experience and for early-stage startups where shipping velocity matters more than infrastructure control. If your workloads require privileged containers or custom DaemonSets, GKE Standard is the only viable path. One trap that surfaces frequently on the exam: GKE Autopilot requires every Pod to declare explicit resource requests. Without requests, GKE Autopilot will refuse to schedule the Pod.
!GKE Standard versus Autopilot
Cluster Topology: Zonal vs Regional, Public vs Private
GKE clusters are classified along two independent axes: availability scope (Zonal vs Regional) and network exposure level (Public vs Private).
| Cluster Type | Control Plane Location | Node Placement | Key Characteristic | |-------------|----------------------|----------------|--------------------| | Zonal | Single zone | Single zone | Lower cost, full outage if zone fails | | Regional | Spread across 3 zones | Spread across 3 zones | High availability, control plane SLA 99.95% | | Public (default) | Public endpoint | Public IP assigned | Internet access allowed | | Private | VPC-internal endpoint | No external IP | Internet isolated, Cloud NAT required |
With a Regional cluster, the control plane is spread across three zones. A single-zone failure leaves kubectl and Pod scheduling fully operational. A Zonal cluster's control plane becomes unresponsive when its zone fails, creating a full operations outage. For production workloads, Regional clusters are the recommended baseline.
Private clusters isolate both the nodes and the control plane inside a VPC. Containers that need outbound internet connectivity require Cloud NAT. To restrict control plane access to internal VPC traffic only, enable the Private Endpoint and use Authorized Networks to lock down which CIDRs can reach the API server.
VPC-native clusters and Alias IP also appear on the exam. In VPC-native mode, Pod IPs are drawn directly from a secondary IP range on the VPC subnet, enabling direct routing between Pod IPs and other VPC resources. Unlike the legacy routes-based cluster model, VPC-level firewall rules can be applied directly to Pod IPs, giving you finer-grained network policy control.
---
Node Pools and Workload Isolation Strategies
A Node Pool is a group of nodes within a cluster that all share the same configuration. Running multiple Node Pools in a single GKE cluster is a common production pattern for isolating different workload types.
| Use Pattern | Node Pool Configuration Example | Purpose | |------------|--------------------------------|--------| | GPU workload isolation | Separate Node Pool with GPU-enabled n1-standard nodes | Pin ML training Pods exclusively to GPU nodes | | Spot cost reduction | Spot Node Pool alongside On-demand Node Pool | Use Spot for batch jobs that tolerate interruption | | High-memory workloads | Separate memory-optimized instance Node Pool | Isolate in-memory databases and cache servers | | Windows containers | Add Windows Server Node Pool | Support Windows-based applications | | System component isolation | Small dedicated Node Pool | Run only kube-system components |
To pin a Pod to a specific Node Pool, use nodeSelector or nodeAffinity. Taints and Tolerations are another common mechanism. By tainting a GPU Node Pool, you prevent regular Pods without the matching Toleration from consuming GPU resources.
One critical Node Pool characteristic to internalize: GKE Node Pools are immutable. You cannot change the machine type of an existing Node Pool in place. The standard procedure for changing machine types is to create a new Node Pool with the desired configuration, cordon the old Node Pool to block new Pod scheduling, drain it to gracefully evict and reschedule existing Pods, and then delete the old pool. Setting a PodDisruptionBudget during this process guarantees a minimum number of available replicas throughout the migration.
---
The Scaling Trio: HPA, VPA, and Cluster Autoscaler
Scaling in GKE operates at three independent layers, each addressing a different dimension of the problem.
| Scaler | What It Adjusts | Key Metric | Primary Use Case | |--------|----------------|------------|------------------| | HPA (Horizontal Pod Autoscaler) | Pod replica count | CPU, memory, custom metrics | Handling traffic spikes | | VPA (Vertical Pod Autoscaler) | Pod CPU/memory requests and limits | Historical usage data | Automatically right-sizing resource requests | | Cluster Autoscaler | Node count | Unschedulable Pods / idle nodes | Reducing cost during off-peak hours |
HPA scales the number of running Pod replicas up or down. The default metric is CPU utilization, but you can also drive HPA from Stackdriver-based custom metrics or external metrics. The minimum replica count for HPA is 1 — scaling to zero requires KEDA (Kubernetes Event-driven Autoscaling).
VPA analyzes a Pod's actual CPU and memory consumption patterns and either recommends updated request values or applies them automatically. Three modes exist: Off (recommendations only), Auto (applies updates by restarting Pods), and Initial (applies only at first scheduling). Avoid running VPA and HPA on the same Pod when HPA is scaling on CPU — the two controllers conflict. Reserve VPA for workloads where HPA is not using CPU as its scaling signal.
Cluster Autoscaler operates at the Node Pool level, with configurable minimum and maximum node counts. When Pods cannot be scheduled due to insufficient resources, Cluster Autoscaler provisions additional nodes. When nodes sit idle, it migrates the Pods and removes the idle nodes. Combining HPA with Cluster Autoscaler creates a two-layer automation loop: HPA adjusts replica count, and Cluster Autoscaler adjusts node count to match.
The most common confusion on the exam breaks down clearly: automatic node count adjustment uses Cluster Autoscaler; automatic Pod count adjustment uses HPA; automatic right-sizing of Pod resource requests uses VPA. Keep these three roles distinct.
---
Building Secure Identity Integration with Workload Identity
Applications running in GKE frequently need to access GCP services such as Cloud Storage, BigQuery, and Cloud SQL. How you configure authentication for that access is a central security question.
| Authentication Method | Description | Security Risk | |----------------------|-------------|---------------| | Service account key file mount | Mount a JSON key file into a Pod as a Kubernetes Secret | Key exposure compromises all permissions; key rotation is a management burden | | Shared node service account | All Pods on a node inherit the node's service account permissions | Violates least-privilege; grants excessive permissions cluster-wide | | Workload Identity | Bind a Kubernetes Service Account (KSA) to a GCP IAM Service Account (GSA) | No key files; per-Pod least-privilege access |
Workload Identity links a Kubernetes Service Account to a GCP IAM Service Account. When a Pod calls a GCP API, the GKE metadata server automatically provides credentials belonging to the GSA bound to that Pod's KSA. Because no key file is ever created or stored, there is no file to leak and no rotation lifecycle to manage.
The setup involves three steps. First, enable Workload Identity on the cluster. Second, grant the KSA the roles/iam.workloadIdentityUser role on the target GSA in GCP IAM. Third, annotate the Kubernetes Service Account with the iam.gke.io/gcp-service-account annotation pointing at the GSA.
On the exam, Workload Identity problems typically appear as: "What is the most secure way for a GKE Pod to access a GCP service?" Mounting a service account key file is presented as the insecure distractor, and Workload Identity is the correct answer — it is Google's officially recommended approach.
---
GKE Operational Scenarios That Frequently Cause Confusion on the Exam
Here is a consolidated reference for the scenarios where candidates most often choose the wrong answer.
| Scenario | Correct Answer | Common Wrong Answer | |---------|---------------|--------------------| | Auto-adjust