Spinning Up a VM Is Never As Simple As It Looks
Launching a VM instance in the Google Cloud Console takes less than a minute. But behind that clean interface lies a maze of decisions: machine type, boot disk, network tags, service account, startup script, availability policy, and maintenance behavior — dozens of choices that shape the runtime characteristics of your instance. In a development environment, the defaults are usually fine. In production, every one of those choices directly determines cost, security, and availability.
The GCP-ACE exam does not ask you how to click through the VM creation wizard. It asks which combination of options is most appropriate for a given set of constraints. This guide walks through the core Compute Engine concepts you need to know — from the perspective of an engineer who actually operates these resources.
---
Instance Creation Options and Machine Type Selection
The first decision when creating a VM is the machine type. Compute Engine offers both predefined machine types and custom machine types.
| Machine Family | Purpose | Characteristics | |---------------|---------|----------------| | E2 | General purpose (cost-optimized) | Spot/preemptible workloads, low-cost web servers, dev environments | | N2/N2D | General purpose (balanced) | Mid-scale applications, databases | | C2/C3 | Compute-optimized | HPC, game servers, high-throughput processing | | M1/M2/M3 | Memory-optimized | In-memory databases, SAP HANA, large-scale analytics | | A2/A3 | Accelerator-optimized | ML training, GPU-intensive workloads | | T2D | Tau (scale-out efficiency) | Scale-out workloads, web serving |
Custom machine types let you independently adjust vCPU count and memory, allowing cost optimization for workloads that don't fit neatly into predefined families.
The table below summarizes the most commonly used instance creation options and their operational implications.
| Option | Description | Gotcha | |--------|-------------|--------| | Network tags | Target VM for specific firewall rules | Typos in tags silently prevent firewall rules from applying | | Service account | Grants the VM permission to call GCP APIs | Avoid using the default compute service account in production | | Startup script | Script executed automatically at boot | Can be stored in Cloud Storage and referenced via | | Availability policy | Controls preemptibility and restart behavior | Spot VMs cannot be set to always restart | | Maintenance behavior | Live Migration or Terminate during host maintenance | GPU instances do not support Live Migration |
Startup scripts are specified either inline via the metadata key or by pointing to a Cloud Storage path with . From inside the instance, the metadata server exposes runtime values like instance ID, zone, and project ID — useful for self-configuration without hardcoding.
---
Persistent Disk and Snapshots — The Data Lifecycle
Storage in Compute Engine is managed independently from the VM lifecycle. When you delete a VM, attached Persistent Disks are retained by default — but the boot disk is deleted along with the VM. This asymmetry is one of the most reliable exam traps.
Persistent Disk Types
| Disk Type | Max IOPS | Throughput | Cost | Primary Use Case | |-----------|----------|------------|------|------------------| | Standard HDD (pd-standard) | Low | Low | Lowest | Cold data, backups | | Balanced (pd-balanced) | Medium | Medium | Medium | General-purpose applications | | SSD (pd-ssd) | High | High | High | Databases, high-performance workloads | | Extreme (pd-extreme) | Very high | Very high | Highest | High-performance DBs, real-time analytics |
Extreme disks use provisioned IOPS — you specify the exact IOPS you need. They are overkill for typical web apps but appropriate for large databases like SAP or Oracle.
Snapshots vs Machine Images vs Custom Images
These three backup and replication mechanisms are frequently confused. Know the distinction cold.
| Mechanism | Scope | Location | Primary Use | |-----------|-------|----------|-------------| | Snapshot | Single disk | Global | Incremental backup, disk replication | | Machine image | All disks + VM config | Global | Full VM clone, disaster recovery | | Custom image | Boot disk OS state | Global | Standard base images, template-based provisioning |
Exam trap: Snapshots are global resources — they are not bound to a region. Persistent Disks themselves, however, are zonal resources. To move a disk across regions, you create a snapshot and then restore it into a new disk in the target region. The boot disk is deleted by default when a VM is deleted, so if data persistence matters, you must explicitly disable the "Delete boot disk when instance is deleted" option.
!Snapshot versus Machine Image versus Custom Image
Instance Templates and Managed Instance Groups (MIG)
The foundation of scalability and high availability in Compute Engine is the Managed Instance Group (MIG). A MIG manages a fleet of identical VMs based on an instance template. The template is an immutable blueprint that defines machine type, boot disk, network tags, service account, and other configuration — it cannot be modified in place; changes require creating a new template version.
Zonal MIG vs Regional MIG
| Attribute | Zonal MIG | Regional MIG | |-----------|-----------|-------------| | Instance placement | Single zone | Distributed across up to 3 zones in a region | | Availability | Full outage if zone fails | Continues serving if one zone fails | | SLA | Lower | 99.99% (multi-zone) | | Resource exhaustion | Cannot scale if zone runs out | Can draw from other zones | | Best for | Stateful, single-zone requirements | Stateless, high-availability requirements |
MIGs expose three core capabilities: autoscaling, Autohealing, and rolling updates.
Autoscaling adjusts VM count based on CPU utilization, HTTP load, Pub/Sub queue depth, or custom Cloud Monitoring metrics. Setting minimum and maximum instance counts prevents runaway cost during traffic spikes.
Autohealing works by continuously checking a health endpoint. If a VM fails HTTP, HTTPS, or TCP health checks beyond a threshold, the MIG automatically deletes and replaces it with a fresh instance. Do not confuse Cloud Monitoring alerts with Autohealing. Alerts notify humans; Autohealing takes automated remediation action. Any exam question mentioning automatic VM recovery without manual intervention points to MIG + Autohealing.
Rolling updates replace VMs incrementally with a new instance template. controls how many additional instances can be created simultaneously during the update. controls how many instances can be taken offline at once. For canary deployments, you can update only a subset of instances to the new template, keeping the rest on the old version.
---
Spot VM and Cost Optimization Strategies
Spot VM (formerly Preemptible VM) is Compute Engine's way of offering Google's excess compute capacity at discounts of up to 60–91% off on-demand pricing. The catch: Google can reclaim the instance at any time with only 30 seconds of warning.
| Attribute | Spot VM | Standard VM | |-----------|---------|-------------| | Price | Up to 60–91% off | On-demand rate | | Preemptible | Yes (30-second warning) | No | | Max runtime | Unlimited (but can be interrupted at any time) | Unlimited | | Automatic restart | Not available | Available | | Provisioning model | or | Default |
Spot VMs are well-suited for workloads that can tolerate interruption: batch processing, data pipelines, CI/CD builds, and ML training jobs. They are not appropriate for transactional systems or services that must maintain continuous availability.
A common pattern is to mix Spot VMs and standard VMs within a single MIG. Spot VMs handle burst capacity while standard VMs maintain a baseline, so preemption of Spot instances does not bring the service down entirely.
Committed use discounts (CUDs) are another lever for cost optimization. By committing to a specific amount of vCPU and memory for one or three years, you receive discounts of up to 57%. CUDs apply at the project level within a region — not to specific VMs — so any VM running in that region automatically benefits from the commitment as long as the resource usage stays within the committed amount.
---
OS Login and the Fundamentals of Instance Security
There are two ways to SSH into a Compute Engine VM: the SSH key metadata approach and OS Login.
| Attribute | SSH Key Metadata | OS Login | |-----------|-----------------|----------| | Key management | Manual (register public keys in metadata) | Automatic via Google Cloud IAM | | Audit trail | Limited | Fully integrated with Cloud Audit Logs | | User account | Creates a local Linux account | Google account maps to a Linux account | | Org policy integration | Not supported | Supported (IAM role-based) | | sudo access | Requires separate configuration | role | | Key revocation | Manual deletion from metadata | Instant via IAM access removal |
OS Login is enabled by setting metadata at the instance or project level. After that, users need either for standard SSH access or for sudo-level access. For organizations managing large fleets, OS Login is dramatically simpler: when an employee leaves, disabling their Google account immediately revokes access to every VM — no need to hunt down and manually remove SSH keys from hundreds of instances.
Shielded VM and Confidential VM
| Type | What It Protects | Key Features | |------|-----------------|-------------| | Shielded VM | Boot process integrity | Secure Boot, vTPM, Integrity Monitoring | | Confidential VM | In-memory data | AMD SEV-based memory encryption |
Shielded VM uses Secure Boot to block rootkits and bootkits from loading during the boot process. Confidential VM encrypts data in memory while the VM is running, ensuring that even the hypervisor cannot read the contents. Both are common in regulated industri