When you search for "serverless" on GCP, three services appear simultaneously: App Engine, Cloud Run, and Cloud Functions. All three share the property that you never manage the underlying infrastructure directly, yet their execution models and ideal workloads are sharply different. This guide distills patterns extracted from 53 real GCP-ACE exam questions and maps each service to the scenarios where it wins.
---
Three Services Behind the Word "Serverless"
| Service | Abstraction Level | Primary Target | Billing Basis | |---------|------------------|----------------|---------------| | App Engine | Platform (PaaS) | Full-stack web apps | Instance-hours | | Cloud Run | Container serverless | HTTP microservices | CPU, memory, request count | | Cloud Functions | Function serverless | Short-lived event-driven functions | Invocation count + execution time |
With App Engine you push code and a runtime config, and Google manages the load balancer, HTTPS termination, and autoscaling for you. Cloud Run accepts any Docker container as-is, removing all language or framework constraints. Cloud Functions wires a single function to an event source and represents the lightest serverless footprint. The single most important distinction among the three is what unit you deploy: an application, a container image, or a function.
For platform engineers who routinely think in containers, Cloud Run will feel the most natural. For backend engineers who want to avoid even the concept of an image build, Cloud Functions removes that layer. App Engine sits in the middle — the deployment artifact is your source code plus an , not a container image, yet the abstraction is higher than a raw VM.
---
App Engine Standard vs Flexible
App Engine offers two distinct execution environments. Understanding precisely where they differ is what separates a correct answer from a wrong one on most App Engine exam questions.
| Attribute | Standard Environment | Flexible Environment | |-----------|---------------------|---------------------| | Runtimes | Python, Java, Go, Node.js, Ruby, PHP (pinned versions) | Any language (custom Dockerfile) | | Instance startup | Seconds (sometimes milliseconds) | Several minutes | | Scale to zero | Yes — instances drop to 0 when idle | No — minimum 1 instance always running | | Direct VPC access | No — requires Serverless VPC Access Connector | Yes | | Cost model | Per-second billing, zero cost when idle | Minimum 1 instance billed continuously | | Filesystem | Read-only (local /tmp is writable) | Full read/write |
App Engine Standard's advantages are fast cold starts and the ability to scale to zero. It is cost-efficient for apps with bursty or unpredictable traffic, but the pinned runtime versions can be a constraint if you depend on a library that requires a newer interpreter.
App Engine Flexible lets you bring a custom Docker image and connect directly into a VPC, but the always-on minimum instance means you pay even when no traffic is flowing. The exam heuristic is simple: if the question mentions direct VPC connectivity, the answer is Flexible; if the question mentions zero cost at zero traffic, the answer is Standard.
One critical constraint that appears on the exam: each GCP project can have exactly one App Engine application, and the region you choose at first deployment is permanent. Changing the region requires creating an entirely new project. This is a hard operational limit, not a soft guideline.
---
Cloud Run: The Default Container Serverless
Cloud Run is the most general-purpose serverless option in GCP. It accepts any Docker container that listens on an HTTP port, which means you are free to use any language, any framework, and any dependency stack.
The core mechanism of Cloud Run is request-driven scaling. When no requests are arriving, the service scales down to zero instances. When traffic spikes, Cloud Run automatically expands up to the configured maximum — 1,000 by default.
| Cloud Run Setting | Default | Description | |-------------------|---------|-------------| | Concurrency | 80 | Maximum simultaneous requests handled by one instance | | Minimum instances | 0 | Instances kept warm at all times | | Maximum instances | 1,000 | Scale-out ceiling | | Request timeout | Up to 3,600 seconds | Maximum duration for a single request | | CPU allocation | During request processing only | Use --cpu-always-on to keep CPU allocated between requests |
Cloud Run operates in two modes. A Cloud Run Service receives HTTP traffic continuously and is the right choice for APIs and web applications. A Cloud Run Job has no HTTP endpoint; it runs to completion and exits, making it ideal for batch processing workloads.
For asynchronous processing, Pub/Sub Push subscriptions that invoke Cloud Run are a standard GCP architecture. Google's recommended authentication approach in this pattern is OIDC tokens. API key authentication is not the correct answer on the exam and is not the recommended approach in production either.
For CPU-bound services, set concurrency lower (closer to 1). For I/O-bound services — those waiting on database queries or external API calls — higher concurrency means fewer instances are needed to sustain the same throughput, which directly reduces cost.
---
Cloud Functions and the Event Trigger Model
Cloud Functions deploys code at the granularity of a single function and wires it directly to an event source. If Cloud Run's unit is a container, Cloud Functions' unit is a function. You do not manage routing, do not define ports, and do not think about request lifecycle beyond the function boundary.
| Trigger Type | Example Event | Gen 1 Support | Gen 2 Support | |--------------|---------------|---------------|---------------| | HTTP | Inbound HTTP request | Yes | Yes | | Cloud Storage | Object upload (finalize), deletion, metadata change | Yes | Yes | | Cloud Pub/Sub | Message published to a topic | Yes | Yes | | Firestore | Document created, updated, or deleted | Yes | Yes | | Eventarc (others) | BigQuery, Cloud Audit Logs, 90+ sources | No | Yes |
The most consequential differences between Gen 1 and Gen 2 are the execution time ceiling and concurrency model. Gen 2 runs on top of Cloud Run internally, which is why it inherits Cloud Run's concurrency configuration and can run HTTP-triggered functions for up to 60 minutes.
| Attribute | Cloud Functions Gen 1 | Cloud Functions Gen 2 | |-----------|-----------------------|-----------------------| | Max execution time | 9 minutes | 60 min (HTTP), 9 min (events) | | Underlying infrastructure | Proprietary | Cloud Run | | Concurrency | 1 request per instance (fixed) | Configurable (same as Cloud Run) | | Eventarc integration | No | Yes | | Traffic splitting | No | Yes |
The canonical use case for Cloud Functions is responding to Cloud Storage events. When an image is uploaded to a bucket, the event fires and the connected function executes automatically to perform resizing or metadata extraction. When no events arrive, cost is zero. Billing is based on invocation count plus execution time rounded up to the nearest 100 milliseconds. The free tier includes 2 million invocations and 360,000 GB-seconds per month.
---
Traffic Splitting and Zero-Downtime Deployment Patterns
App Engine's Traffic Splitting feature distributes incoming requests across multiple deployed versions of the same service at configurable percentages. This enables canary deployments, A/B testing, and blue-green switchovers natively, without an external load balancer configuration.
| Split Method | Characteristic | Best Scenario | |--------------|---------------|---------------| | IP-based | Same client IP always routes to the same version | When session consistency matters | | Cookie-based | HTTP cookie pins a user to a version (more stable than IP) | Logged-in user sessions | | Random | Each request independently selects a version | Pure percentage-based testing |
Both setting traffic splits and performing rollbacks are single gcloud commands:
Canary deploy (10% to new version): Immediate rollback:
When you want to deploy a new version without shifting any traffic to it, use the flag. The version is deployed and starts, but receives zero traffic until you explicitly adjust the split. This is the safest deployment pattern for production services with strict uptime requirements.
Cloud Run manages versions at the revision level and also supports traffic splitting across revisions through the same gcloud tooling.
| Service | Version Unit | Traffic Splitting | Instant Rollback | |---------|-------------|-------------------|------------------| | App Engine | Version | Yes (IP/cookie/random) | Yes | | Cloud Run | Revision | Yes | Yes | | Cloud Functions Gen 1 | Not supported | Not supported | Requires redeploy | | Cloud Functions Gen 2 | Supported | Supported | Yes |
If the exam scenario mentions "instant rollback without redeployment," the answer is App Engine traffic splitting. Deployed versions persist until explicitly deleted, so routing 100% of traffic back to a previous version completes the rollback in seconds.
---
The Concurrency, Minimum Instances, and Cold Start Tradeoff
In serverless architecture, cost and response latency move in opposite directions. Every configuration choice that reduces cold start latency adds idle-instance cost, and every choice that eliminates idle cost reintroduces cold start risk. Understanding this tradeoff is essential both for production system design and for answering exam scenarios correctly.
A cold start occurs when the first request arrives at a service that currently has zero running instances. The new instance must be provisioned, the container or runtime must be initialized, and only then can the request be processed. Both App Engine Standard and Cloud Run control this behavior through the minimum instances setting.
| Setting | Effect | Cost Impact | |---------