Every engineer running cloud infrastructure eventually faces the same question: is my system actually healthy right now? Answering that question requires a system for collecting and analyzing three types of signals — metrics, logs, and traces — which together form what the industry calls an observability stack. On Google Cloud, Cloud Monitoring and Cloud Logging sit at the center of that stack.
---
What It Means to Run Cloud Without Observability
In an environment where dozens of VMs and managed services are tightly coupled, walking over to a server rack to check status simply is not an option. When something breaks, not knowing where it broke, when it started, or why it happened is called black-box operations. Observability is the ability to infer what is happening inside a system from signals observable from the outside.
The three pillars of observability map directly to GCP tools:
| Signal | What It Tells You | GCP Tool | |--------|-------------------|----------| | Metrics | Numeric system state — CPU at 80%, request latency 200ms | Cloud Monitoring | | Logs | What happened at a specific moment — error messages, audit events | Cloud Logging | | Traces | How a request flowed across services — pinpoint bottlenecks | Cloud Trace |
Metrics answer "what is the current state," logs answer "what happened," and traces answer "why is it slow." Exam scenarios frequently require distinguishing which signal type — and which tool — applies to a given situation. Getting these three locked in early pays dividends on every operations scenario you encounter.
---
Core Components of Cloud Monitoring
Cloud Monitoring was formerly known as Stackdriver Monitoring. It ingests thousands of metrics automatically from GCP resources, stores them in a time-series database, and provides dashboards and Alerting Policies so operations teams can assess system health at a glance.
Metric Collection Architecture
For Compute Engine VMs, installing the Ops Agent enables collection of system-level metrics — CPU, memory, disk, network — alongside application logs from the same host. Managed services such as Cloud SQL, GKE, and Cloud Run emit metrics automatically with no agent required. Custom application metrics can be sent via the OpenTelemetry SDK or Prometheus. In GKE environments, Managed Service for Prometheus ingests Prometheus metrics natively without requiring changes to existing instrumentation.
Dashboards and Alerting Policy
The two headline features of Cloud Monitoring are dashboards for metric visualization and Alerting Policy for automated notification. An Alerting Policy has three components: a condition that defines which metric and threshold triggers the alert, a notification channel that specifies who gets notified and how, and documentation that provides runbook guidance alongside the alert notification.
| Component | Role | Example | |-----------|------|---------| | Condition | Which metric, which threshold, for how long | disk/read_latencies > 100ms for 5 min | | Notification Channel | Who gets notified and how | Email, Slack, PagerDuty, Webhook, Pub/Sub | | Duration | Ignore transient spikes, trigger only on sustained violations | 1 min, 5 min, 10 min | | Documentation | Runbook or guidance message attached to the alert | "If this fires, check X first" |
The duration window is critical for reducing alert fatigue. Without it, a CPU spike that resolves in seconds still fires a page. Setting a five-minute duration filters out transient noise and ensures that only genuine, sustained problems generate notifications. This single configuration detail is the difference between an on-call rotation that trusts its alerts and one that ignores them.
---
How Cloud Logging and Log Router Work
Cloud Logging is GCP's centralized log management service. Two audit log categories are especially important for the exam. Admin Activity audit logs record resource creation, deletion, and modification — including IAM changes and schema alterations — automatically and cannot be disabled. Data Access audit logs record data-plane operations such as BigQuery queries and Cloud Storage reads. They are disabled by default and must be explicitly enabled per service.
Log Router and Sinks
The architectural backbone of Cloud Logging is the Log Router and its sinks. The Log Router inspects every log entry as it arrives and evaluates it against all configured sink filters. Entries that match a filter are exported to the sink's destination. The flow is: log source → Cloud Logging ingestion → Log Router → sink filter matching → destination.
Choosing the correct sink destination is the most frequently tested Log Router concept. The table below maps scenario keywords to the right answer:
| Sink Destination | Use Case | Keyword Hints | |-----------------|----------|---------------| | Cloud Storage | Long-term retention, compliance archiving, audit log preservation | "long-term," "archive," "compliance" | | BigQuery | Log-based analytics, SQL queries, BI dashboard integration | "analyze," "SQL query," "BigQuery" | | Pub/Sub | Real-time processing, external system integration, streaming pipelines | "real-time," "external system," "streaming" | | Splunk | Integration with existing SIEM tooling | "Splunk," "SIEM" | | Another Cloud Logging | Centralize logs across organizations, route audit logs to a security project | "centralize," "another project" |
---
Log-based Metrics and SLO-based Alerting
Log-based Metrics
One of Cloud Logging's most powerful features is the ability to derive metrics directly from log data. A counter metric that counts log entries containing the string "ERROR" becomes a time-series that Cloud Monitoring can alert on, just like any infrastructure metric. Two metric types are available: counter metrics (count of log entries matching a filter) and distribution metrics (statistical distribution of a numeric value extracted from log entries). When an exam scenario describes alerting when a specific log pattern exceeds a threshold, the answer is Log-based Metrics combined with an Alerting Policy.
SLO-based Alerting
A Service Level Objective defines a reliability target for a service. Cloud Monitoring can define SLOs directly and alert based on error budget burn rate rather than raw metric thresholds. For a monthly 99.9% availability SLO, an alert fires immediately when 10% of the error budget burns in one hour, and routes through a slower channel when 1% burns over 24 hours. This is the SRE-idiomatic approach to alerting — it surfaces real user impact rather than transient infrastructure noise.
| Alerting Approach | Characteristics | Best Used For | |-------------------|-----------------|---------------| | Threshold-based | Fires when a metric exceeds a numeric value | Infrastructure metrics requiring immediate response | | Log-based Metric Alert | Fires when a log pattern frequency exceeds a threshold | Error log spikes, security event detection | | SLO-based Alert | Fires based on error budget burn rate | SLO-managed services, minimizing alert fatigue |
---
Measuring Availability with Uptime Check
Uptime Check is Cloud Monitoring's mechanism for probing service availability from the outside in. Cloud Monitoring sends HTTP, HTTPS, or TCP requests from multiple global regions to a specified URL or IP, then inspects the response code, response body, and response time.
| Component | Description | |-----------|-------------| | Target | URL, IP address, or GCP resource endpoint (e.g., https://example.com/health) | | Protocol | HTTP, HTTPS, TCP | | Check Interval | 1 to 15 minutes (1 minute recommended) | | Content Matcher | Response body must contain a specific string (e.g., "OK", "healthy") to pass | | Check Locations | Multiple regions simultaneously — us-east1, europe-west1, asia-east1, and more |
When an Uptime Check fails, it integrates with Alerting Policy to send immediate notifications. This is the fastest way to detect that real users cannot reach a service. One constraint: VMs with internal-only IPs cannot be checked by default. Uptime Check applies only to externally reachable endpoints.
---
Cloud Trace, Error Reporting, and Cloud Profiler — One-Line Summaries
Beyond Cloud Monitoring and Cloud Logging, the GCP observability stack includes three additional tools. Understanding each tool's distinct role prevents confusion on scenario-based exam questions.
| Service | Role | When to Use | |---------|------|-------------| | Cloud Trace | Distributed tracing — follows a request as it moves across services | High API latency, identifying the bottleneck service | | Error Reporting | Automatically aggregates and groups errors, alerts on new error types | Quickly detecting error spikes after a deployment | | Cloud Profiler | Flame graph analysis of CPU and memory consumption in production | Performance optimization, memory leak detection |
Cloud Trace is instrumented via the OpenTelemetry SDK or Cloud Trace API directly. GKE and Cloud Run offer automatic trace collection options that provide baseline latency data without any code changes. Error Reporting extracts and groups error-level log entries automatically from Cloud Logging and fires an alert the first time a new error type appears. Cloud Profiler uses a lightweight agent with negligible overhead, making it safe to run continuously in production environments.
!Cloud Trace versus Error Reporting versus Cloud Profiler
Monitoring and Logging Scenarios That Trip Up Exam Takers
This section consolidates the scenario types that appear most frequently on the exam and cause the most confusion.
Scenario 1: Choosing the Right Sink Destination
This is the most common Log Router question. Match the keyword in the scenario to the correct sink before evaluating answer choices.
| Keyword | Correct Sink | Common Wrong Answer | |---------|-------------|---------------------| | Long-term retention, compliance, archive, reta