CDL cost management and operational excellence covers about 14% of the exam. Understanding Cloud Billing structure, cost optimization strategies, and operational principles like SRE and DevOps are key.
Resource Hierarchy
Google Cloud organizes resources in a four-level hierarchy, similar to a corporate org chart:
Project is the fundamental billing unit. IAM policies set at a higher level inherit downward through the hierarchy.
Cloud Billing
Billing Accounts
A Cloud Billing account links a payment method (credit card, invoice) to your Google Cloud usage. One billing account can be connected to multiple projects, allowing centralized cost management.
| Relationship | Description | |---|---| | 1:Many | One billing account can cover multiple projects | | Required | Every project must link to exactly one billing account |
Budget Alerts
Budget alerts notify you by email when your spending approaches or exceeds a threshold you define. For example, set alerts at 50%, 90%, and 100% of a $1,000 monthly budget.
Important: Budget alerts do NOT automatically stop resources — a person must take action after receiving the alert. (Automated response requires Cloud Functions integration.)
Cloud Billing Reports
A visual dashboard showing cost trends over time, filterable by service, region, or project. This helps you identify which services are driving costs.
Cost Export
Export detailed billing data to BigQuery for advanced SQL analysis or custom Looker Studio dashboards. Enables complex queries like "which team spent the most on Compute Engine last quarter?"
Labels
Key-value pairs (e.g., , ) attached to resources. Labels appear in Billing Reports, enabling cost breakdown by team, environment, or cost center.
Cost Optimization Methods
| Method | Requirement | Savings | Notes | |---|---|---|---| | Sustained Use Discounts | None — automatic | Up to 30% | Auto-applied based on monthly usage | | Committed Use Discounts | 1 or 3 year commitment | Up to 57% | Cannot cancel mid-term | | Spot VMs | Fault-tolerant workload | Up to 91% | Can be preempted anytime by Google | | Right-sizing | Usage analysis | Varies | Recommender tool provides suggestions |
Sustained Use Discounts: Automatically applied when Compute Engine VMs run a significant portion of the month — no sign-up needed. Committed Use Discounts: Commit to specific resource amounts (vCPU, memory) for 1–3 years. Best for predictable, steady workloads. Spot VMs: Use Google's spare capacity at steep discounts. Suitable for batch processing, data analysis, or rendering — workloads that tolerate interruption. Google Cloud Pricing Calculator: Estimate costs before provisioning resources.
SRE (Site Reliability Engineering)
SRE is Google's operations philosophy: apply software engineering principles to infrastructure and operations. The three key metrics:
| Metric | Meaning | Example | Audience | |---|---|---|---| | SLA (Service Level Agreement) | Customer contract | "99.99% availability, or we refund" | External (customer) | | SLO (Service Level Objective) | Internal target | "We aim for 99.95% availability" | Internal (team) | | SLI (Service Level Indicator) | Actual measurement | "This month we achieved 99.97%" | Observed value |
Order to remember: SLI (actual) → SLO (target) → SLA (contract)
SLO is always set higher than SLA to provide a buffer. Error budgets track remaining "allowed downtime" within an SLO period — teams use remaining budget to decide when it's safe to ship new features.
DevOps
DevOps breaks down the wall between Development and Operations teams, enabling faster and more reliable software delivery.
CI (Continuous Integration): Developers frequently merge code into a shared repo with automated builds and tests. CD (Continuous Delivery/Deployment): Automatically deploy tested code to production. Automation: Replace manual tasks with code to reduce errors. Monitoring: Quickly detect and respond to post-deployment issues.
Cloud Operations Suite
| Service | Role | |---|---| | Cloud Monitoring | Collect metrics, build dashboards, set alerts | | Cloud Logging | Centralize log collection, storage, and search | | Error Reporting | Auto-detect and group application errors | | Cloud Trace | Distributed tracing to find latency bottlenecks | | Cloud Debugger | Inspect variable state in live production code |
High Availability and Disaster Recovery
Multi-Zone deployment: Spread VMs across zones in the same region — one zone fails, others continue. Multi-Region deployment: Highest availability — survives entire region outage. Auto-scaling: Automatically add or remove VMs based on demand. Managed services: Use Cloud SQL, Cloud Spanner, etc. — Google handles HA for you.
Customer Care Tiers
| Tier | Cost | Key Feature | |---|---|---| | Basic | Free | Docs, community forums, billing support | | Standard | Paid | Business-hours technical support | | Enhanced | Paid | 24/7 support, faster response times | | Premium | Paid | Dedicated TAM, highest priority support |
Exam Key Points
"Fundamental billing unit" -- Project
"Budget threshold email, no auto-stop" -- Budget Alerts
"Detailed cost analysis via SQL" -- Cost Export → BigQuery
"Key-value pairs for cost tracking" -- Labels
"Automatic discount based on monthly usage" -- Sustained Use Discounts
"1–3 year commitment, up to 57% off" -- Committed Use Discounts
"Interruptible, up to 91% discount" -- Spot VM
"Customer availability contract" -- SLA
"Internal availability target, higher than SLA" -- SLO
"Actual measured availability" -- SLI
"SLI → SLO → SLA" -- SRE metric order
"Metrics, dashboards, alerts" -- Cloud Monitoring
"Centralized log collection and search" -- Cloud Logging
"Dedicated TAM, highest support tier" -- Premium