Choosing the Right Storage and Database Service on GCP

From Cloud Storage's four storage classes to Spanner, Bigtable, and BigQuery — a practical breakdown of the data service selection criteria that trip up engineers on the GCP-ACE exam.

One of the first moments of overwhelm for engineers new to GCP is realizing how many data services exist — and how similar they all sound on the surface. Cloud Storage, Cloud SQL, Spanner, Firestore, Bigtable, BigQuery, Memorystore. Each has a distinct design purpose, and picking the wrong one can mean a failed architecture review or a missed exam question. The goal of this guide is not to hand you a memorization list but to build pattern recognition: when you spot certain keywords in a scenario, the right service should come to mind automatically.

---

 

Cloud Storage: Four Storage Classes and Lifecycle Automation

Cloud Storage is GCP's foundational object storage service. It stores unstructured data — images, videos, log files, pipeline source data, backups — at unlimited scale, with no server management and no capacity pre-provisioning required. You pay for what you use, and the cost varies significantly based on how frequently you need to access your data.

The most important design decision in Cloud Storage is choosing the right storage class. The pricing model rewards infrequent access with lower per-GB storage costs, but charges retrieval fees when you pull data out.

| Class | Primary Use | Minimum Duration | Storage Cost | Retrieval Fee | |-------|------------|------------------|--------------|---------------| | Standard | Hot data, frequent access | None | Highest (~$0.020/GB/mo) | None | | Nearline | Less than once a month | 30 days | ~$0.010/GB/mo | Yes | | Coldline | Less than once a quarter | 90 days | ~$0.004/GB/mo | Yes | | Archive | Once a year or less, long-term retention | 365 days | Lowest (~$0.0012/GB/mo) | Highest |

On the exam, access frequency and retention requirements always appear together. If the scenario says "retain for 7 years, accessed only in the event of a legal dispute," Archive is the answer. If the scenario describes quarterly disaster recovery drills requiring backup restoration, Coldline is the right fit. Violating the minimum storage duration incurs an early deletion charge.

Lifecycle policies let you automate class transitions without manual intervention. A single rule that reads "transition to Nearline after 30 days, Coldline after 90 days, Archive after 365 days" handles cost optimization automatically once it is configured. This is exactly the kind of operational efficiency pattern the exam rewards.

---

 

Storage Types Compared: Object vs. Block vs. File

Cloud Storage is object storage, but GCP also offers block storage and file storage. These three categories represent fundamentally different storage architectures and serve completely different workloads.

| Storage Type | GCP Service | Access Method | Primary Use Case | |-------------|------------|---------------|------------------| | Object Storage | Cloud Storage | HTTP/HTTPS (REST API) | Images, video, backups, pipeline source data | | Block Storage | Persistent Disk | Block-level (mounted to VM) | VM OS disk, database servers | | File Storage | Filestore | NFS (shared mount across VMs) | Shared file systems, legacy NAS migrations |

Persistent Disk attaches to Compute Engine VMs and behaves like a traditional disk drive. Regional Persistent Disk replicates synchronously across two zones within the same region, enabling failover without data loss in the event of a zone outage. This is the recommended option for stateful workloads that need zone-level resilience.

Filestore provides managed NFS shares that multiple VMs can mount simultaneously. When you are migrating on-premises NAS workloads to GCP and the application expects a standard file system interface, Filestore is the natural landing zone. It eliminates the complexity of running your own NFS server on a VM.

!Object versus Block versus File storage

Cloud SQL: Where Managed Relational Databases Fit

Cloud SQL is GCP's fully managed relational database service, supporting MySQL, PostgreSQL, and SQL Server. Because it maintains wire-protocol compatibility with each engine, migrating an on-premises database usually requires nothing more than updating the connection string. You get the same SQL dialect, the same drivers, and the same query behavior — but without the operational overhead of managing the underlying infrastructure.

| Feature | Cloud SQL | |---------|-----------| | Supported Engines | MySQL 8.0, PostgreSQL 15, SQL Server 2019 | | High Availability | Automatic Primary + Standby in a different zone, same region | | Automatic Failover | Standby promoted within ~60 seconds of zone failure | | Read Replicas | Same-region and cross-region replicas for read scaling | | PITR | Point-in-time recovery up to 7 days back | | Maximum Storage | 64 TB |

Cloud SQL's architectural constraint is its single-instance write path. Read replicas can distribute read traffic, but all writes funnel through one Primary instance. When your dataset grows past tens of terabytes or write throughput approaches single-instance limits, that is the signal to evaluate Cloud Spanner. Cloud SQL is the right tool for conventional web applications, internal business systems, and any workload that fits comfortably within a single region and a single primary.

---

 

Spanner: A New Category Called Global Strong Consistency

Cloud Spanner occupies a category that did not exist before Google built it: a relational database that supports full SQL and ACID transactions while scaling horizontally across multiple regions. No other GCP service combines those three properties simultaneously.

| Property | Cloud SQL | Cloud Spanner | |----------|-----------|---------------| | Scaling Model | Vertical (resize instance) | Horizontal (add nodes, automatic sharding) | | Maximum Scale | ~64 TB | Unlimited (multi-petabyte) | | Regional Scope | Single region | Single region or multi-region | | Availability SLA | 99.95% (HA config) | 99.999% (multi-region) | | Write Scaling | Single Primary | Distributed across nodes automatically | | Cost | Low | High (per-node pricing) |

The technology that makes Spanner possible is TrueTime. Google uses GPS receivers and atomic clocks to synchronize time across globally distributed nodes, enabling external consistency — a guarantee stronger than serializable isolation — even across geographically separated regions. Replication uses Paxos consensus, and leader election is automatic when a region fails. This gives Spanner an RPO of zero and an RTO measured in single-digit seconds.

Exam pattern recognition: when a single scenario sentence contains the words "global," "strong consistency," "ACID transactions," "region failure tolerance," and "horizontal write scaling" together, Cloud Spanner multi-region is the answer. Any service that handles only one or two of those requirements is a distractor.

---

 

Firestore and Bigtable: Two Faces of NoSQL

GCP's NoSQL landscape is split between two services with very different design philosophies. Firestore is a document-oriented NoSQL database. Bigtable is a wide-column NoSQL database. Choosing between them depends entirely on workload characteristics, not on a preference for one NoSQL style over another.

| Property | Firestore | Bigtable | |----------|-----------|----------| | Data Model | Collections → Documents (JSON-like) | Row key + column families (wide-column) | | Throughput | Serverless auto-scaling | Millions of operations per second, millisecond latency | | Optimal Workload | Mobile/web apps, user profiles | IoT time-series, gaming, ad tech, analytics pipelines | | Offline Support | Yes (automatic sync) | No | | Transactions | Single-document atomicity | Single-row atomicity only | | HBase Compatible | No | Yes |

Firestore includes real-time listeners and offline synchronization baked into its client SDKs. This makes it the natural choice for mobile application backends where the client needs to stay in sync with server state even during network interruptions. Firestore Native mode supports mobile SDKs and real-time updates; Datastore mode exists for teams migrating applications from the older Cloud Datastore API.

Cloud Bigtable performance is almost entirely determined by row key design. For IoT time-series data, the recommended pattern is — placing the most recent data at the top of the row range makes lookups fast. If you use a forward timestamp, every write lands at the bottom of the table, creating a write hotspot that degrades throughput across the entire cluster.

---

 

BigQuery: The Standard for Analytical Workloads

BigQuery is GCP's serverless data warehouse. You run SQL queries against petabytes of data without provisioning any clusters, managing any infrastructure, or pre-warming any caches. The architecture fully decouples storage from compute, which is what makes the on-demand pricing model possible — you pay for storage continuously, and you pay for compute only when a query runs.

| Property | Details | |----------|---------| | Architecture | Storage and compute fully decoupled, serverless | | Data Format | Columnar storage — optimized for aggregation queries | | Pricing Model | On-demand: ~$5 per TB scanned / Flat-rate: reserved slots | | Processing Scale | Query petabyte-scale datasets in seconds to minutes | | External Tables | Query JSON/CSV/Parquet files in Cloud Storage directly |

BigQuery is designed for OLAP workloads — batch analytical queries that aggregate hundreds of millions or billions of rows. High-frequency small transactions belong to Cloud SQL or Spanner. The flag lets you estimate how many bytes a query will scan before you execute it, which is the standard cost governance pattern for large-scale analysis. Running a dry run on an unfamiliar query before committing compute spend is a practical habit the exam scenario may describe as a requirement.

---

 

Scenario-Based Decision Patterns That Commonly Appear on the Exam

Data service questions on the GCP-ACE exam are resolved by keyword combinations. Train yourself to recognize th

Back to blog list