Performance and Cost Optimization

A practical guide to Global Accelerator vs CloudFront, CloudFront Origin Shield, Placement Groups, Spot Fleet, Compute Optimizer, and S3 Intelligent-Tiering for SAP-C02 D3 performance and cost optimization.

Performance improvement and cost optimization appear together in SAP-C02 Domain D3. The goal is not simply to make things faster but to evaluate your ability to make architecture decisions that achieve optimal performance per dollar.

Two core questions anchor this domain: "Which service eliminates the performance bottleneck in this scenario?" and "Which cost optimization option best fits this workload pattern?"

 

Global Accelerator vs CloudFront — When to Use Which

Both services improve global performance, but they work in fundamentally different ways. This distinction is one of the most frequently tested topics in the exam.

AWS Global Accelerator is an Anycast IP-based TCP/UDP transport accelerator. It uses 2 fixed Anycast IPv4 addresses. User traffic enters the nearest AWS edge location and travels to its destination over AWS's private global network instead of the public internet, which often improves latency by 60% or more.

CloudFront is an HTTP/HTTPS caching CDN. It caches static and dynamic content at 400+ edge locations to serve users from nearby points of presence. It operates at the HTTP layer and uses DNS for routing.

| Criterion | Global Accelerator | CloudFront | |-----------|-------------------|-----------| | Protocol | TCP/UDP (all protocols) | HTTP/HTTPS | | Caching | None (routing only) | Yes (content caching) | | IP Address | 2 fixed Anycast IPs | Dynamic IP (domain-based) | | DNS caching issue | None (BGP routing) | Possible TTL delay during cutover | | Best fit | Gaming, IoT, VoIP, non-HTTP apps | Websites, APIs, media streaming | | Traffic dial | Instant cross-region weight adjustment | Not available |

When you need DNS caching bypass for blue/green deployments, TCP-based applications, or fixed IPs for allowlisting, Global Accelerator is the answer.

 

CloudFront Deep Dive — Origin Shield, Lambda@Edge, CloudFront Functions

CloudFront Origin Shield places an additional caching layer between the CloudFront distribution and the origin. Instead of multiple edge locations hitting the origin directly, requests are funneled through Origin Shield, dramatically reducing origin load. This is especially valuable when the origin is an on-premises server or a high-cost API.

Lambda@Edge and CloudFront Functions are two ways to run code at the edge.

Lambda@Edge runs at all 4 CloudFront event stages (Viewer Request, Origin Request, Origin Response, Viewer Response). It supports Node.js and Python with execution times up to 30 seconds, enabling complex logic. Common exam patterns include using User-Agent header analysis to serve different content for mobile vs desktop, and querying DynamoDB to perform dynamic redirects.

CloudFront Functions runs only at Viewer Request and Viewer Response stages with execution time under 1 millisecond. It suits lightweight tasks like URL normalization, query parameter sorting, and header injection, and is much cheaper than Lambda@Edge. In cache key normalization patterns, CloudFront Functions runs before the cache check, so it normalizes requests before any cache lookup occurs.

 

EC2 Placement Groups — Three Types

Placement groups control the physical placement of instances to optimize either performance or availability.

Cluster placement groups place instances close together on the same rack within a single AZ, enabling 25–100 Gbps ultra-low-latency, high-bandwidth inter-instance communication. They are optimal for tightly coupled HPC workloads and distributed ML training. Combined with FSx for Lustre, parallel file I/O is maximized as well.

Spread placement groups distribute instances across different physical racks. There is a hard limit of 7 instances per AZ, but this ensures a single hardware failure cannot affect multiple instances simultaneously. This is the high-availability pattern for a small number of critical instances.

Partition placement groups divide instances into partitions, each using an independent rack group. Up to 7 partitions per AZ are supported, with hundreds of instances per partition. This is the right choice for loosely coupled distributed systems like Hadoop, Cassandra, and HDFS.

 

Advanced Auto Scaling — Predictive Scaling + Warm Pools

Basic Auto Scaling reacts to current CPU or network metrics. Predictive Scaling uses ML to analyze historical patterns and scales out before a traffic surge arrives. If traffic spikes every morning at 9 AM, Predictive Scaling adds instances at 8:50 AM so they are ready when the surge hits.

Warm Pools maintain pre-started, pre-initialized instances in a waiting state. When a scale-out event occurs, these instances are immediately available without a cold start. For applications with long initialization times (post-boot data downloads, cache warming), Warm Pools dramatically reduce response-time degradation during scale-out.

 

Cost Optimization — Spot Instance Strategies

Spot Instances cost up to 90% less than On-Demand but can be reclaimed with 2 minutes of notice. They are ideal for fault-tolerant workloads that can handle interruptions.

ECS Fargate Spot brings Spot savings (up to 70%) to ECS workloads. It is especially suitable for stateless services and batch processing. A SIGTERM signal is sent 2 minutes before reclamation, allowing graceful shutdown.

Spot Fleet uses a diversified allocation strategy across multiple instance types and AZs. This reduces reclamation risk compared to relying on a single instance type and helps maintain desired capacity more reliably.

RDS Reserved Instances offer up to 72% discount with a 1-year or 3-year commitment. They are ideal for predictable, always-on workloads. Note that RDS Savings Plans do not exist (Savings Plans are EC2/Fargate only). Aurora Serverless v2 is optimized for variable workloads and is inefficient for fixed, always-on workloads.

 

Right-sizing — Using Compute Optimizer

Right-sizing adjusts instance sizes to match actual usage and eliminates unnecessary costs.

AWS Compute Optimizer provides right-sizing recommendations for EC2, Auto Scaling Groups, EBS volumes, and Lambda functions. For Lambda, it analyzes 14–93 days of invocation metrics to recommend optimal memory and execution time settings. The API returns current and recommended configurations along with estimated cost savings. An automation pattern combining EventBridge Scheduler and Lambda to regularly save recommendations to S3 as CSV also appears in the exam.

Cost Explorer provides Reserved Instance and Savings Plans recommendations but does not support Lambda memory optimization.

Eliminating unused resources is equally important. Use Cost Explorer to identify underutilized EC2 instances, and Trusted Advisor to find unattached EBS volumes and unused Elastic IPs.

 

S3 Cost Optimization — Intelligent-Tiering and Special Features

S3 Intelligent-Tiering is best for data with unpredictable access patterns. Objects not accessed for 30 days automatically move to the infrequent access tier; after 90 days they move to the archive instant access tier. There are no retrieval charges between tiers, making it economical for large data sets with low access frequency.

S3 Glacier Instant Retrieval is optimal for archive data accessed less than once per quarter that still needs immediate retrieval in milliseconds. It costs roughly 68% less than S3 Standard. Glacier Flexible Retrieval is even cheaper but requires 1 minute to 12 hours for restoration.

Incomplete multipart upload costs are an easy-to-miss waste. The S3 Lifecycle rule automatically deletes incomplete upload parts after a defined number of days.

 

Exam Key Points

"Non-HTTP (TCP/UDP) global acceleration, fixed IP required, DNS caching bypass" -- Global Accelerator

"HTTP content caching, CDN" -- CloudFront

"Reduce origin server load, additional caching layer" -- CloudFront Origin Shield

"Lightweight Viewer Request processing (URL normalization, query parameter sorting)" -- CloudFront Functions (not Lambda@Edge)

"HPC ultra-low latency high bandwidth, tightly coupled" -- Cluster Placement Group

"Hardware failure isolation for a small number of critical instances" -- Spread Placement Group

"Hadoop/Cassandra distributed systems" -- Partition Placement Group

"Stateless service up to 70% cost savings" -- ECS Fargate Spot

"RDS 24/7 steady workload up to 72% discount" -- RDS Reserved Instances (RDS Savings Plans do not exist)

"Lambda memory right-sizing recommendations" -- AWS Compute Optimizer (Cost Explorer does not support Lambda memory)

"Large data with unpredictable access patterns" -- S3 Intelligent-Tiering

Back to blog list