The Deployment Decision: Matching Inference Pattern to Infrastructure
Deploying a trained model reliably into production is often more complex than training it in the first place. AWS SageMaker offers four distinct inference modes, each optimized for a specific workload pattern. A large portion of the MLA-C01 exam tests your ability to read a scenario and select the right deployment type. The three key criteria to evaluate first are latency tolerance, payload size, and traffic pattern.
Real-Time Endpoints: When Every Millisecond Counts
SageMaker real-time endpoints are the right choice when an application needs predictions in hundreds of milliseconds. Recommendation engines in web apps, financial fraud detection, and real-time image analysis are classic examples. You specify the instance type and count when creating the endpoint.
For instance selection, GPU-based instances (ml.g4dn, ml.g5, ml.p3) are necessary for deep learning models, while lighter models and traditional ML often run efficiently on CPU instances (ml.c5, ml.m5). AWS Inferentia-based instances (ml.inf1, ml.inf2) can deliver up to 70% cost savings compared to GPU instances at high throughput.
Production variants allow you to deploy multiple model versions to a single endpoint with configurable traffic splits. This enables A/B testing and canary deployments — for example, routing 10% of traffic to a new model version before promoting it fully.
Batch Transform: Offline Scoring at Scale
Batch Transform is ideal when you need to score large volumes of data already stored in S3 and real-time response is not required. Nightly batch scoring, preprocessing in data pipelines, and generating offline recommendations are typical use cases. Instances are automatically terminated when the job completes, so there is no ongoing cost. Multiple instances can be specified to process data in parallel, with SageMaker handling data splitting and result reassembly automatically.
Serverless Inference: Cost Optimization for Intermittent Traffic
Serverless Inference is designed for workloads with unpredictable or sparse traffic. You pay only for the compute time used and the data processed — no cost during idle periods. The main tradeoff is cold start latency: when the container has been idle, the first request after a period of inactivity will experience a delay of several seconds while the container restarts. This makes it unsuitable for user-facing services requiring consistent low latency, but it is cost-effective for internal tools, development environments, and infrequently called prediction APIs. You configure the memory size (1GB to 6GB), which directly affects pricing.
Async Inference: Large Payloads and Long Processing Times
Async Inference is the right choice when request payloads are large (up to 1GB) or when processing time is long enough that synchronous waiting is impractical. Requests are placed in a queue, SageMaker processes them asynchronously, and results are stored in S3. Clients can poll the result location or receive SNS notifications. Medical image analysis, video processing, and large document analysis are typical scenarios. Async endpoints support Auto Scaling based on queue depth.
Comparison of All Four Inference Modes
| Mode | Latency | Max Payload | Cost Model | Typical Use Case | |------|---------|-------------|------------|-----------------| | Real-Time Endpoint | Milliseconds | 6MB | Instance uptime | Real-time recommendations, fraud detection | | Batch Transform | N/A (offline) | Unlimited | Job duration | Large-scale offline scoring | | Serverless Inference | Tens to hundreds of ms (cold start) | 4MB | Invocations + duration | Intermittent traffic, internal tools | | Async Inference | Seconds to minutes | 1GB | Instance uptime | Large file processing |
!The 4 SageMaker inference modes
Multi-Model Endpoints and Multi-Container Endpoints
Multi-Model Endpoints (MME) allow hundreds or thousands of similar models to be served from a single endpoint. Models are stored in S3 and dynamically loaded on demand, with frequently used models cached in memory and inactive ones unloaded. This is cost-effective when you have many per-customer or per-region models, each with low individual traffic.
Multi-Container Endpoints run up to 15 different containers within a single endpoint. Each container can host a different framework or model, and requests can be routed sequentially (pipeline mode) or directly. This suits ML inference pipelines where preprocessing, classification, and postprocessing models are chained together.
SageMaker Neo and AWS Inferentia: Hardware Optimization
SageMaker Neo compiles models for a specific hardware target, improving inference performance. It supports TensorFlow, PyTorch, MXNet, and XGBoost, with targets including x86/ARM EC2 instances, edge devices (NVIDIA Jetson, Raspberry Pi), and AWS Inferentia. Compiled models run on the Neo Runtime (DLR).
AWS Inferentia is Amazon's purpose-built ML inference chip. Inf1 instances (Inferentia1) are widely used in production, while Inf2 instances (Inferentia2) deliver more memory and bandwidth for large models and generative AI workloads. Models are compiled and deployed using the Neuron SDK.
Auto Scaling Policies
SageMaker endpoints integrate with Application Auto Scaling and support three scaling policy types.
Target Tracking is the most common approach, maintaining a CloudWatch metric such as InvocationsPerInstance or CPUUtilization at a target value by automatically adding or removing instances. It requires minimal configuration and works well for most use cases.
Step Scaling applies different scaling amounts based on how far a metric has breached a threshold. For example, add one instance if InvocationsPerInstance exceeds 70, add three if it exceeds 90. This is useful for workloads with predictable traffic surge patterns.
Scheduled Scaling pre-warms instances before known traffic spikes. If traffic predictably surges every Monday morning, instances can be scaled up in advance to avoid cold start delays. Custom CloudWatch metrics — including queue depth or business-level signals — can also serve as scaling triggers.
Container Options: Built-in vs BYOC
SageMaker provides pre-built containers for major frameworks (TensorFlow, PyTorch, XGBoost, Scikit-learn). Using these eliminates the need to manage containers directly — you supply model artifacts and AWS handles security patches and framework updates. For specialized libraries, custom preprocessing logic, or proprietary software, you can bring your own container (BYOC). Build a Docker image, push it to ECR (Amazon Elastic Container Registry), and register it with SageMaker. ECR supports image versioning, vulnerability scanning, and replication across regions.
VPC Configuration and Edge Deployment
In production, SageMaker endpoints are typically deployed inside a VPC so they are not directly exposed to the internet. In VPC mode, access to SageMaker services goes through PrivateLink, and traffic between endpoints and S3 stays within the VPC. For edge devices, combine AWS IoT Greengrass with SageMaker Edge Manager — Neo-compiled models are deployed to devices, and Edge Manager handles fleet-wide model versioning and performance monitoring. Using CloudFormation or CDK to define SageMaker endpoints as infrastructure-as-code ensures consistency across environments and tracks all changes.
Exam Tips
"Millisecond latency, continuous traffic" -- Real-time endpoint "Process large dataset in bulk, results to S3" -- Batch Transform "Intermittent traffic, minimize cost" -- Serverless Inference (if cold start acceptable) "1GB payload, long processing time" -- Async Inference "Thousands of models, low per-model traffic" -- Multi-Model Endpoint (MME) "Reduce GPU inference cost" -- AWS Inferentia (Inf1/Inf2) "Compile model for edge device" -- SageMaker Neo + IoT Greengrass "Auto scale on InvocationsPerInstance" -- Target Tracking Auto Scaling "Route 10% traffic to new model version" -- Production Variant "Custom framework, proprietary library" -- BYOC + ECR