Amazon SageMaker Deep Dive

Master all 16 SageMaker features in one post — inference types, Clarify, Ground Truth, Pipelines, JumpStart, and more. The definitive SageMaker study guide.

Amazon SageMaker Deep Dive

Amazon SageMaker is AWS's fully managed ML platform covering the entire lifecycle from data preparation to model monitoring. The AIF-C01 exam tests your ability to map each SageMaker feature to the right scenario.

---

 

What is Amazon SageMaker?

A fully managed service that lets data scientists and developers build, train, and deploy ML models without managing servers. AWS handles the infrastructure automatically.

SageMaker covers: data collection and preparation, model build and training, hyperparameter tuning, model deployment, and performance monitoring.

---

 

Built-in ML Algorithms

Ready-to-use algorithms without writing code from scratch: Supervised learning: linear regression, classification, KNN Unsupervised learning: PCA, K-means, anomaly detection Specialized: NLP text summarization, image classification and detection

---

 

Automatic Model Tuning (AMT)

AMT automates hyperparameter optimization. Define a target metric, and AMT searches the hyperparameter space to find the best combination. Use this whenever the exam mentions automated hyperparameter optimization.

---

 

4 Deployment Types Compared

| Type | Latency | Max Payload | Best For | |------|---------|-------------|----------| | Real-Time | ms–seconds | 6 MB | Fraud detection, live recommendations | | Serverless | low (cold start) | 4 MB | Intermittent chatbots | | Asynchronous | minutes–hours | 1 GB | Medical imaging, large file processing | | Batch Transform | slowest | unlimited | Whole-dataset predictions |

---

 

Key Features Summary

Studio: unified IDE for all SageMaker features Data Wrangler: no-code data preparation and feature engineering Feature Store: centralized feature storage and sharing across teams Clarify: bias detection and model explainability Ground Truth: human labeling of training data (RLHF) Model Monitor: production drift detection Model Registry: version control and team sharing Pipelines: ML CI/CD automation JumpStart: pre-trained model hub (more models than Bedrock) Canvas: no-code ML for non-technical users

---

 

Exam Quick Reference

4 deployment types and when to use each Ground Truth = label training data; A2I = human review of low-confidence predictions Clarify = bias + explainability; Model Monitor = production drift JumpStart = developer hub; Canvas = no-code Fine-tuning cost order: Prompt Engineering < RAG < Instruction < Domain Adaptation

Back to blog list