Model Selection: A Question of Judgment, Not Just Algorithm Knowledge
One of the biggest mindset shifts from junior to senior ML engineer is realizing that knowing the latest algorithms matters far less than knowing when and why to use each approach. You might be able to implement a transformer from scratch in PyTorch, but if you cannot decide whether XGBoost is better than a foundation model for a given tabular classification problem, you will waste budget and time on every project.
AWS MLA-C01 Domain 2 (ML Model Development), Task 2.1 tests exactly this judgment. It asks you to match problem types, data characteristics, and operational constraints to the right modeling strategy on AWS. Not "implement XGBoost" but "which approach fits this scenario and why."
Starting with Problem Type
The right starting point for model selection is always the problem type. Business requirements must be translated into ML problem types before any algorithm discussion begins.
| ML Problem Type | Examples | SageMaker Approach | |----------------|----------|--------------------| | Binary Classification | Churn prediction, fraud detection | XGBoost, Linear Learner | | Multi-class Classification | Product categorization, sentiment | XGBoost, BlazingText | | Regression | Price prediction, demand forecasting | XGBoost, Linear Learner, DeepAR | | Time Series Forecasting | Revenue, traffic forecasting | DeepAR, Amazon Forecast | | Anomaly Detection | Network intrusion, manufacturing defects | Random Cut Forest | | Clustering | Customer segmentation | K-Means | | Dimensionality Reduction | Feature compression, visualization | PCA | | Image Classification | Product images, medical imaging | SageMaker Image Classification, Rekognition | | Object Detection | Autonomous vehicles, inventory | SageMaker Object Detection | | Text Classification / Embeddings | Review sentiment, document search | BlazingText | | Machine Translation / Summarization | Multilingual support, document summary | Seq2Seq, Amazon Translate | | Generative AI | Chatbots, code generation | Amazon Bedrock, SageMaker JumpStart | | Recommendation | Product recommendations, personalization | Amazon Personalize |
This mapping is a starting point, not a fixed rule. Data characteristics and operational requirements drive further choices.
SageMaker Built-in Algorithms in Depth
SageMaker Built-in Algorithms are AWS-optimized implementations that eliminate the need to write training code from scratch. You specify an algorithm image URI, configure hyperparameters, and start a training job. These algorithms are designed for SageMaker's distributed training infrastructure, so they scale better than equivalent custom implementations.
For supervised learning, XGBoost is the workhorse of tabular ML. It handles missing values natively, supports sparse data, and outputs feature importance. It is consistently the first thing to try on tabular classification and regression tasks. Linear Learner is faster and more interpretable but requires linearly separable relationships. Use it as a fast baseline, especially when you need coefficient-level explanations.
Computer vision built-ins include Image Classification (ResNet-based, supports transfer learning with custom labels), Object Detection (SSD/YOLO-style bounding box prediction), and Semantic Segmentation (pixel-level object identification, the most granular computer vision task).
For natural language, BlazingText operates in two modes: Word2Vec for word embeddings and supervised text classification. Sequence to Sequence handles encoder-decoder tasks like translation and summarization, though it is increasingly replaced by JumpStart foundation models for most practical NLP workloads.
Unsupervised algorithms include K-Means (clustering, often paired with PCA for high-dimensional data) and Random Cut Forest, which outputs anomaly scores without labeled data and works well on both point anomalies and time series patterns.
DeepAR is SageMaker's time series specialist. It trains across multiple related time series simultaneously using RNN architecture, capturing complex seasonality patterns and covariate effects that simpler ARIMA models miss.
Built-in Algorithms vs BYOC: When to Go Custom
Despite the strength of built-in algorithms, there are clear signals that custom training containers (BYOC) are the right choice. Use built-in algorithms when the problem type maps cleanly to an available algorithm, when you want AWS-managed distributed training, and when iteration speed matters more than customization. Switch to BYOC or Script Mode when your team has existing model code, when you need specific library versions or frameworks, when the problem type has no built-in equivalent (reinforcement learning, for example), or when you are implementing a recent research paper architecture.
SageMaker Script Mode is the practical middle ground. AWS provides managed containers for PyTorch, TensorFlow, and scikit-learn, and you inject your custom training script into them. This covers most custom training scenarios without the overhead of fully managing a Docker image.
!Built-in algorithms versus BYOC
SageMaker JumpStart and Foundation Models
JumpStart is a hub of pre-trained models, solution templates, and example notebooks. Since 2024, its primary value proposition is as the entry point for foundation models that you want to run in your own AWS account infrastructure.
JumpStart model categories include text generation LLMs (Llama 2/3, Mistral, Falcon), text-to-image models (Stable Diffusion, Amazon Titan Image Generator), and embedding models for RAG systems. Deployment and fine-tuning are one-click operations. For fine-tuning, you upload domain-specific data to S3 and configure a fine-tuning job — the training code is automatically generated and managed.
The key distinction from Amazon Bedrock is the deployment model. JumpStart runs model weights in your own AWS account on SageMaker endpoints, giving you full control over the infrastructure, custom fine-tuning capability, and clear data isolation. Bedrock is fully serverless — AWS manages all infrastructure, and you access models via API calls only.
Amazon Bedrock and Generative AI Scenarios
Amazon Bedrock provides access to foundation models from multiple providers (Anthropic Claude, Amazon Titan, Meta Llama, Stability AI) through a single API, with a guarantee that your data is never used for model training. This serverless, API-first approach is the right choice when you need generative AI capabilities without provisioning or managing model infrastructure.
Bedrock's key features include text generation and summarization (Claude, Titan Text, Llama), image generation (Stable Diffusion, Amazon Titan Image Generator), embeddings (Amazon Titan Embeddings for RAG), Knowledge Bases for Bedrock (automated document ingestion, embedding, and storage in OpenSearch Serverless or Aurora for RAG pipelines), and Agents for Bedrock (multi-step workflows where LLMs interact with external APIs).
Exam questions on Bedrock typically take the form: "The team wants to add generative AI features with minimum infrastructure management — what should they use?" The serverless, API-only access model is what distinguishes Bedrock from SageMaker endpoints.
Transfer Learning, Fine-Tuning, and Hugging Face Integration
Transfer learning — adapting a pre-trained model to domain-specific data — is one of the most practical patterns in modern ML. It is especially powerful when labeled data is scarce, because the pre-trained model has already learned general representations from massive datasets.
SageMaker supports transfer learning through three primary paths. JumpStart fine-tuning provides pre-built fine-tuning code alongside each foundation model, making it the easiest starting point. The SageMaker HuggingFace Estimator allows you to load any model from the Hugging Face Hub and fine-tune with the transformers library on SageMaker infrastructure, with full distributed training support. Bedrock fine-tuning enables custom adaptation for supported models (certain Titan and Claude versions) entirely through API, with no infrastructure management.
For large models, PEFT (Parameter-Efficient Fine-Tuning) techniques like LoRA are important to understand. Rather than updating all model parameters, LoRA trains a small number of adapter weights, dramatically reducing GPU memory requirements and training time. JumpStart uses LoRA by default for LLM fine-tuning.
Interpretability vs Accuracy Tradeoff
Model selection is not purely about maximizing accuracy. Interpretability is a first-class requirement in regulated industries (finance, healthcare, insurance) and anywhere decisions must be explained to stakeholders or regulators. Linear models and tree ensembles like XGBoost are inherently more interpretable. SageMaker Clarify's SHAP-based feature importance makes this even stronger.
Deep learning models dominate on unstructured data (images, text, audio) but carry black-box characteristics. Post-hoc explanation techniques (SHAP, LIME, Grad-CAM) partially address this but do not eliminate fundamental opacity. Foundation models are even less interpretable. Knowing when interpretability requirements should override raw accuracy is a key judgment skill tested on MLA-C01.
Exam Tips
"Tabular data, classification or regression" -- XGBoost or Linear Learner first "Anomaly detection, no labels" -- Random Cut Forest (RCF) "Time series forecasting, multiple series" -- DeepAR or Amazon Forecast "Generative AI, no infrastructure management" -- Amazon Bedrock "LLM fine-tuning on custom data, own infrastructure" -- SageMaker JumpStart "Fast baseline, low ML expertise on team" -- SageMaker Autopilot "Standard capability: translation, speech, image analysis" -- Amazon AI managed services "Custom training code,