AIF-C01 covers the ML development lifecycle at approximately 20% of the exam. Key topics: data collection, preprocessing, model training, overfitting, evaluation metrics, and MLOps.
---
The ML Pipeline — A Cooking Analogy
| ML Stage | Cooking Analogy | Goal | |----------|----------------|------| | 1. Data collection | Sourcing ingredients | Sufficient, diverse data | | 2. EDA | Checking freshness | Understand distributions, missing values, outliers | | 3. Preprocessing | Cleaning and prepping | Handle nulls, normalize, encode | | 4. Feature engineering | Developing the recipe | Transform data into model-friendly form | | 5. Model training | Cooking | Algorithm learns patterns from data | | 6. Hyperparameter tuning | Seasoning | Optimize model configuration | | 7. Evaluation | Tasting | Measure precision, recall, F1, AUC | | 8. Deployment | Serving | Provide predictions in production | | 9. Monitoring | Ongoing quality control | Detect performance degradation over time |
!The ML development lifecycle from data collection to deployment
Data Collection
Data Augmentation Artificially expanding training data by transforming existing examples (flip, rotate, adjust brightness for images). Useful when training data is scarce.
Class Imbalance If 99% of transactions are normal and 1% are fraud, a model that always predicts "normal" achieves 99% accuracy — but is useless. Use oversampling or undersampling to address this.
---
EDA and Preprocessing
EDA checks data before training: distributions, missing values, outliers, correlations.
Preprocessing fixes issues found: fill missing values, normalize/standardize numeric features, remove outliers, encode categorical variables as numbers.
---
Overfitting
A model that memorizes training data exactly performs well on training data but poorly on new data — like a student who memorizes practice questions but fails on a slightly different test.
| Problem | Cause | Solution | |---------|-------|---------| | Overfitting | Model too complex, too little data | Regularization, dropout, data augmentation, cross-validation | | Underfitting | Model too simple | More complex model, add features |
Data splits: Training (~70%) / Validation (~15%) / Test (~15%)
---
Evaluation Metrics
| Metric | What it measures | Best when | |--------|-----------------|-----------| | Accuracy | Overall correct predictions | Balanced classes | | Precision | Of predicted positives, how many are correct | False positives are costly (spam filter) | | Recall | Of actual positives, how many were caught | False negatives are costly (cancer detection) | | F1 Score | Harmonic mean of precision and recall | Both matter equally | | AUC-ROC | Overall classifier performance | Threshold-independent comparison |
---
Deployment and MLOps
| Tool | Role | |------|------| | SageMaker Endpoints | Real-time inference API | | SageMaker Batch Transform | Large-scale batch inference | | SageMaker Model Monitor | Detect data drift and performance degradation | | SageMaker Pipelines | Automate the full ML workflow (MLOps) | | SageMaker Experiments | Track and compare multiple experiment runs |
---
Exam Key Points
"Check distributions, missing values, outliers" — EDA "Expand training data by transforming existing examples" — Data Augmentation "Model memorizes training data, fails on new data" — Overfitting "Of predicted positives, actual positives" — Precision "Of actual positives, how many caught" — Recall "Balanced precision + recall metric" — F1 Score "Detect performance degradation after deployment" — SageMaker Model Monitor "ML pipeline automation, DevOps for ML" — SageMaker Pipelines / MLOps