ML Project Lifecycle & Hyperparameter Tuning
A machine learning project follows a clear sequence of steps — like building a house where you cannot put up walls before laying the foundation. The AIF-C01 exam presents these steps as real business scenarios.
---
The 8-Stage ML Project Lifecycle
ML projects are not linear. Like developing a recipe, if the result is bad you go back and adjust — it is a cycle.
Stage 1: Define Business Goals
Start with the problem, not the technology. Clarify what you want to achieve, set KPIs, and get stakeholder agreement on budget and constraints.
Stage 2: Frame the ML Problem
Convert the business problem into an ML task (classification, regression, etc.). Confirm that ML is actually the right solution.
Stage 3: Data Processing
"Garbage In, Garbage Out." Collect, clean, encode, and engineer features from raw data.
Stage 4: Exploratory Data Analysis (EDA)
Visualize data distributions, check correlations, and decide which features to keep or drop.
Stage 5: Model Development
Train multiple models, evaluate with metrics (Accuracy, F1, AUC, RMSE), and tune hyperparameters iteratively.
Stage 6: Retraining
Return here when performance drops or drift is detected. Add new data, adjust features, or change hyperparameters.
Stage 7: Deployment
Choose a deployment strategy: Real-time: low-latency online predictions Batch: large-scale offline predictions Serverless: intermittent, unpredictable traffic Asynchronous: large payloads with longer processing time
Stage 8: Monitoring
Track model performance, data quality, and latency in production. Detect model drift and trigger retraining when needed.
---
Hyperparameters vs Parameters
Parameters (weights): learned automatically during training Hyperparameters: set by humans before training begins
---
4 Key Hyperparameters
Learning Rate: controls the step size for updating weights. Too high = overshooting; too low = slow convergence. Batch Size: number of samples processed before updating weights. Smaller = stable but slow; larger = fast but less stable. Number of Epochs: how many times the model sees the full dataset. Too few = underfitting; too many = overfitting. Regularization: penalizes model complexity to reduce overfitting.
---
Overfitting — Causes and Fixes
Overfitting means the model memorizes training data but fails on new data. Five fixes:
| Fix | Description | |-----|-------------| | Data expansion | More and more diverse training data | | Early stopping | Stop training when validation performance starts dropping | | Data augmentation | Generate new samples by transforming existing data | | Hyperparameter tuning | Adjust learning rate, epochs, regularization | | Feature reduction | Remove low-value features |
---
Exam Quick Reference
8 stages in order: Define Goals → Frame Problem → Data Processing → EDA → Model Development → Retraining → Deployment → Monitoring Hyperparameters are set before training; parameters are learned during training SageMaker AMT = automated hyperparameter tuning Too few epochs = underfitting; too many = overfitting Drift detected → go back to retraining