Building MLOps and CI/CD Pipelines

MLOps applies software engineering best practices to ML model development, deployment, and operations. This guide covers SageMaker Pipelines, Model Registry, SageMaker Projects, and integration with AWS CI/CD tools to build reproducible and automated ML workflows.

Why MLOps Matters: Reproducibility, Automation, Governance

In traditional software development, CI/CD automatically builds, tests, and deploys code changes. In machine learning, you must version not just code but also data and models, and the non-deterministic nature of training results requires more complex validation gates. MLOps is the systematic approach to managing this complexity.

Reproducibility is the foundational principle: you need to be able to recreate the exact results of a model trained six months ago, trace exactly what data was used, what hyperparameters were applied, and which evaluation thresholds were passed. Automation accelerates the iterative experimentation and retraining cycle while reducing human error. Governance creates an auditable record of which models are deployed in production, who approved them, and how they are performing.

 

SageMaker Pipelines: The Backbone of ML Workflows

SageMaker Pipelines is a fully managed service for defining and running ML workflows as directed acyclic graphs (DAGs). Input/output relationships and step dependencies are defined in code, and execution history and artifact lineage are tracked automatically.

The key step types are: ProcessingStep for data preprocessing, feature engineering, and post-processing (wraps SageMaker Processing Jobs with processors like SKLearnProcessor or PySparkProcessor); TrainingStep for running training jobs and capturing model artifacts; TuningStep for hyperparameter optimization (HPO); TransformStep for integrating Batch Transform into the pipeline; ModelStep for creating SageMaker model resources; ConditionStep for conditional branching (for example, only proceed to deployment if model accuracy exceeds a threshold); and CallbackStep for invoking external services or Lambda functions and waiting for a completion signal, useful for human review loops or external system integration.

 

Pipeline Parameters and Caching

Pipeline parameters allow values such as instance type, training epochs, and data S3 paths to be declared at definition time and injected at runtime. This allows the same pipeline definition to be reused across development, staging, and production environments.

Step caching reduces pipeline execution cost and time. If a step has already been successfully run with identical inputs, the cached result is reused instead of re-executing. Preprocessing steps do not need to rerun if the data has not changed, and training steps can leverage the cache if code and data are identical. Cache expiration can be configured to invalidate stale caches.

 

SageMaker Model Registry: Versioning and Approval Workflow

The Model Registry is a catalog of trained models. Models are registered as versions within model groups, with each version containing the model artifact location, training container, evaluation metrics, and metadata.

Every model version has one of three approval statuses: PendingManualApproval (the default, awaiting review), Approved (cleared for production deployment), or Rejected (failed to meet criteria). The approval workflow can combine automated evaluation and human review. A ConditionStep automatically checks metrics (for example, RMSE below 0.05, AUC above 0.90) and registers passing models in the Registry. A data scientist or ML engineer then reviews detailed metrics, explainability reports, and bias assessment results before manually approving. Approval events can trigger automated deployment pipelines via EventBridge.

 

SageMaker Projects: MLOps Templates

SageMaker Projects provide templates with MLOps best practices baked in. The built-in templates include a complete CI/CD structure connecting SageMaker Pipelines, CodePipeline, and the Model Registry. Custom templates reflecting organizational standards can be distributed and shared via AWS Service Catalog.

Creating a project automatically generates a model build repository (CodeCommit or GitHub) and a model deployment repository. Changes to the build repository trigger a SageMaker Pipeline execution, and when a model reaches Approved status, a CodePipeline in the deployment repository automatically updates the production endpoint.

 

Integration with AWS CI/CD Tools

While SageMaker Pipelines handles ML-specific workflows, AWS's general-purpose CI/CD tools manage code change detection, builds, and infrastructure deployment.

CodePipeline serves as the CI/CD orchestrator, connecting source stages (CodeCommit, GitHub, or ECR image change detection), build stages (CodeBuild for container builds, unit tests, and SageMaker Pipeline execution), and deploy stages (CloudFormation for endpoint updates). CodeBuild is a serverless build environment responsible for installing Python packages, running model integration tests, and building and pushing container images to ECR. CodeDeploy manages deployment strategies for EC2 or ECS-based inference servers.

 

Deployment Strategies: Blue/Green, Canary, Linear

SageMaker endpoint updates support three deployment strategies.

Blue/green deployment fully provisions the new version (green) before switching all traffic at once. The green environment can be thoroughly validated before cutover, and rollback is immediate if problems arise. The temporary cost of running two full sets of instances is the main tradeoff.

Canary deployment sends a fraction of traffic (for example, 5%) to the new version first, monitors metrics, and then completes the transition once the deployment is deemed safe. This allows progressive validation against real production traffic.

Linear deployment increases the traffic percentage in fixed steps at regular intervals — for example, 10% more every 10 minutes. More gradual than canary, but the full transition takes longer.

!3 SageMaker deployment strategies

Step Functions: Complex ML Workflows

Step Functions is useful when broader AWS service integration is needed beyond what SageMaker Pipelines provides. Lambda functions, ECS/Fargate tasks, AWS Glue jobs, SNS/SQS, and DynamoDB can be natively integrated, and complex error handling and retry logic can be expressed as state machines.

SageMaker Pipelines excels at ML experiment tracking and artifact lineage, while Step Functions is stronger for heterogeneous service composition, parallel processing, and complex branching logic. In practice, the two can be combined — Step Functions handles overall orchestration and calls SageMaker Pipelines as one of its tasks.

 

EventBridge and Automated Retraining Triggers

EventBridge acts as the hub for event-driven ML pipeline triggering. Possible patterns include: schedule-based retraining using cron expressions; detecting Model Registry approval status changes to trigger CodePipeline execution; receiving data drift alarms from SageMaker Model Monitor to start retraining pipelines; and automatically running pipelines when new data arrives in S3. This enables a fully automated continuous training loop — data drift is detected, a model is automatically retrained, passes evaluation, and is deployed to production without human intervention.

 

Testing Strategy and Model Approval Gates

Like software CI/CD, ML pipelines need multiple test levels. Unit tests verify the correctness of preprocessing functions, feature engineering logic, and custom training code. Integration tests run the full pipeline on a small dataset to confirm the end-to-end flow works. Model validation tests check whether performance metrics (accuracy, F1, RMSE) meet thresholds on a holdout dataset, whether bias evaluation passes, and whether inference latency falls within SLA bounds.

Model approval gates consist of two stages: automated validation and human review. Models passing the ConditionStep's automated validation are registered in the Registry. Organizations then choose between manual approval by a data scientist and fully automated approval based on metrics alone. Regulated industries (healthcare, finance) typically require human review — this is referred to as a human-in-the-loop.

 

Version Control and Continuous Training

Git-based version control should encompass not only code but also pipeline definitions, hyperparameter configurations, and evaluation criteria. SageMaker Experiments automatically links code version, data version, parameters, and metrics for each run, tracking full artifact lineage.

Continuous Training triggers automatic retraining when model performance drops below a threshold or when sufficient new data accumulates. The standard pattern is: SageMaker Model Monitor detects data drift or model quality degradation, an EventBridge alarm starts the retraining pipeline, the pipeline evaluates the new model, and if it passes, the model is automatically deployed.

 

Exam Tips

"Define ML workflow as DAG, track artifact lineage" -- SageMaker Pipelines "Model versioning, approved/rejected status" -- SageMaker Model Registry "Pending → Approved → trigger automated deployment" -- Model Registry + EventBridge + CodePipeline "Conditional step, branch on metric threshold" -- ConditionStep in SageMaker Pipelines "Heterogeneous service workflow (Lambda, Glue, ECS)" -- AWS Step Functions "MLOps best-practice templates, quick start" -- SageMaker Projects "Prevent re-execution of step with same inputs" -- Pipeline step caching "Data drift detected → automated retraining" -- Model Monitor + EventBridge "Deployment approval with human review" -- Human-in-the-loop (CallbackStep or manual approval) "Route 5% traffic to new version first" -- Canary deployment strategy

Back to blog list