Validates skills in building, deploying, and operating ML workloads on AWS. Covers SageMaker and MLOps pipelines.
Data Ingestion & Storage (S3, Kinesis, Glue), Feature Engineering (SageMaker Data Wrangler, Feature Store), Data Quality & Governance (Clarify, Lake Formation)
Choosing a Modeling Approach (Built-in Algorithms, JumpStart, Bedrock), Model Training & Tuning (Distributed Training, HPO, Debugger), Model Analysis & Evaluation (Clarify, SHAP, Performance Metrics)
Selecting Deployment Infrastructure (Real-time, Batch, Serverless Inference), ML Infrastructure Scaling (Auto Scaling, Containers, VPC), CI/CD Pipelines (SageMaker Pipelines, CodePipeline, MLOps)
Model Inference Monitoring (Model Monitor, Drift Detection), Infrastructure & Cost Optimization (CloudWatch, Inference Recommender), Securing ML Solutions (IAM, KMS, VPC Isolation)
Try real exam-style questions from the free sample set. Each answer comes with a full explanation.
An MLOps team at a financial company is experiencing frequent deployment delays and configuration inconsistencies across environments because data scientists manually approve and deploy models trained in Amazon SageMaker.
The team wants to build a pipeline that automatically validates model artifacts after training, deploys to a staging environment, goes through an approval step, and then automatically deploys to production endpoints. Infrastructure configuration must be managed as code, and reproducibility of the deployment environment must be guaranteed.
Which CI/CD pipeline configuration BEST minimizes operational overhead while improving consistency and automation of model deployment?
Answer: B. CodePipeline + CloudFormation + approval stage
The core of deployment automation is a pipeline where the entire process from training completion to production is defined as code and executed without human intervention. AWS CodePipeline orchestrates model artifact validation, staging deployment, manual approval gate, and production deployment sequentially, which directly matches the requirements in this scenario. AWS CloudFormation declares infrastructure as YAML/JSON templates to provision staging and production environments identically, preventing configuration inconsistencies at the source.
AWS CodePipeline is a fully managed continuous delivery service that inserts a Manual Approval Action so that production deployment only proceeds after a designated reviewer approves the staging validation. Integrating with the Amazon SageMaker Model Registry allows governance configurations where only approved model versions enter the pipeline. AWS CloudFormation ensures no environment drift occurs by reproducing environments repeatedly from the same template.
On the exam, when IaC + approval stage + automated deployment appear together, CodePipeline + CloudFormation is the correct combination. Remember that Amazon SageMaker Pipelines is dedicated to ML training and evaluation workflows, while CodePipeline handles the overall CI/CD orchestration.
A fintech company needs to deploy a fraud detection model trained on SageMaker so that it can be called via HTTPS from their own on-premises applications. Per security policy, all inference requests must be encrypted with TLS, and the on-premises network must access through AWS internal networks without going through the public internet.
Which inference endpoint access configuration BEST meets the security requirements while minimizing changes to existing infrastructure?
Answer: B. Direct Connect + PrivateLink VPC endpoint
To select the correct answer, you need to identify the architecture that simultaneously meets 'no public internet traversal' and 'direct on-premises access'. AWS Direct Connect provides a dedicated physical circuit between the on-premises data center and AWS so that traffic never goes through the internet. Creating a PrivateLink-based VPC interface endpoint for SageMaker Runtime allows on-premises traffic entering via Direct Connect to reach the SageMaker endpoint through the VPC internal network, with TLS encryption provided automatically via HTTPS.
AWS Direct Connect provides dedicated physical circuits from 1 Gbps to 100 Gbps, guaranteeing consistent low latency and reliable bandwidth without internet traversal. AWS PrivateLink (VPC Interface Endpoint) assigns VPC internal IP addresses to AWS services for private access. For SageMaker Runtime, you create the com.amazonaws.{region}.sagemaker.runtime endpoint service within the VPC. Combining both services creates a fully private path for on-premises access to SageMaker inference endpoints without any internet exposure.
VPN is an internet-based encrypted tunnel, so it traverses the public internet, and combining it with a public endpoint as in Option 1 fails to meet security requirements. API Gateway + Lambda proxy requires additional infrastructure components, violating the 'minimize existing infrastructure changes' requirement. CloudFront is a fully internet-based CDN that fundamentally conflicts with the non-internet access requirement. Key exam pattern: 'on-premises to AWS fully private connection' → Direct Connect; 'private access to AWS services from VPC' → PrivateLink; their combination is the standard pattern for accessing AWS services from on-premises without internet.
A media streaming company operates a content recommendation service using a SageMaker real-time inference endpoint.
Traffic is low during weekday daytime hours, but spikes 3-5x during evenings and weekends. The current fixed instance count causes cost waste when traffic is low and response latency when traffic is high.
Which auto scaling policy type BEST automatically adjusts the number of instances based on a specific metric while preventing excessive scaling costs?
Answer: C. Target Tracking Scaling policy
This question tests whether you can identify, among three auto scaling policy types, the one that 'automatically adjusts based on a specific metric while preventing excessive costs'. Target Tracking Scaling sets a target value for a metric such as InvocationsPerInstance, and SageMaker Application Auto Scaling automatically adjusts the instance count to maintain that target. It scales out when the metric exceeds the target and scales in when it falls below, precisely responding to traffic changes while automatically removing unnecessary instances.
SageMaker endpoint auto scaling operates on AWS Application Auto Scaling. The representative metric for Target Tracking Scaling is SageMakerVariantInvocationsPerInstance (inference calls per minute), and setting a target value (e.g., 70 requests per minute) causes the system to automatically adjust instances to maintain that value. Separate ScaleInCooldown and ScaleOutCooldown settings prevent excessive frequent scaling. AWS official documentation also specifies Target Tracking as the default recommended auto scaling policy for SageMaker endpoints.
Manual instance scaling fails to meet the automation requirement at all. Simple Scaling takes no additional action during cooldown after a single threshold trigger, making it inefficient for rapid traffic changes. Step Scaling enables automation but requires manual design of threshold ranges and adjustment amounts, resulting in higher operational burden and lower precision than Target Tracking. On the exam, when the core requirement is 'automatic adjustment based on a specific metric', choose Target Tracking Scaling.
An online gaming company wants to train a churn prediction model in Amazon SageMaker using hundreds of millions of user behavior event data stored in Amazon DynamoDB.
The DynamoDB data is continuously growing and contains nested JSON attributes, making it difficult to use directly for ML training. To regularly extract large-scale data without affecting the read performance of the operational DynamoDB table and transform it into a structured format suitable for training to load into Amazon S3,
Which ETL pipeline design approach is MOST appropriate?
Answer: C. Glue ETL + DynamoDB Export to S3
The most critical requirement in this scenario is extracting data without affecting the read performance of the operational DynamoDB table. The Scan API and EMR DynamoDB Connector directly read the table and consume Read Capacity Units (RCU), violating this requirement. DynamoDB Export to S3 exports data based on PITR (Point-in-Time Recovery) snapshots, having absolutely no impact on the operational table's read capacity, and AWS Glue ETL subsequently transforms the nested JSON structure into a structured format suitable for ML.
DynamoDB Export to S3 is a fully managed export feature available when PITR is enabled on the DynamoDB table. It exports the entire table or a point-in-time snapshot to S3 in DynamoDB JSON or Amazon Ion format, and the export operation itself does not consume table RCU. AWS Glue ETL is a serverless Apache Spark-based service that uses the DynamicFrame API to flatten nested JSON structures and convert them into Parquet or CSV formats readable by SageMaker. Integration with Amazon EventBridge Scheduler allows configuration of periodic automated pipeline execution.
On the exam, when large-scale extraction without affecting DynamoDB operational performance appears, DynamoDB Export to S3 is the key. Distinguish that DynamoDB Streams is dedicated to change data capture (CDC) and unsuitable for full batch extraction, while Scan API and EMR Connector directly consume RCU and affect operational performance.
An autonomous driving technology company is training deep learning models in Amazon SageMaker using millions of small image files (each a few KB to several MB) and large sensor log files of tens of gigabytes simultaneously. Currently, all data is read directly from Amazon S3, causing data loading time at the start of training to account for a significant portion of total training time, with GPU utilization remaining low.
Which data access strategy is MOST appropriate for small files and large files respectively?
Answer: B. S3 Fast File Mode + FSx for Lustre
To select the correct answer, you first need to understand the I/O characteristic differences between many small files and large single files. Millions of small image files have an IOPS (input/output operations per second) bottleneck as the core issue, and Amazon S3 Fast File Mode minimizes loading latency by streaming directly from S3 in a POSIX-compatible manner without downloading files to the training instance. Tens-of-gigabyte large sensor logs require sequential throughput, and Amazon FSx for Lustre is an HPC-optimized parallel file system that directly integrates with S3 buckets to deliver hundreds of GB/s of aggregate throughput.
SageMaker has three data input modes: File Mode, Pipe Mode, and Fast File Mode. Fast File Mode mounts S3 objects like a file system and streams data simultaneously with training start, eliminating pre-download wait time and optimized for random access to small files. Amazon FSx for Lustre is a fully managed parallel file system where multiple GPU instances in distributed training can access simultaneously, and auto-integrates with the S3 data layer so costs can be controlled by provisioning only during training.
When you see many small files + low GPU utilization, remember Fast File Mode; when you see large files + high throughput, remember FSx for Lustre. Amazon EBS can only be attached to a single instance, which makes it fatally limited in multi-instance distributed training scenarios where data sharing is impossible.
The AWS Certified Machine Learning Engineer – Associate (MLA-C01) exam consists of 65 questions with a 130-minute time limit.
The passing score for the MLA-C01 exam is 720 out of 1000.
The MLA-C01 sample set (20 questions with full explanations) is free. The full bank of 340+ questions is available with a CloudMasterIT subscription.
The MLA-C01 certification is valid for 3 years after passing. Recertification is required after that.