FM Deployment Strategy: Bedrock vs SageMaker
When integrating foundation models into an application on AWS, the first architectural decision is whether to use Amazon Bedrock or Amazon SageMaker. The right choice depends on how much model customization you need, how much infrastructure management you can accept, and your cost model preference.
| Criterion | Amazon Bedrock | Amazon SageMaker | |-----------|--------------|----------------| | Infrastructure | Fully managed (serverless) | Choose instance type and count | | Model support | AWS partner models (Claude, Titan, Llama, etc.) | Open-source and custom models | | Fine-tuning | Bedrock Fine-tuning (limited) | Full training and fine-tuning pipeline | | Billing | Per-token pay-as-you-go | Instance uptime | | Best for | Immediate API use, prototyping to production | Custom models, high-performance specialized workloads |
The AIP-C01 exam tests this decision frequently. If you can use a pre-trained FM as-is or with limited fine-tuning, Bedrock is the answer. If you need a fully custom model or specific GPU instances, SageMaker is the answer.
SageMaker Inference Endpoint Types
SageMaker provides four inference endpoint types for different workload characteristics. The exam tests which type fits which scenario.
Real-time Inference keeps instances always running for immediate responses. Use when you need consistently low latency (p50 < 100ms) or predictable 24/7 traffic. Supports auto-scaling but incurs costs even when idle.
Serverless Inference provisions compute only when requests arrive. Has cold start latency (several seconds), making it unsuitable for ultra-low-latency requirements. Cost-effective for sporadic or low-average-traffic workloads.
Asynchronous Inference queues requests in S3 and stores results in S3 when complete. Use for long-running inference (1-15 minutes) or large payloads (up to 1GB). Completion notifications via SNS or S3 events. Ideal for document processing or video analysis pipelines.
Batch Transform runs inference on large S3 datasets without a persistent endpoint. Cost-effective for periodic batch jobs like generating embeddings for an entire catalog or classifying large datasets.
Bedrock API: InvokeModel vs Converse
Bedrock offers two invocation patterns. uses each model's native request/response schema — giving you full control over model-specific parameters but requiring schema changes when switching models. provides a unified interface across all Bedrock models with standardized multi-turn conversation, system prompts, image inputs, and tool use.
For new projects, start with API. It abstracts model-specific differences, making it easier to experiment with different models and swap them later without changing application code. Only drop down to when you need features that does not yet expose.
Streaming Responses
Streaming LLM responses token-by-token dramatically improves perceived performance. Bedrock supports streaming via and . Both return Server-Sent Events (SSE) that your application processes as they arrive.
Key implementation considerations:
API Gateway has a 29-second integration timeout by default. If streaming responses can exceed 30 seconds, use WebSocket API or Lambda direct invocation instead. Lambda streaming requires and Lambda Web Adapter. AppSync GraphQL subscriptions are another option for real-time streaming in mobile/web applications.
API Gateway + Lambda for GenAI APIs
The standard pattern for exposing Bedrock calls as a managed API is API Gateway + Lambda. HTTP API (not REST API) is recommended for most GenAI use cases — lower latency and lower cost. Lambda invokes Bedrock and returns results. Lambda execution role needs permission.
Authentication options:
Cognito User Pool: JWT-based user authentication for external user-facing APIs Lambda Authorizer: Custom token validation logic (API keys, external IdP) IAM Authorization: Service-to-service communication with SigV4 signing
Add API Gateway Usage Plans to set per-API-key request quotas, preventing any single consumer from exhausting your Bedrock account-level throughput limits.
Event-Driven Architecture
Not all AI processing needs real-time responses. For email summarization, document classification, or nightly batch analysis, event-driven async architectures are more cost-effective. The standard pattern: S3 upload → S3 event notification → SNS → SQS → Lambda (Bedrock invocation) → S3 result → SNS completion notification.
SQS as a buffer is critical: without it, a burst of S3 upload events can hit Lambda concurrency limits or Bedrock throttling limits simultaneously. SQS smooths the processing rate through concurrency and batch size settings. Configure a Dead Letter Queue (DLQ) for failed messages to prevent infinite retry loops and enable later reprocessing.
EventBridge enables event-driven GenAI pipelines triggered by events from AWS services and SaaS partners — for example, automatically generating a reply draft when a new Salesforce case is created.
!Event-driven GenAI processing pipeline
Error Handling and Retry Strategies
| Error Type | Cause | Strategy | |-----------|-------|---------| | ThrottlingException | Request rate exceeded | Exponential backoff with jitter | | ModelTimeoutException | Response timeout | Retry or fall back to faster model | | ValidationException | Invalid request | Fix request — do not retry blindly | | ServiceUnavailableException | Transient error | Retry (up to 3 times recommended) |
AWS SDK auto-retries retriable errors with exponential backoff by default. is not retriable — the same invalid request will fail again. Build fallback strategies: on primary model failure, fall back to a smaller/faster model, return a cached response, or gracefully degrade with a user message.
Amazon Q Developer
Amazon Q Developer (formerly CodeWhisperer) is an AI coding assistant integrated directly into IDEs. For the AIP-C01 exam, it represents the "developer productivity" category of AI tools.
Key capabilities:
Real-time code suggestions based on comments and function signatures Security vulnerability scanning against AWS security best practices Open-source reference tracking for license compliance Natural language to code generation
The Pro plan enables customization — training Q Developer on your organization's internal codebase for more relevant suggestions. The exam frames Amazon Q Developer as the answer to "reduce time writing AWS SDK code" or "automatically detect security vulnerabilities in code."
Exam Quick Reference
"Fully managed FM API, per-token billing" -- Amazon Bedrock "Custom models, direct instance control" -- Amazon SageMaker "Always-on, ultra-low latency" -- SageMaker Real-time Endpoint "No cost when idle, cold start acceptable" -- SageMaker Serverless Inference "Long processing time, large payload" -- SageMaker Async Inference "Batch inference on S3 dataset" -- SageMaker Batch Transform "Unified interface for all models" -- Converse API "Model-native schema control" -- InvokeModel API "Real-time token streaming" -- InvokeModelWithResponseStream / ConverseStream "Expose GenAI API externally" -- API Gateway + Lambda "Async GenAI processing queue" -- SQS + Lambda + Bedrock "Throughput exceeded error" -- ThrottlingException → exponential backoff retry "AI coding assistant with security scanning" -- Amazon Q Developer