AIP-C01: FM Model Selection and Solution Design

AIP-C01 exam essentials — compare Amazon Bedrock model catalog options, tune inference parameters, and understand the trade-offs between Provisioned Throughput and on-demand inference.

The AIP-C01 exam tests your ability to map business requirements to FM characteristics — not just memorizing model names. This guide covers the Amazon Bedrock model catalog, inference parameter tuning, and Provisioned Throughput selection criteria from a practical engineering perspective.

 

Analyzing GenAI Solution Requirements

Before selecting a model, break down requirements along three axes: input/output modality, the cost-latency-quality trade-off triangle, and context window size.

Modality determines feasibility. If you need image analysis, a text-only model simply cannot do the job. If the use case is text-only, using a multimodal model wastes budget.

The trade-off triangle is always present. Real-time chatbots prioritize latency. Overnight batch report generation prioritizes quality and cost. Clearly defining which constraint dominates before model selection prevents costly mistakes.

!The Cost-Latency-Quality trade-off

Context window size is decisive for long document workloads. Claude 3.5 Sonnet supports 200K tokens, enabling a 300-page contract to be analyzed in a single API call without chunking.

 

Amazon Bedrock Model Catalog

Amazon Bedrock is a fully managed service that provides access to FMs from multiple vendors through a single API, with no infrastructure management required.

Anthropic Claude Family

Claude is the most widely used model family on Bedrock. The family spans from ultra-low-latency Haiku (real-time classification, keyword extraction) to the highest-quality Opus (complex legal and medical reasoning). Claude 3.5 Sonnet offers the best balance of performance and cost for code generation and agentic tasks. All Claude 3 models support a 200K token context window.

Amazon Titan Family

AWS-native models optimized for integration within the AWS ecosystem. Titan Embeddings V2 is the default embedding model in Amazon Bedrock Knowledge Bases. Titan Image Generator handles text-to-image generation. Titan Multimodal Embeddings supports joint image and text indexing.

Meta Llama Family

Open-source foundation with low fine-tuning costs. Llama 3 (8B/70B) and Llama 3.1 (up to 405B) offer strong multilingual and code capabilities. Best suited for domain-specific fine-tuning scenarios.

Mistral Family

European provider with efficient architectures. Mixtral 8x7B uses a Mixture of Experts (MoE) architecture, delivering ~70B-class quality at a fraction of the inference cost. MoE activates different expert sub-networks per token, reducing active parameter count during inference.

Cohere and Stability AI

Cohere Command excels at multilingual business workflows. Cohere Embed provides high-quality multilingual embeddings. Stable Diffusion XL is the go-to for high-quality image generation.

 

Inference Parameters

Temperature

Controls output diversity. Low (0.0–0.2) for deterministic tasks like code generation or classification. High (0.8–1.0) for creative brainstorming or marketing copy. The mid-range (0.3–0.7) works well for most conversational and summarization use cases.

top-p (Nucleus Sampling)

Samples from the smallest set of tokens whose cumulative probability mass reaches p. At top-p=0.9, only tokens covering the top 90% of probability mass are considered. Avoid setting both temperature and top-p to extremes simultaneously — the combined effect can be overly restrictive.

top-k

Restricts sampling to the top-k most probable tokens at each step. Useful for suppressing rare tokens in large-vocabulary models.

max_tokens

The single most direct cost control parameter. Set conservatively for chat responses (256–512 tokens) and more generously for code generation (1024–4096 tokens).

stop_sequences

Halts generation when a specified string is produced. Invaluable for structured output parsing — setting as a stop sequence when expecting JSON prevents runaway generation after the payload closes.

 

Bedrock Model Evaluation Job

Model Evaluation provides two modes: automatic evaluation using built-in metrics (Accuracy, Robustness, Toxicity, BERTScore, ROUGE) against labeled datasets in S3, and human evaluation using Amazon SageMaker Ground Truth Plus or AWS Marketplace labeling partners. Use human evaluation for brand tone alignment or subtle quality differences that automated metrics miss.

 

Provisioned Throughput vs On-Demand

On-demand charges per token with no reservation, suitable for unpredictable or low-volume traffic. Provisioned Throughput reserves Model Units (MU) at an hourly rate, guaranteeing throughput — essential for production workloads with SLA requirements and for fine-tuned custom models, which cannot be deployed on on-demand mode.

6-month commitments for Provisioned Throughput are approximately 40% cheaper than 1-month commitments. The break-even point versus on-demand is typically around 60–70% sustained utilization of the reserved capacity.

 

Cross-Region Inference Profile

Inference Profiles logically group model endpoints across multiple AWS regions behind a single endpoint. System-defined profiles are AWS-preconfigured; application inference profiles let you specify region lists and routing weights. Primary use cases: capacity overflow routing, geographic proximity optimization for global users, and multi-region disaster recovery failover.

Back to blog list