AIF-C01 covers RAG architecture and FM app design at approximately 24% of the exam. Key topics: RAG, Bedrock Knowledge Bases, Bedrock Agents, vector databases, and inference parameters.
---
What Is RAG? — The Library Analogy
RAG stands for Retrieval-Augmented Generation. Before answering, the AI first searches for relevant documents, then uses those as the basis for its response.
Think of a librarian. When asked a question, they don't guess — they find the relevant books first, then answer based on what they read. RAG works the same way.
| Problem | Without RAG | With RAG | |---------|-------------|----------| | Hallucination | AI fabricates plausible-sounding facts | Answer grounded in real documents | | Knowledge cutoff | Unaware of events after training | Can search latest documents in real time | | Domain knowledge | Only general knowledge | Can use internal company documents | | Cost | Requires expensive fine-tuning to add knowledge | No fine-tuning needed |
---
How RAG Works — 5 Steps
Step 1 — User submits a question
Example: "What is our company's vacation policy?"
Step 2 — Question is converted to an embedding (number vector)
Text becomes a numerical array that allows semantic comparison. "Vacation policy" and "leave regulations" produce similar vectors.
Step 3 — Vector database is searched for similar chunks
Documents stored in advance as embeddings are compared against the query. The most relevant chunks are retrieved.
Step 4 — Retrieved documents plus the original question are sent to the FM
The FM receives: "Using the following document as context, answer the question: [policy content]... Question: What is the vacation policy?"
Step 5 — FM generates a grounded answer
Because the answer is based on actual documents, hallucination is significantly reduced.
!How RAG works in 5 steps
Vector Databases
A regular database finds exact matches. A vector database finds semantically similar content.
| Service | Key trait | Best when | |---------|-----------|-----------| | Amazon OpenSearch Serverless | Managed, serverless, default for Bedrock Knowledge Bases | Default choice for most use cases | | Aurora PostgreSQL (pgvector) | Integrates with existing RDS workloads | Already using RDS | | Amazon Neptune Analytics | Combines graph DB with vector search | Relational + vector data needed | | Redis OSS (ElastiCache) | In-memory, ultra-fast | Sub-millisecond response required |
---
Bedrock Knowledge Bases — Automated RAG Pipeline
Bedrock Knowledge Bases removes the complexity of building RAG manually. Upload documents to S3, connect Knowledge Bases, and the service handles chunking, embedding, storage, and retrieval automatically.
Chunking strategies:
| Strategy | Description | Best for | |----------|-------------|----------| | Fixed size | Splits at a consistent token count | General text documents | | Hierarchical | Maintains parent-child chunk structure | Long, structured documents | | Semantic | Splits at meaningful boundaries | Quality-critical use cases |
---
Bedrock Agents — AI That Takes Action
Standard chatbots only answer questions. Bedrock Agents can interact with external systems and complete multi-step tasks autonomously.
Example: "Analyze last month's sales and email the summary to the team lead."
An Agent would: call the database API → analyze the data → write the summary → send the email via API — all without human intervention at each step.
| Concept | Description | |---------|-------------| | Action Groups | Set of APIs and functions the Agent can call | | Tool Use | FM's ability to invoke external tools directly | | Memory | Retains conversation history for context continuity | | Orchestration | Automatically plans and sequences multi-step tasks |
---
Inference Parameters
| Parameter | Effect | Low value | High value | |-----------|--------|-----------|------------| | Temperature | Creativity vs consistency | Predictable, consistent answers | Creative, varied answers | | Top-p | Range of token candidates | Only top-probability tokens | Wider candidate range | | Max tokens | Maximum response length | Short answers | Long answers |
---
Exam Key Points
"Search external docs to ground FM responses" — RAG "Convert text to numeric vectors for similarity search" — Embedding "Database optimized for similarity search" — Vector database "Default vector DB for Bedrock Knowledge Bases" — Amazon OpenSearch Serverless "Automated end-to-end RAG pipeline" — Bedrock Knowledge Bases "FM calls external APIs to complete multi-step tasks" — Bedrock Agents "Controls creativity vs consistency in responses" — Temperature RAG advantages: reduces hallucination + current knowledge + domain specialization