RAG (Retrieval-Augmented Generation) is the architectural pattern for injecting knowledge that wasn't in the FM's training data — internal documents, post-training events, proprietary data — into model responses at inference time. This guide covers the entire RAG pipeline from data preprocessing to retrieval optimization, from a practical engineering standpoint.
FM Input Data Preprocessing
RAG quality is determined far more at the data quality stage than at the retrieval stage. Key preprocessing steps include: noise removal (headers, footers, page numbers, ads), encoding normalization (UTF-8, special characters), structure preservation (tables, code blocks, lists), and deduplication across document versions. AWS Glue is the standard choice for batch ETL pipelines: S3 raw bucket → Glue Job (extract, clean, enrich) → S3 clean bucket → Bedrock Knowledge Base sync.
!Indexing Stage (Offline) pipeline diagram
For real-time document streams, use Kinesis Data Streams → Lambda preprocessing → S3. For complex multi-step pipelines with branching logic, Step Functions orchestrates PDF extraction → language detection → translation → embedding → vector DB insertion with built-in retry and error handling.
Vector Database Selection
Amazon OpenSearch Serverless
The default vector store for Bedrock Knowledge Bases. Serverless auto-scaling eliminates capacity planning. Supports ANN-based vector search, hybrid search (BM25 + vector), and full-text search in a single engine. Minimum OCU cost applies even with zero traffic.
Amazon Aurora PostgreSQL + pgvector
Add the extension to an existing Aurora cluster. Best when you need SQL-based combined vector + relational filtering, or when reusing existing RDS infrastructure. Slower than OpenSearch at very large scale (hundreds of millions of vectors).
Amazon Neptune Analytics
Graph database with vector search integration. Use when queries require both semantic similarity and graph traversal — e.g., "find papers similar to this one, written by authors connected to this institution."
Third-party: Pinecone, Redis Enterprise
Bedrock Knowledge Bases also supports Pinecone and Redis Enterprise Cloud. Suitable when these are already in use in your infrastructure.
Embedding Models
Titan Embeddings V2 supports variable dimensions (256/512/1024) using Matryoshka embeddings, allowing storage and speed optimization with minimal quality loss. Cohere Embed Multilingual covers 100+ languages and is the clear choice for multilingual RAG. Always match the embedding model used at index time with the one used at query time — dimension and normalization must be consistent.
Chunking Strategies
Fixed-size chunking is simplest: split by token count (typically 256–512 tokens) with a 50–100 token overlap. Fast and consistent but may break semantic units mid-sentence.
Semantic chunking splits at semantic boundaries by computing cosine similarity between adjacent sentence embeddings and splitting where similarity drops sharply. Higher retrieval quality but more computationally expensive and produces variable-size chunks.
Hierarchical chunking (supported natively by Bedrock Knowledge Bases) creates parent chunks (full sections for context) and child chunks (granular paragraphs for precision). At retrieval time, child chunks identify the relevant passage; the parent chunk is passed to the FM as context. This "small-to-big retrieval" pattern significantly improves coherence of FM responses.
Amazon Bedrock Knowledge Bases
Bedrock Knowledge Bases provides a fully managed RAG indexing and retrieval layer. Data source connectors (S3, web crawler, Confluence, SharePoint, Salesforce) handle ingestion. Chunking strategy and embedding model are configurable. The API performs retrieval and generation in a single call. The API returns raw retrieved chunks for custom generation logic. Metadata filtering narrows the search space — setting a category filter reduces irrelevant results and improves precision.
!Retrieval and Generation Stage pipeline diagram
Advanced Retrieval: Hybrid Search, Reranking, Query Rewriting
Hybrid search combines vector similarity (semantic) with BM25 keyword matching, fused via Reciprocal Rank Fusion (RRF). This covers both precise keyword queries and broad semantic intent queries robustly.
Reranking uses a cross-encoder model (Cohere Rerank on Bedrock) to re-evaluate the top 20–50 initial results and select the top 5–10 most relevant for FM context injection. Cross-encoders jointly encode the query and document, providing higher accuracy than bi-encoder similarity alone.
Query rewriting expands the original user query using an FM to improve retrieval recall. "How to reduce AWS costs?" expands to "AWS cost optimization, Reserved Instances, Savings Plans, Spot Instances, right-sizing" before embedding.