hardMultiple Select
AIF-C01 Practice Question: Designing a RAG pipeline for a legal document…
A company is designing a RAG pipeline for a legal document review system. They need to ingest hundreds of documents, create embeddings, and store them for retrieval. Which THREE steps are essential in the ingestion phase of the RAG pipeline?
⚠ Common exam trap
AIF-C01 often tests the components of RAG, and candidates might include fine-tuning or guardrails as part of ingestion. The trap is confusing training-time activities with ingestion-time activities.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Chunk the documents into smaller pieces
The ingestion phase of a RAG pipeline is responsible for preparing source documents so they can later be retrieved: option B (chunk the documents into smaller pieces) is essential because splitting long legal documents into appropriately sized chunks keeps each unit within the embedding model's token limit and improves retrieval granularity and relevance. Option D (generate embeddings for each chunk using an embedding model) is essential because RAG retrieval works by comparing the vector representation of a query against vector representations of the content, so every chunk must be converted into an embedding. Option E (store the embeddings in a vector store) is essential because the resulting vectors must be persisted in a vector database or index so that similarity search can be performed at query time. Option A (Bedrock Guardrails) is a runtime safety and content-filtering concern applied to model inputs/outputs, not a data ingestion step, and option C (fine-tuning the base model) is a separate model-customization activity that is not required to ingest documents into a RAG pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure Bedrock Guardrails for the application
Why it's wrong here
Bedrock Guardrails filter and moderate model inputs and outputs at inference time; they do not chunk, embed or index documents, so they sit outside the ingestion phase. It is tempting as a safety control for legal content, but guardrails belong to the retrieval and generation stages, not ingestion.
- ✓
Chunk the documents into smaller pieces
Why this is correct
Chunking splits lengthy legal documents into smaller passages before embedding, which is essential because embedding models have fixed token limits and retrieval precision degrades when a single vector represents an entire contract. Smaller chunks let the retriever return the specific clause relevant to a query, satisfying the ingestion requirement to prepare documents for accurate semantic search.
- ✗
Fine-tune the base model on legal documents
Why it's wrong here
Fine-tuning alters model weights using labelled examples; it does not chunk documents, generate embeddings or write vectors to a store, so it is not an ingestion step. It is tempting for domain adaptation in legal work, but RAG grounds answers through retrieved context, making fine-tuning unnecessary here.
- ✓
Generate embeddings for each chunk using an embedding model
Why this is correct
Each chunk must be passed through an embedding model to produce a dense vector representation. This converts text into the numeric form required for semantic similarity search, satisfying the ingestion phase's need to create embeddings.
- ✓
Store the embeddings in a vector store
Why this is correct
Storing embeddings in a vector store is essential because retrieval requires similarity search across the embedded corpus. The vector store indexes high-dimensional vectors, enabling nearest-neighbour queries that fetch relevant chunks for the generator. Without it, the ingested documents could not be retrieved, breaking the RAG pipeline's grounding constraint.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.