easyMultiple Select
AIF-C01 Practice Question: A data scientist is setting up a RAG pipeline…
A data scientist is setting up a RAG pipeline using Amazon Bedrock Knowledge Bases. They need to ingest documents, split them into chunks, generate embeddings, and store them for retrieval. Which TWO steps are part of the ingestion process? (Choose TWO)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Parse the documents and split them into chunks
Option D is correct because the ingestion (data source sync) process in Amazon Bedrock Knowledge Bases begins by parsing the source documents and splitting them into manageable chunks during the chunking step, which is required before embeddings can be created. Option E is correct because after chunking, the service invokes a selected Amazon Bedrock embedding model (for example, Amazon Titan Embeddings or Cohere Embed) to convert each chunk into a vector representation that is then written to the vector store. Option A is not part of the standard ingestion flow; while Lambda functions can be used with custom transformations or orchestration, defining a Lambda for post-processing is not a required ingestion step. Option B is incorrect because similarity search is a retrieval-time operation performed against the vector store, not part of ingestion. Option C is also incorrect because augmenting the user prompt with retrieved chunks happens during the retrieval/generation (query) phase, after ingestion is complete.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Define a Lambda function for post-processing
Why it's wrong here
Post-processing via Lambda runs after retrieval, not during ingestion; ingestion covers parsing, chunking, embedding and vector storage. It tempts because Lambda customisation is genuinely useful for transforming retrieved chunks or responses in a RAG pipeline, but that belongs to the query or orchestration stage, not the ingestion workflow the question asks about.
- ✗
Perform a similarity search on the vector store
Why it's wrong here
Similarity search queries the vector store at retrieval time, after ingestion has already embedded and stored the chunks. It tempts because retrieval is the visible purpose of the pipeline, yet the question asks only for ingestion steps; searching presupposes the embeddings already exist, so it cannot be part of building the index.
- ✗
Augment the user prompt with retrieved chunks
Why it's wrong here
Augmentation happens at query time, after retrieval, when the user prompt is combined with the chunks returned from the knowledge base. Ingestion covers parsing, chunking, embedding and storage only. Augmentation is the correct step when discussing the retrieval-and-generation half of the pipeline.
- ✓
Parse the documents and split them into chunks
Why this is correct
Parsing documents and splitting them into chunks is a core ingestion step: source files are read, text extracted, and content divided into manageable segments before embedding. This chunking stage satisfies the pipeline's requirement to prepare documents for vectorisation and subsequent retrieval.
- ✓
Generate embeddings for each chunk using a Bedrock embedding model
Why this is correct
Generating embeddings for each chunk using a Bedrock embedding model converts text segments into vectors, which are then written to the vector store. This is a required ingestion step, satisfying the pipeline's need to produce searchable representations for retrieval.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.