mediumMultiple Select
AIF-C01 Practice Question: A machine learning engineer is implementing a RAG…
A machine learning engineer is implementing a RAG system using Amazon Bedrock and a vector database. They need to chunk a large set of PDF documents before embedding. Which THREE considerations are important for chunking strategy? (Select THREE.)
⚠ Common exam trap
AWS AI Practitioner exams often test the distinction between constraints (like model limits) and active strategy choices (like chunk size and overlap), so candidates mistakenly select the embedding model's maximum input length as a chunking strategy consideration instead of recognizing it as a boundary condition.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Overlap between consecutive chunks
Options A, B, and C are the three core design parameters of any chunking strategy. A is correct because adding overlap between consecutive chunks preserves context that would otherwise be lost at chunk boundaries, improving retrieval quality when a relevant passage spans two chunks. B is correct because the splitting method (sentence, paragraph, semantic, or fixed token count) determines how coherently meaning is preserved within each chunk and directly affects embedding quality. C is correct because chunk size in tokens controls the trade-off between retrieval granularity and the amount of context each embedding captures; too small loses context, too large dilutes relevance. Option D is not a chunking-strategy consideration per se — the embedding model's maximum input length is a hard constraint that caps chunk size, but it is a model property rather than a strategic choice. Option E is about document metadata and filtering, which relates to retrieval and indexing, not to how documents are split into chunks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Overlap between consecutive chunks
Why this is correct
Overlap between consecutive chunks preserves context across boundaries, so sentences split mid-thought still appear intact in at least one chunk. This improves retrieval quality in the RAG pipeline by preventing meaning loss at chunk edges.
- ✓
The method used to split chunks (e.g., by sentence, paragraph, or fixed token count)
Why this is correct
The splitting method determines chunk boundaries: sentence or paragraph splitting preserves semantic units, whereas fixed token counts ignore structure. This choice directly affects retrieval relevance, making it a core chunking consideration for the Bedrock RAG pipeline.
- ✓
Chunk size (e.g., number of tokens)
Why this is correct
Chunk size directly governs retrieval granularity: oversized chunks dilute embedding specificity and waste context-window tokens, while undersized chunks fragment semantic meaning across boundaries. Since the stem requires chunking before embedding, token count per chunk determines what each vector represents, making it a primary design consideration for the RAG pipeline.
- ✗
The embedding model's maximum input length
Why it's wrong here
The embedding model's maximum input length constrains chunk size, so chunks exceeding the token limit are truncated and lose content. It is tempting because it is a genuine constraint, but it is a hard limit to respect rather than a strategy consideration; chunk size, overlap, and boundaries are the actual decisions.
- ✗
The metadata associated with each document
Why it's wrong here
Chunking strategy turns on chunk size, overlap, and splitting boundaries such as headings or paragraphs; document metadata is stored alongside chunks for filtering and citation, but it does not determine how text is divided. It is tempting because metadata improves retrieval, yet chunking decisions concern segmentation, not metadata.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.