Courseiva
hardMultiple Choice

AIF-C01 Practice Question: Building a RAG application that indexes thousands…

A company is building a RAG application that indexes thousands of PDF documents. They notice that some documents are very long (hundreds of pages) and the vector search often returns irrelevant chunks. Which configuration change would MOST improve retrieval relevance?

⚠ Common exam trap

AIF-C01 often tests the misconception that retrieval quality is improved by upgrading the model or vector database, when the actual root cause is usually data preparation — chunking, embedding model choice, or metadata filtering.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Adjust the chunk size and overlap to better capture context from the documents

Chunk size and overlap directly control how much semantic context each embedded vector carries. When documents are hundreds of pages long, a fixed chunk size (e.g., 512 tokens) can split related sentences across chunks, producing vectors that represent fragments rather than complete ideas. Tuning chunk size and overlap ensures each chunk is semantically self-contained, which is the single most impactful lever for retrieval precision in a RAG pipeline.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch from Amazon OpenSearch Serverless to Pinecone

    Why it's wrong here

    Swapping vector stores changes hosting and cost characteristics, not chunk boundaries; irrelevant chunks persist because long PDFs still yield coarse segments. The tempting appeal is that a different engine may offer tuning knobs, but relevance here depends on chunking strategy. Pinecone would be the right choice when managed scaling or hybrid search is the actual requirement.

  • ✗

    Increase the embedding dimension from 1024 to 4096

    Why it's wrong here

    Raising embedding dimensions enlarges each vector but does not separate semantically distinct passages inside a hundreds-page PDF; irrelevant chunks still match. It is tempting because higher dimensions are marketed as greater fidelity. Increasing dimensions would be correct when the corpus is short-form and subtle semantic distinctions are being lost.

  • ✗

    Use a larger, more capable foundation model for response generation

    Why it's wrong here

    A larger foundation model only rewrites the answer from whatever chunks retrieval supplies; it cannot change which chunks are returned. It is tempting because output quality visibly improves, masking poor retrieval. Upgrading the generator would be correct when answers are fluent but shallow despite accurate context being retrieved.

  • ✓

    Adjust the chunk size and overlap to better capture context from the documents

    Why this is correct

    Long PDFs produce chunks that mix unrelated topics, so embeddings drift and retrieval returns irrelevant passages. Reducing chunk size with sensible overlap keeps each vector semantically focused, directly improving relevance without changing the embedding model or index.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.