Courseiva
hardMultiple Choice

AIF-C01 Practice Question: Building a real-time document analysis tool using…

A company is building a real-time document analysis tool using Amazon Bedrock. Their documents average 15,000 tokens each. Users submit a document and ask a single question about it. The team wants to minimize latency while maintaining answer quality. Which approach is MOST suitable?

⚠ Common exam trap

The AIF-C01 exam often tests the misconception that RAG or chunking is always necessary for large documents, but the trap here is that 15,000 tokens is well within the context window of modern foundation models, making direct prompting the most latency-efficient and quality-preserving approach for single-document Q&A.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Amazon Bedrock's Converse API with a prompt that includes the full document text and the user's question in a single invocation.

Using Amazon Bedrock's Converse API with a prompt that includes the full document text and the user's question in a single invocation minimizes latency by avoiding the overhead of multiple API calls or retrieval steps. This approach maintains answer quality because the model has access to the entire document context (up to 15,000 tokens) in one pass, which is well within the context window of models like Claude 3 or Llama 2, ensuring accurate responses without the need for chunking or external retrieval.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Amazon Bedrock's Knowledge Base (RAG) to retrieve relevant chunks from the document and then answer the question.

    Why it's wrong here

    Streaming does not solve the context window limitation.

  • ✗

    Split the document into chunks of 1,000 tokens each, process each chunk with the model separately, and aggregate results.

    Why it's wrong here

    This limits the document to 8K tokens, requiring truncation or chunking.

  • ✓

    Use Amazon Bedrock's Converse API with a prompt that includes the full document text and the user's question in a single invocation.

    Why this is correct

    A single Converse invocation passes the full 15,000-token document and question together, avoiding the extra retrieval round trip and embedding lookup that a RAG pipeline would add. With one question per document, prompt-stuffing satisfies the latency constraint directly while preserving answer quality.

  • ✗

    Use a fine-tuned model that has been trained on similar documents to avoid context processing.

    Why it's wrong here

    This would require chunking and retrieval, adding latency.

About these practice questions

This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.