Courseiva
hardMultiple Select

AIF-C01 Practice Question: An AWS AI practitioner is designing a document…

An AWS AI practitioner is designing a document processing pipeline using Amazon Textract and Amazon Comprehend. The pipeline must extract text from PDFs, detect entities, and classify documents into categories (e.g., invoice, contract, report). Which THREE steps should be included in the pipeline? (Choose three.)

⚠ Common exam trap

Watch out — candidates often confuse Amazon Personalize (a recommendation engine) with Amazon Comprehend (a natural language processing service) for classification tasks, and assuming Amazon Rekognition can extract text from PDFs when that is the role of Amazon Textract.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Amazon Comprehend to train a custom classifier for document type

Option C is correct because Amazon Textract is the AWS service purpose-built to extract text and structured data from PDFs and scanned documents via synchronous or asynchronous APIs such as DetectDocumentText and StartDocumentTextDetection. Option D is correct because Amazon Comprehend's entity detection (DetectEntities) natively identifies entities like DATE, QUANTITY, and other standard or custom entity types needed to pull dates and amounts from the extracted text. Option A is correct because Amazon Comprehend custom classification lets you train a custom classifier on labeled documents to categorize them into classes such as invoice, contract, or report, which is exactly the classification requirement. Option B is wrong because Amazon Personalize is a recommendation service for personalized user experiences, not document categorization. Option E is wrong because Amazon Rekognition performs image and video analysis (objects, faces, labels) and does not extract document text or classify document types.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Amazon Comprehend to train a custom classifier for document type

    Why this is correct

    Amazon Comprehend custom classification trains a model on labelled examples to assign documents to categories such as invoice, contract or report. This satisfies the stem's classification requirement, which entity detection alone cannot fulfil, since categorisation needs supervised training on document types.

  • ✗

    Use Amazon Personalize to recommend document categories

    Why it's wrong here

    Amazon Personalize builds recommendation systems from user interaction data, not document classification; it cannot categorise invoices or contracts from extracted text. It is tempting because it also uses machine learning on AWS, but its purpose is product or content recommendations for individual users, so it would be correct for a retail personalisation scenario instead.

  • ✓

    Use Amazon Textract to extract text from the PDF

    Why this is correct

    Amazon Textract uses optical character recognition to extract raw text and structured data from PDFs, satisfying the pipeline's first requirement. Downstream Comprehend entity detection and classification depend on this extracted text, so Textract must run before the analytics stages.

  • ✓

    Use Amazon Comprehend to detect entities such as dates and amounts

    Why this is correct

    Amazon Comprehend's entity detection identifies named entities including dates, amounts, people and organisations within extracted text. This satisfies the stem's entity-detection requirement, operating on the text Textract produced rather than performing extraction or document categorisation itself.

  • ✗

    Use Amazon Rekognition to analyze images in the PDF

    Why it's wrong here

    Amazon Rekognition performs image and video analysis such as object, face and label detection; it does not extract text or classify documents by category. It is tempting because PDFs can contain images, but Rekognition would be the right choice for detecting objects, faces or inappropriate content within those images, not for this text pipeline.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.