Courseiva
easyMultiple Choice

AIF-C01 Practice Question: Which AWS service is BEST suited for extracting…

Which AWS service is BEST suited for extracting text from scanned PDF documents, such as invoices and receipts?

⚠ Common exam trap

The trap is that candidates often confuse Amazon Rekognition's text-in-image capability with Amazon Textract's document-specific extraction. However, Rekognition lacks the ability to extract text from structured documents like invoices and receipts, as it is designed for image analysis rather than document processing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Amazon Textract

Amazon Textract is specifically designed to extract text, handwriting, and data from scanned documents like invoices and receipts. It uses machine learning to recognize and extract printed text, forms, and tables from document images, going beyond simple OCR by understanding the structure of the document.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Amazon Comprehend

    Why it's wrong here

    Amazon Comprehend analyses existing text for entities, sentiment and key phrases; it cannot perform optical character recognition on scanned images. It is tempting because invoice extraction ultimately yields structured entities, but Comprehend requires text already extracted — Amazon Textract supplies that OCR step first.

  • ✓

    Amazon Textract

    Why this is correct

    Amazon Textract uses optical character recognition with layout analysis to extract text, forms and tables from scanned documents such as invoices and receipts. This directly satisfies the requirement for extracting text from scanned PDFs, unlike generic storage or compute services.

  • ✗

    Amazon Rekognition

    Why it's wrong here

    Amazon Rekognition performs image and video analysis — object, face and label detection — and does not extract text from documents. It is tempting because scanned PDFs are images, but Rekognition's label detection returns categories, not transcribed invoice fields; Amazon Textract is purpose-built for document text and form extraction.

  • ✗

    Amazon Transcribe

    Why it's wrong here

    Amazon Transcribe converts speech audio into text and accepts no PDF input, so it cannot read scanned invoices. The temptation is that both Transcribe and Textract perform extraction, but the axis of difference is input modality: Transcribe handles audio streams, whereas Textract handles document images and PDFs.

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.