Courseiva
Implementing AI Solutions →mediumMultiple Select

AI0-001 Implementing AI Solutions Practice Question

An organisation is developing a document intelligence system that extracts information from scanned invoices. Which THREE data preparation steps are critical to ensure high extraction accuracy? (Choose THREE.)

⚠ Common exam trap

AI0-001 often tests the distinction between general NLP preprocessing steps (like stopword removal and lowercasing) and document-specific preprocessing (like OCR correction and image enhancement), causing candidates to incorrectly select B or C as critical for extraction accuracy.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cleaning and correcting OCR output

Option A (Cleaning and correcting OCR output) is correct because OCR on scanned invoices inevitably introduces character-level errors, and fixing those errors before feeding text to the extraction model directly improves field-level accuracy. Option D (Annotating bounding boxes and field labels) is correct because supervised document intelligence models need labelled ground truth that ties specific fields (e.g., invoice number, total) to their spatial locations to learn accurate extraction. Option E (Image preprocessing such as deskewing and binarisation) is correct because scanned invoices often suffer from rotation, noise, and uneven lighting, and correcting these at the image level raises OCR quality and downstream extraction accuracy. Option B (Removing punctuations and stopwords) is not appropriate because invoice fields such as dates, currency amounts, and vendor names rely on punctuation and specific tokens, so removing them would destroy critical information. Option C (Normalising all text to lowercase) is not appropriate because case can carry meaning in invoice data (e.g., currency codes, product identifiers, proper names), and lowercasing everything can reduce extraction fidelity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Cleaning and correcting OCR output

    Why this is correct

    Cleaning and correcting OCR output directly addresses the scanned-invoice constraint: OCR introduces character errors on low-quality scans, and those errors propagate into extraction models. Normalising recognised text before training or inference raises accuracy, since the system's inputs are images rather than typed digital text.

  • ✗

    Removing punctuations and stopwords

    Why it's wrong here

    Stripping punctuation and stopwords removes delimiters such as colons, currency symbols and dates that anchor invoice fields, degrading extraction. This suits bag-of-words text classification, where such tokens add little, but not structured document field extraction.

  • ✗

    Normalising all text to lowercase

    Why it's wrong here

    Lowercasing destroys case cues that distinguish vendor names, codes and field labels on invoices, harming extraction accuracy. Normalisation is useful for matching free-text or training language models, but invoice extraction depends on preserving original casing and formatting.

  • ✓

    Annotating bounding boxes and field labels

    Why this is correct

    Annotating bounding boxes and field labels creates the ground-truth training data that supervised invoice extraction models learn field locations from. Without labelled coordinates for each key-value field, the model cannot learn to detect or extract them accurately.

  • ✓

    Image preprocessing (e.g., deskewing, binarisation)

    Why this is correct

    Deskewing and binarisation normalise scanned invoice images before optical character recognition, correcting rotation and separating text from background noise. This directly satisfies the stem's requirement for high extraction accuracy, since skewed or low-contrast scans cause character misrecognition that propagates into the extracted fields.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.