Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A machine learning engineer is preparing a dataset for a natural language processing task. The dataset contains text reviews with varying lengths, and the engineer plans to use a transformer model. Which preprocessing step is most critical to ensure the model can handle the input effectively?

⚠ Common exam trap

The trap here is focusing on traditional text normalization steps like lowercasing or stemming while overlooking the transformer-specific need for uniform sequence lengths via tokenization and padding.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Tokenize the text into subword units and pad or truncate sequences to a fixed maximum length.

Transformer models require fixed-length input sequences, so tokenizing into subword units and padding/truncating to a uniform length is essential. This enables batch processing and handles out-of-vocabulary words. Other steps like lowercasing, one-hot encoding, or stemming are either not critical or incompatible with transformer input expectations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert all text to lowercase and remove punctuation to reduce vocabulary size.

    Why it's wrong here

    While lowercasing and punctuation removal can reduce vocabulary, they may discard useful information such as sentiment cues (e.g., exclamation marks) and proper nouns. Transformers can learn from case and punctuation. This step is not critical and may harm performance; the critical step is handling variable sequence lengths through tokenization and padding.

  • ✗

    Perform stemming or lemmatization to reduce words to their base forms.

    Why it's wrong here

    Stemming or lemmatization can reduce vocabulary but may lose nuanced meaning. Transformer models, especially those pretrained on raw text, can handle inflected forms and benefit from subword tokenization. This step is not critical and may even degrade performance by removing morphological information that the model could learn from.

  • ✓

    Tokenize the text into subword units and pad or truncate sequences to a fixed maximum length.

    Why this is correct

    Transformer models require input sequences of uniform length within a batch. Tokenization into subword units (e.g., WordPiece or BPE) handles out-of-vocabulary words and reduces vocabulary size. Padding shorter sequences and truncating longer ones to a fixed maximum length ensures batch processing. This is essential for efficient training and inference with transformers.

  • ✗

    Apply one-hot encoding to each word in the vocabulary to create binary vectors.

    Why it's wrong here

    One-hot encoding produces high-dimensional, sparse vectors and does not capture semantic relationships. Transformer models expect dense embeddings and typically use learned token embeddings. One-hot encoding is impractical for large vocabularies and incompatible with the model's input requirements. It is not a critical preprocessing step for transformers.

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.