Courseiva
mediumMultiple Choice

AIF-C01 Practice Question: A company has a large dataset of customer support…

A company has a large dataset of customer support emails labeled with issue categories. They need to classify new emails automatically. Which algorithm is BEST suited for this task?

⚠ Common exam trap

AWS certification exams often test the distinction between regression and classification by placing linear regression as a distractor, exploiting the common misconception that 'regression' implies any predictive modeling, when in fact it is only for continuous outputs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Logistic regression

Logistic regression is the best choice because it is a supervised learning algorithm specifically designed for binary or multi-class classification tasks. Given labeled emails with issue categories, logistic regression models the probability that a new email belongs to each category using a logistic (sigmoid) function, making it ideal for this classification problem.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Linear regression

    Why it's wrong here

    Linear regression predicts continuous values, not categories.

  • ✗

    K-means clustering

    Why it's wrong here

    K-means is unsupervised and not appropriate for labeled classification.

  • ✓

    Logistic regression

    Why this is correct

    Logistic regression is used for classification, including multiclass via softmax.

  • ✗

    Gradient boosting

    Why it's wrong here

    Gradient boosting is a powerful ensemble method renowned for its high accuracy in various classification and regression tasks, particularly with structured, tabular data. This makes it tempting, as it excels at finding complex patterns. However, it is not best suited for directly classifying raw, unstructured text data like customer support emails. Gradient boosting models require numerical input features, meaning the text would first need extensive pre-processing and vectorisation (e.g., TF-IDF or word embeddings) before the algorithm could be applied, adding a significant, separate step not inherent to the algorithm itself. It would be an excellent choice if the email data had already been transformed into numerical features.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.