Courseiva
hardMultiple Choice

MLA-C01 Shape Mismatch Error Practice Question

Exhibit

{
    "TrainingJobName": "fraud-detection-model-20241015",
    "TrainingJobStatus": "Failed",
    "FailureReason": "AlgorithmError: Encountered an unexpected error during training: ValueError: Expected 2D array, got 1D array instead. Reshape your data using array.reshape(-1, 1) if your data has a single feature.",
    "AlgorithmSpecification": {
        "TrainingImage": "382416733822.dkr.ecr.us-west-2.amazonaws.com/sagemaker-scikit-learn:1.0-1-cpu-py3",
        "TrainingInputMode": "File"
    },
    "ResourceConfig": {
        "InstanceType": "ml.m5.large",
        "InstanceCount": 1
    },
    "InputDataConfig": [
        {
            "ChannelName": "training",
            "DataSource": {
                "S3DataSource": {
                    "S3DataType": "S3Prefix",
                    "S3Uri": "s3://my-bucket/train/data.csv",
                    "S3DataDistributionType": "FullyReplicated"
                }
            },
            "ContentType": "text/csv",
            "CompressionType": "None"
        }
    ]
}

Refer to the exhibit. A data scientist used a SageMaker training job with a custom Scikit-learn script. The training job failed with the error shown. What is the most likely cause of this failure?

⚠ Common exam trap

In AWS SageMaker, a shape mismatch error often occurs when the CSV file contains extra columns (e.g., an index column from pandas during save) that increase the feature count beyond what the model expects. This is different from missing values or content type issues.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The training script is reading the CSV file incorrectly, causing a shape mismatch.

The error 'shape mismatch' typically occurs when the number of columns in the CSV file does not match the number of features expected by the Scikit-learn model. In SageMaker, the training script often loads data using pandas or numpy, and if the CSV has extra columns (e.g., an index column, header row misinterpreted, or trailing delimiter), the feature matrix shape will be inconsistent with the model's input dimensions, causing a ValueError during fit or transform.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The training script is reading the CSV file incorrectly, causing a shape mismatch.

    Why this is correct

    The failure stems from the script reading the CSV without specifying that the first row contains column headers, so Scikit-learn treats header text as data and raises a value-conversion error. Supplying `header=0` (or `skiprows=1`) to pandas aligns the parsed shape with the model's expected feature count.

  • ✗

    The InputDataConfig specifies ContentType text/csv but the actual file is not CSV.

    Why it's wrong here

    A ContentType mismatch usually triggers a parsing or deserialisation error when SageMaker reads the channel, not the failure shown. The exhibit's traceback points elsewhere, so the declared MIME type is not implicated. This would be correct if the log showed a CSV parse failure.

  • ✗

    The SageMaker training image is outdated and does not support Scikit-learn 1.0.

    Why it's wrong here

    An outdated image would surface as a package or import error, not the specific exception in the exhibit. The failure stems from the script's code or data handling, so the image version is irrelevant. This would be correct only if the log showed a missing Scikit-learn module or version conflict.

  • ✗

    The training data contains missing values that need to be imputed.

    Why it's wrong here

    Missing values typically raise a ValueError inside the estimator or preprocessing step, producing a traceback naming that function. The exhibit's error does not reference NaN handling, so imputation is not the cause. This would be correct if the log showed NaN or null-related exceptions.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.