A data scientist trains a linear regression model to predict housing prices. The model achieves a low training error but a high test error. Which concept does this BEST illustrate?
Trap 1: Bias-variance tradeoff
The bias-variance tradeoff describes the general tension between underfitting and overfitting, not the specific diagnosis of a model that fits training data well and generalises poorly. It is tempting as the umbrella concept, but the stem's low training error and high test error names variance directly.
Trap 2: Regularization
Regularization deliberately penalises model complexity to reduce variance, so it would lower the gap rather than cause it. It is tempting because it addresses overfitting, but the stem asks which concept the observed behaviour illustrates, and no penalty term is described in the training process.
Trap 3: Underfitting
Underfitting produces high error on both training and test data, because the model is too simple to capture the underlying pattern. It is tempting as the opposite failure mode, but the stem's low training error with high test error indicates the model memorised the training set.
- A
Bias-variance tradeoff
Why it fails: The bias-variance tradeoff describes the general tension between underfitting and overfitting, not the specific diagnosis of a model that fits training data well and generalises poorly. It is tempting as the umbrella concept, but the stem's low training error and high test error names variance directly.
- B
Regularization
Why it fails: Regularization deliberately penalises model complexity to reduce variance, so it would lower the gap rather than cause it. It is tempting because it addresses overfitting, but the stem asks which concept the observed behaviour illustrates, and no penalty term is described in the training process.
- C
Underfitting
Why it fails: Underfitting produces high error on both training and test data, because the model is too simple to capture the underlying pattern. It is tempting as the opposite failure mode, but the stem's low training error with high test error indicates the model memorised the training set.
- D
Overfitting
Low training error with high test error means the model has memorised training noise rather than learning generalisable patterns, so it fails on unseen data. Overfitting is the specific term for this train-test performance gap, which the stem describes directly.