A data engineer is designing a pipeline to train a linear regression model on a dataset with 10 million rows and 50 features. The dataset fits in memory. Which approach should the engineer use to train the model efficiently?
Trap 1: Normal equation
Normal equation requires computing (X^T X)^{-1}, which is computationally expensive for large datasets.
Trap 2: Batch gradient descent
Batch gradient descent uses the whole dataset for each update, which is slow for large datasets.
Trap 3: Principal component analysis
PCA reduces dimensionality but does not train a model.
- A
Normal equation
Why it fails: Normal equation requires computing (X^T X)^{-1}, which is computationally expensive for large datasets.
- B
Batch gradient descent
Why it fails: Batch gradient descent uses the whole dataset for each update, which is slow for large datasets.
- C
Principal component analysis
Why it fails: PCA reduces dimensionality but does not train a model.
- D
Stochastic gradient descent
SGD updates weights per sample, making it efficient for large datasets.