A data analyst is preparing a dataset for a machine learning model and notices that several numerical features have different scales (e.g., age ranges from 18-90, income from 20,000-200,000). The analyst decides to apply feature scaling. Which TWO of the following statements about scaling techniques are correct? (Choose two.)
Standardization subtracts the mean and divides by the standard deviation, resulting in a distribution with mean 0 and standard deviation 1. It does not bound values to a specific range, so outliers have less impact on the scaling of the majority of data compared to min-max normalization.
Why this answer
Min-max normalization scales to a fixed range but is sensitive to outliers, while standardization centers data with unit variance and is less affected by outliers. Both are common scaling techniques. The other statements misrepresent properties of scaling or incorrectly claim that distance-based algorithms are scale-invariant.
Exam trap
The trap here is confusing the properties of min-max normalization and standardization, especially regarding outlier sensitivity and output range.