A data engineer is implementing a data classification process using Snowflake's Data Classification feature. The engineer wants to automatically classify columns containing sensitive data and then use the results to apply masking policies. After running the classification, the engineer notices that some columns that should be classified as 'EMAIL' are not being tagged. The engineer has verified that the data contains valid email addresses. What is the most likely reason for the missing classification?
Data Classification samples a limited number of rows to infer the semantic category. If the sample does not contain a sufficient number of email addresses (e.g., due to low frequency or sampling randomness), the column may not be classified as EMAIL. This is a common reason for missed classifications. The engineer can adjust the sampling or manually tag the column. This option correctly identifies the sampling limitation as the likely cause.
Why this answer
Data Classification uses sampling to analyze column data. If the sample does not contain a sufficient number of email addresses, the column may not be classified as EMAIL. This is a known limitation.
The other options are incorrect because classification does not depend on column names, can handle NULLs, and is not restricted to a specific schema.
Exam trap
The trap here is assuming that Data Classification scans all rows or relies on column names, when it actually samples data and uses pattern recognition.