A data engineer is preparing a large corpus of support tickets stored in a Unity Catalog volume for fine-tuning a Llama model on Databricks. They must remove personally identifiable information (PII) before the data reaches the training cluster. Which TWO approaches are appropriate for detecting and redacting PII at scale in this pipeline? (Choose two.)
A pandas UDF running a PII detection library like Presidio or Spark NLP processes partitions in parallel and can redact entities before data is written to the training table. This keeps the redaction inside the lakehouse, scales with the cluster, and integrates with Unity Catalog governance. It is a common pattern for pre-training data sanitization.
Why this answer
PII must be removed from the content itself before training. Distributed detectors such as Presidio or Spark NLP in pandas UDFs, and model-based rewriting with ai_query(), both transform the text at scale inside Databricks. Column masks, encryption, and retention policies govern access or lifecycle but leave the underlying tokens intact, so they do not satisfy the preprocessing requirement.
Exam trap
The trap here is confusing access controls like column masks or encryption with actual content redaction, which must alter the text before training.