An ML team trains a model using a dataset stored in a BigQuery table. They want to ensure that the exact data snapshot used for training is recorded and can be reproduced later for auditing. Which approach should they take?
BigQuery table snapshots provide a point-in-time, immutable copy of the table. By capturing the snapshot ID as a pipeline parameter, the training run is linked to the exact data version. This ensures reproducibility and auditability, as the snapshot can be restored or queried later. It directly addresses the need to record the data snapshot.
Why this answer
Using BigQuery table snapshots captures an immutable, point-in-time copy of the training data. Recording the snapshot ID in the pipeline run creates a direct link between the model and the exact data version, enabling reproducibility and audit compliance. Other methods either do not capture the precise data state or lack automatic linkage to the training run.
Exam trap
The trap here is assuming that exporting data or enabling audit logs provides a reproducible snapshot, when only a table snapshot guarantees an immutable, point-in-time copy.