Courseiva
mediumMultiple Choice

PDE Practice Question: A media company uses Cloud Data Loss Prevention…

A media company uses Cloud Data Loss Prevention (DLP) API to inspect and de-identify sensitive data before loading into BigQuery. They want to reduce costs by sampling the data during inspection. Which configuration should they use?

⚠ Common exam trap

Test-takers frequently confuse 'limiting rows/bytes' (which scans sequentially from the start) with 'random sampling' (which distributes inspection across the entire dataset), leading them to pick options A or D, which do not achieve representative cost reduction.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set the sample method to 'RANDOM' with a percentage.

The Cloud DLP API supports a 'sample_method' of 'RANDOM' with a 'sampling_percentage' to inspect only a random subset of rows. This directly reduces the volume of data scanned, lowering costs while still providing statistically representative coverage for sensitive data discovery.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the 'ROWS' limit in the inspection job.

    Why it's wrong here

    The ROWS limit applies to BigQuery table inspection, restricting how many rows are scanned, which is the sampling mechanism for structured data. It is the correct choice when inspecting BigQuery tables directly, but the stem's de-identification workflow before loading implies a different sampling control.

  • ✓

    Set the sample method to 'RANDOM' with a percentage.

    Why this is correct

    The RANDOM sample method with a percentage inspects only a statistical subset of records, reducing both the volume of data processed and the associated DLP API charges. This satisfies the stated cost-reduction goal while still providing representative inspection coverage before loading into BigQuery.

  • ✗

    Use a hybrid inspection with a BigQuery sample table.

    Why it's wrong here

    A hybrid inspection with a sample table requires extracting and staging a representative subset in BigQuery before scanning, adding pipeline complexity rather than configuring sampling within the inspection job itself. It suits recurring audits of stable datasets, not one-off cost reduction during pre-load inspection.

  • ✗

    Use the 'BYTES_LIMIT' parameter.

    Why it's wrong here

    BYTES_LIMIT caps the number of bytes scanned per content item, which limits inspection volume but does not implement row-based sampling of tabular BigQuery data. It is appropriate for bounding cost on unstructured content such as text blobs, where byte ceilings map naturally to scan scope.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.