Courseiva

PDE Ingesting and Processing the Data Practice Question

A data platform team uses Cloud Data Fusion to move data from an on-premises relational database into BigQuery. They need the pipeline to run on a fixed schedule, capture only rows changed since the last successful run, and avoid re-reading the entire source table each night. The source table has an updated_at column that is reliably populated. Which two approaches should they use? (Choose two.)

⚠ Common exam trap

The trap here is treating lineage or truncate-and-reload as change-capture mechanisms, when only a stored watermark plus a filtered source query reads just the changed rows.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the pipeline to run on a schedule using the Data Fusion scheduler or a Cloud Composer trigger

Scheduling the pipeline and filtering the source query on a persisted updated_at watermark together deliver a recurring job that reads only changed rows. The watermark must be written after a successful load so failures do not advance it prematurely, and the query-based Database source plugin is what makes the filter possible without custom code.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a Replicator pipeline with the default full-table snapshot mode for each run

    Why it's wrong here

    A Replicator pipeline in default snapshot mode reads the entire source table on every execution, so each nightly run would re-scan all rows regardless of the updated_at column. That directly conflicts with the requirement to capture only changed rows and avoid full-table reads, and it would increase load on the source database unnecessarily.

  • ✗

    Enable the Cloud Data Fusion lineage feature to track which rows were previously loaded

    Why it's wrong here

    Lineage records metadata about data movement and field-level transformations for governance and impact analysis; it does not store row-level load state or act as a watermark. Relying on lineage to decide what to read next would not prevent a full-table scan, so it cannot satisfy the incremental extraction requirement described in the scenario.

  • ✓

    Configure the pipeline to run on a schedule using the Data Fusion scheduler or a Cloud Composer trigger

    Why this is correct

    Cloud Data Fusion pipelines can be scheduled directly through the built-in scheduler, or triggered externally by Cloud Composer, which provides the fixed nightly cadence the team requires. Without a schedule the pipeline would have to be started manually, so this is necessary to meet the recurring-run requirement while keeping the incremental logic in the pipeline itself.

  • ✗

    Use the BigQuery sink plugin with write disposition set to truncate before each nightly load

    Why it's wrong here

    Truncating the destination before each load would discard all previously loaded history and force the pipeline to re-read the entire source table to repopulate it, which is the opposite of the incremental behavior requested. It also adds risk of a partial or empty table if the load fails midway, so it does not satisfy the change-capture requirement.

  • ✓

    Use the Database source plugin with a query that filters on updated_at greater than the last recorded watermark

    Why this is correct

    The Database source plugin accepts a custom query, so filtering on updated_at greater than the stored watermark reads only rows changed since the previous successful run. Persisting the watermark after a successful load keeps the job incremental rather than full-table, which is exactly what the team asked for and is supported natively by the plugin.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.