Courseiva

PDE Preparing and Using Data for Analysis Practice Question

You need to track the lineage of data in BigQuery, showing how tables are derived from other tables via queries. Which service provides this capability?

⚠ Common exam trap

PDE often tests the confusion between Data Catalog (metadata discovery) and the Lineage API (derivation tracking) — candidates who equate 'catalog' with 'lineage' pick C.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

BigQuery Lineage API

The BigQuery Lineage API provides programmatic access to lineage information, showing how tables are derived from other tables through queries, including column-level lineage. It captures dependencies from SQL queries, scheduled queries, and other BigQuery operations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    BigQuery Lineage API

    Why this is correct

    The BigQuery Lineage API exposes column- and table-level lineage captured from query jobs, tracing how each table is derived from upstream tables. This directly satisfies the requirement to track derivation lineage, unlike Dataplex or Data Catalog, which handle discovery and governance metadata.

  • ✗

    Cloud Composer

    Why it's wrong here

    Cloud Composer orchestrates workflows via Airflow DAGs; it schedules and runs pipelines but does not capture or expose table-level lineage between BigQuery datasets. It is correctly chosen for dependency scheduling and task orchestration, not for metadata lineage visualisation.

  • ✗

    Cloud Data Catalog

    Why it's wrong here

    Cloud Data Catalog is a metadata management service for discovering and tagging assets; it does not automatically parse BigQuery SQL to derive table-to-table lineage. It is tempting because it stores metadata, but lineage requires a dedicated service that analyses query jobs and builds dependency graphs.

  • ✗

    Dataflow

    Why it's wrong here

    Dataflow is a managed stream and batch processing service for building pipelines; it does not record or display table-level lineage metadata. It is tempting because Dataflow jobs read and write BigQuery tables, but lineage tracking requires a metadata catalogue rather than a processing engine.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.