PDE Preparing and Using Data for Analysis Practice Question
You need to track the lineage of data in BigQuery, showing how tables are derived from other tables via queries. Which service provides this capability?
⚠ Common exam trap
PDE often tests the confusion between Data Catalog (metadata discovery) and the Lineage API (derivation tracking) — candidates who equate 'catalog' with 'lineage' pick C.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
BigQuery Lineage API
The BigQuery Lineage API provides programmatic access to lineage information, showing how tables are derived from other tables through queries, including column-level lineage. It captures dependencies from SQL queries, scheduled queries, and other BigQuery operations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
BigQuery Lineage API
Why this is correct
The BigQuery Lineage API exposes column- and table-level lineage captured from query jobs, tracing how each table is derived from upstream tables. This directly satisfies the requirement to track derivation lineage, unlike Dataplex or Data Catalog, which handle discovery and governance metadata.
- ✗
Cloud Composer
Why it's wrong here
Cloud Composer orchestrates workflows via Airflow DAGs; it schedules and runs pipelines but does not capture or expose table-level lineage between BigQuery datasets. It is correctly chosen for dependency scheduling and task orchestration, not for metadata lineage visualisation.
- ✗
Cloud Data Catalog
Why it's wrong here
Cloud Data Catalog is a metadata management service for discovering and tagging assets; it does not automatically parse BigQuery SQL to derive table-to-table lineage. It is tempting because it stores metadata, but lineage requires a dedicated service that analyses query jobs and builds dependency graphs.
- ✗
Dataflow
Why it's wrong here
Dataflow is a managed stream and batch processing service for building pipelines; it does not record or display table-level lineage metadata. It is tempting because Dataflow jobs read and write BigQuery tables, but lineage tracking requires a metadata catalogue rather than a processing engine.
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.