Courseiva

PDE Ingesting and Processing the Data Practice Question

A company wants to use dbt (data build tool) to transform data in BigQuery. They have a Cloud Storage bucket containing raw CSV files that are loaded daily into BigQuery via an external table. Which dbt feature should they use to modularize the transformation logic and handle dependencies between models?

⚠ Common exam trap

Candidates often confuse the purpose of dbt components: models with `ref()` manage transformation logic and dependencies, while tests handle data quality, snapshots track historical changes, and seeds load static data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

dbt models with ref()

C is correct because dbt models with the `ref()` function allow you to modularize SQL transformation logic and automatically handle dependencies between models. When you use `ref('model_name')`, dbt builds a dependency graph, ensuring models are executed in the correct order based on their references. This is essential for transforming raw data from an external table into a structured, analytics-ready dataset in BigQuery.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    dbt tests

    Why it's wrong here

    Tests assert data quality expectations on existing models; they do not modularise transformation logic or manage dependencies. Models with ref() build the DAG. Tests are the correct choice when you need to validate uniqueness, not-null constraints, or accepted values before downstream consumption.

  • ✗

    dbt snapshots

    Why it's wrong here

    Snapshots capture slowly changing dimension history by recording row-level changes over time; they do not modularise transformations or wire dependencies. Models with ref() provide that. Snapshots are correct when you must preserve historical versions of mutable records, such as customer address changes.

  • ✓

    dbt models with ref()

    Why this is correct

    The ref() function creates a directed acyclic graph between dbt models, letting each model reference upstream ones by name rather than hard-coded table paths. This satisfies the modularisation and dependency-handling requirement, since dbt resolves build order automatically and materialises each model against the BigQuery external table.

  • ✗

    dbt seeds

    Why it's wrong here

    Seeds load static CSV files into the warehouse as tables; they do not modularise transformation logic or resolve dependencies. That role belongs to dbt models, which use ref() to build a dependency DAG. Seeds suit small, slowly changing lookup data such as country codes or mapping tables.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.