PDE Ingesting and Processing the Data Practice Question
A data engineer is building a Dataflow pipeline in Python that reads from BigQuery, transforms data, and writes to Cloud Storage. The pipeline will be deployed in production. Which approach should they use to ensure the pipeline is reusable across environments with different configuration parameters?
⚠ Common exam trap
Google often tests the distinction between Classic Templates and Flex Templates, where candidates mistakenly choose Classic Templates because they are simpler, but Flex Templates are required for custom environments and parameterized production reuse.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a Dataflow Flex Template
Dataflow Flex Templates allow you to package a Docker container with your pipeline code and dependencies, enabling parameterization at runtime via the Dataflow UI, CLI, or API. This makes the pipeline reusable across environments (e.g., dev, staging, prod) by passing different configuration parameters (like project IDs, table names, or output paths) without modifying the code. Flex Templates support custom container images and are the recommended approach for production pipelines that need environment-agnostic deployment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a separate pipeline for each environment with hardcoded values
Why it's wrong here
Hardcoding values per environment duplicates pipeline code, so a change must be applied in every copy and drift is likely. It is tempting because separate pipelines appear to isolate environments cleanly, but reusable configuration requires runtime parameters supplied at launch instead.
- ✗
Use a Dataflow Classic Template
Why it's wrong here
Classic Templates fix the pipeline graph at creation and accept only a limited set of predefined parameters, so arbitrary environment-specific configuration cannot be injected. It is tempting because templates package reusable pipelines, but Flex Templates are needed when runtime parameters must vary per environment.
- ✓
Use a Dataflow Flex Template
Why this is correct
Flex Templates package the pipeline as a Docker image with a metadata file, so runtime parameters such as input, output and project values are supplied at job launch. This satisfies the reusability requirement across environments without editing code, unlike classic templates with fixed dependencies.
- ✗
Run the pipeline using the DirectRunner for each environment
Why it's wrong here
DirectRunner executes locally on the engineer's machine and ignores Dataflow's managed scaling, so it cannot serve production workloads. It is tempting for quick local testing of pipeline logic, but production deployment across environments requires a runner that runs on Google Cloud infrastructure.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.