PDE Ingesting and Processing the Data Practice Question
You are building a Dataflow pipeline that reads Avro-formatted files from Cloud Storage. The files use a schema that is updated frequently. You want to minimize pipeline restarts due to schema changes. Which approach should you take?
⚠ Common exam trap
The trap here is assuming that Avro schemas must be defined statically in the pipeline code, when in fact Avro's self-describing format allows dynamic schema resolution.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Read the Avro files using a schema inferred at runtime from the Avro file metadata.
Avro files include their schema in the file header, so a Dataflow pipeline can read that schema at runtime and adapt to changes automatically. This avoids the need to hardcode or manually update schemas, reducing pipeline restarts when the source schema evolves.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert all Avro files to CSV before ingestion and use a fixed CSV schema.
Why it's wrong here
Converting to CSV loses Avro's self-describing schema and strong typing, and CSV does not carry schema information, so you would still need to manage schema changes manually. This adds unnecessary processing and does not solve the restart problem.
- ✗
Use a BigQuery load job with schema autodetect instead of Dataflow.
Why it's wrong here
BigQuery load jobs do not process data in a Dataflow pipeline; they are for batch loading into BigQuery. The scenario requires a Dataflow pipeline, so switching to a load job is not applicable and does not address the Avro schema evolution within Dataflow.
- ✓
Read the Avro files using a schema inferred at runtime from the Avro file metadata.
Why this is correct
Avro files embed their schema in the file header, so a Dataflow pipeline can dynamically read that schema and adapt to changes without code changes or restarts. This is the standard approach for handling evolving Avro schemas in a streaming or batch pipeline.
- ✗
Use a fixed Avro schema in your pipeline code and manually update it each time the source schema changes.
Why it's wrong here
Hardcoding an Avro schema forces you to redeploy the pipeline whenever the source schema evolves, causing downtime and defeating the goal of minimizing restarts. It also risks read failures if the incoming data includes new fields not present in the hardcoded schema.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.