Courseiva
easyMultiple Choice

PDE Practice Question: A media company runs a batch data pipeline on…

A media company runs a batch data pipeline on Cloud Dataflow that ingests log files from Cloud Storage, transforms them, and writes results to BigQuery for analytics. The pipeline runs daily and has been stable for months. Recently, the source log format changed: a new optional field was added to some records. The pipeline started failing with ParseErrors for rows that contain the new field. The error logs show that the Dataflow job uses a hardcoded JSON schema that does not include the new field. The Dataflow pipeline logs are written to Stackdriver Logging, but no alerts are configured. The team wants to ensure that future schema changes do not break the pipeline and that failures are detected promptly. The team has limited experience with streaming and wants to keep the batch approach. Which course of action should the team take to improve solution quality?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement schema validation and evolution using a schema registry (e.g., AVRO) in the Dataflow pipeline, and configure Stackdriver alerts on pipeline failure or error logs.

It directly addresses both requirements: schema validation and evolution via a schema registry (e.g., AVRO) allows the Dataflow pipeline to handle added optional fields without hardcoded schema ParseErrors, and configuring Stackdriver (Cloud Monitoring) alerts on pipeline failure or error logs ensures prompt detection of future failures. A schema registry supports backward-compatible schema evolution, which is the standard way to prevent new optional fields from breaking batch ingestion. Option A only adds alerting and a manual runbook, so it detects failures but does not prevent schema changes from breaking the pipeline. Option B's hourly Cloud Function header check is a custom, brittle workaround that does not integrate with Dataflow's parsing or guarantee schema compatibility. Option C's BigQuery dry runs validate BigQuery load compatibility, not the Dataflow pipeline's JSON parsing schema, so it would not prevent the ParseErrors.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create a Cloud Monitoring alert on any PipelineError log entries from the Dataflow job, and set up a runbook to manually fix schema mismatches within one hour.

    Why it's wrong here

    Alerting on PipelineError entries plus a manual runbook detects failures but leaves the hardcoded schema rejecting the new optional field, so every run still fails until someone edits code. It is tempting because it addresses the missing alerts, but it does not make the pipeline tolerate schema evolution.

  • ✗

    Schedule a Cloud Function to run every hour that checks the latest log file headers and compares them to the pipeline schema, sending an alert if differences are found.

    Why it's wrong here

    A Cloud Function comparing file headers to the schema detects drift but cannot prevent ParseErrors, since the pipeline's hardcoded schema still rejects new fields. It is tempting as lightweight monitoring, yet the correct approach makes the pipeline schema-flexible and alerts on actual job failures rather than inspecting headers.

  • ✗

    Use BigQuery dry run queries to validate the schema before loading data, and if a mismatch is detected, block the pipeline run and notify the team via email.

    Why it's wrong here

    BigQuery dry runs validate queries against BigQuery tables, not the JSON parsing stage in Dataflow, so ParseErrors occur before any load. It is tempting because dry runs are cheap validation, but blocking the run still leaves the hardcoded schema rejecting the new optional field.

  • ✓

    Implement schema validation and evolution using a schema registry (e.g., AVRO) in the Dataflow pipeline, and configure Stackdriver alerts on pipeline failure or error logs.

    Why this is correct

    AVRO schema registry enforces compatibility and permits optional-field evolution, so added fields no longer trigger ParseErrors against the hardcoded schema. Stackdriver alerts on failure logs give prompt detection. This preserves the existing batch approach, matching the team's limited streaming experience and stated preference.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.