Courseiva

PDE Ingesting and Processing the Data Practice Question

A data engineer is building a Dataflow pipeline that reads newline-delimited JSON files from a Cloud Storage bucket and writes them to BigQuery. The files arrive continuously, some with malformed records, and the team wants the pipeline to keep running while routing bad records to a dead-letter location for later analysis. The team also wants the schema to be inferred from the files during development but fixed in production. Which two approaches should the data engineer use? (Choose two.)

⚠ Common exam trap

The trap here is conflating schema inference with pipeline execution settings, when inference is a development-time client concern and dead-lettering is a transform-level pattern.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define an explicit TableSchema in the pipeline for production and use schema autodetect only in a development branch.

Dead-letter routing is implemented with a side output from a DoFn that catches parse failures, keeping the main pipeline alive. Schema stability in production is achieved by declaring an explicit TableSchema rather than letting autodetect run continuously, with autodetect reserved for development. Together these two choices satisfy both the resilience and schema-governance requirements without changing write dispositions or relying on service-level flags.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Define an explicit TableSchema in the pipeline for production and use schema autodetect only in a development branch.

    Why this is correct

    Fixing the schema explicitly in production makes writes deterministic and prevents an unexpected field in one file from altering the table layout, while schema autodetect remains useful in development for discovering the shape of new data sources. This directly satisfies the requirement that inference be used only during development while production uses a fixed schema.

  • ✗

    Enable the --enableStreamingEngine flag so that schema inference happens automatically at the service level.

    Why it's wrong here

    Streaming Engine affects where pipeline state and shuffle are executed; it has no relationship to schema inference. Schema inference is a client-side, development-time concern handled by tools such as the BigQuery command-line schema autodetect or Beam's schema-aware transforms. Enabling this flag will not infer a schema, so the requirement remains unmet.

  • ✗

    Use BigQueryIO with withSchemaUpdateOptions set to ALLOW_FIELD_ADDITION and rely on runtime inference for every file.

    Why it's wrong here

    Schema update options allow adding nullable fields to an existing table, but they do not infer schemas from raw JSON and they do not handle malformed records. Relying on runtime inference for every file is also fragile because a single unexpected field can change the inferred schema and break downstream queries. This does not satisfy the production requirement for a fixed schema.

  • ✗

    Configure BigQueryIO with withCreateDisposition set to CREATE_NEVER and withWriteDisposition set to WRITE_TRUNCATE.

    Why it's wrong here

    WRITE_TRUNCATE deletes existing table data on every write, which is catastrophic for a continuously appending pipeline, and CREATE_NEVER fails if the destination table does not yet exist. Neither setting helps with malformed records or schema inference. This option addresses write semantics rather than the error-handling and schema requirements described in the scenario.

  • ✓

    Use a DoFn with a try/catch that emits failed parses to a side output tagged as dead-letter data.

    Why this is correct

    A DoFn that catches parse exceptions and emits them on a tagged side output is the canonical Beam pattern for dead-letter handling: the main output continues downstream while bad records are written to a separate sink such as Cloud Storage or a BigQuery error table. This keeps the pipeline running on malformed input and preserves the bad records for later analysis.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.