PDE Ingesting and Processing the Data Practice Question
You are building a Dataflow pipeline that reads from Cloud Storage and writes to BigQuery. The pipeline must handle files that are compressed with gzip and contain JSON data. You need to ensure that the pipeline can process these files efficiently and write to BigQuery with minimal errors. Which approach should you take?
⚠ Common exam trap
The trap here is assuming that BigQueryIO can read files or that manual decompression is necessary, when TextIO already handles gzip seamlessly.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use TextIO to read the gzip files, parse the JSON using a DoFn, and write to BigQuery using BigQueryIO with the STORAGE_WRITE_API method.
TextIO natively supports reading gzip-compressed files, simplifying ingestion. Parsing JSON in a DoFn allows transformation into BigQuery-compatible rows. Using BigQueryIO with the STORAGE_WRITE_API method ensures high-throughput, exactly-once writes, which is critical for efficient and reliable loading. This combination addresses the compressed format and minimizes errors through robust write semantics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use TextIO to read the gzip files, parse the JSON using a DoFn, and write to BigQuery using BigQueryIO with the STORAGE_WRITE_API method.
Why this is correct
TextIO can read gzip-compressed files transparently, and a DoFn can parse the JSON into a TableRow. Writing with BigQueryIO using the STORAGE_WRITE_API method provides high-performance, exactly-once writes. This combination efficiently processes compressed JSON files and handles errors through dead-letter patterns if configured, meeting the requirements.
- ✗
Use TextIO to read the gzip files, but disable compression detection and manually decompress each file in a DoFn before parsing JSON.
Why it's wrong here
TextIO automatically detects and decompresses gzip files based on the file extension. Disabling compression detection and manually decompressing is redundant and error-prone. It adds complexity without benefit, and may lead to performance degradation. The correct approach is to let TextIO handle decompression natively.
- ✗
Use AvroIO to read the gzip files, convert the Avro records to JSON, and write to BigQuery using BigQueryIO with the FILE_LOADS method.
Why it's wrong here
AvroIO is designed for Avro files, not gzip-compressed JSON. Attempting to read gzip JSON with AvroIO would fail because the file format does not match. Converting to Avro and then to JSON adds unnecessary complexity. FILE_LOADS is a batch method and not ideal for streaming, and the approach does not address the file format mismatch.
- ✗
Use BigQueryIO to read the gzip files directly, parse the JSON, and write to BigQuery using the default write method.
Why it's wrong here
BigQueryIO is a sink for writing to BigQuery, not a source for reading from Cloud Storage. It cannot read gzip files directly. Using it as a source is incorrect. The default write method may not provide the same performance or exactly-once guarantees as the Storage Write API, and the approach fundamentally misunderstands BigQueryIO's purpose.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.