PDE Preparing and Using Data for Analysis Practice Question
You are designing a Dataflow pipeline that reads from a Cloud Storage bucket containing thousands of small JSON files, transforms the data, and writes to BigQuery. The pipeline is slow and expensive because of the large number of small files. You want to improve throughput and reduce cost. What should you do?
⚠ Common exam trap
The trap here is assuming that adding workers or using FileIO hints will fix small-file inefficiency, when the real fix is reducing the number of files by compaction.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Preprocess the data outside Dataflow to combine the small files into larger files (e.g., 100-500 MB each) in Cloud Storage, then run the Dataflow pipeline on the consolidated files.
The bottleneck with thousands of small files is the per-file read and metadata overhead. Combining them into larger files before the Dataflow run reduces the number of reads and improves throughput. Increasing workers or using FileIO hints does not address the root cause, and the Storage Read API is for BigQuery tables, not Cloud Storage files.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use the FileIO transform with a match pattern and enable the `withHintMatchesManyFiles` option to optimize reading many small files.
Why it's wrong here
The withHintMatchesManyFiles hint helps FileIO handle a large number of matches more efficiently by using a different expansion strategy, but it does not combine the files into larger reads. The per-file read overhead remains, so throughput may still be limited. It is a partial mitigation, not a complete solution.
- ✗
Increase the number of Dataflow workers and set the autoscaling algorithm to THROUGHPUT_BASED to handle the small files in parallel.
Why it's wrong here
Adding workers increases parallelism but does not solve the fundamental inefficiency of reading thousands of tiny files. Each file requires a separate read operation, and metadata overhead dominates. More workers may increase cost without proportional throughput gains. The root cause is file granularity, not worker count.
- ✓
Preprocess the data outside Dataflow to combine the small files into larger files (e.g., 100-500 MB each) in Cloud Storage, then run the Dataflow pipeline on the consolidated files.
Why this is correct
Consolidating small files into larger ones reduces the number of read operations and metadata overhead, which is the primary cause of slowness. Dataflow can then read each large file efficiently, improving throughput and lowering cost. This is the recommended pattern for many small files.
- ✗
Switch the pipeline to use the BigQuery Storage Read API to read the JSON files directly, bypassing Cloud Storage.
Why it's wrong here
The BigQuery Storage Read API reads data from BigQuery tables, not from Cloud Storage JSON files. It cannot access the source files in this scenario. This option misidentifies the API's purpose and would not work for reading raw JSON from Cloud Storage.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.