Google PCA Practice Question: Managing and Provisioning a Solution Infrastructure
A data engineering team wants to ingest streaming data from Pub/Sub, transform it using Apache Beam, and load it into BigQuery for real-time analytics. They need a fully managed solution that handles autoscaling and does not require managing servers. Which TWO Google Cloud services should they use?
⚠ Common exam trap
PCA often tests the difference between Dataflow (serverless Beam) and Dataproc (managed Spark/Hadoop), catching candidates who assume any data processing service is serverless.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Dataflow
Option B, Cloud Dataflow, is correct because it is Google Cloud's fully managed, serverless runner for Apache Beam pipelines, automatically handling autoscaling of worker instances and eliminating server management for the transform-and-load stage into BigQuery. Option E, Cloud Pub/Sub, is correct because it is the fully managed, serverless messaging service used to ingest the streaming data that Dataflow then reads and processes. Together, Pub/Sub provides the ingestion layer and Dataflow provides the managed Beam processing that writes results to BigQuery for real-time analytics. Option A, Cloud Dataproc, is not appropriate because it is a managed Spark/Hadoop cluster service that still requires cluster provisioning and management rather than being serverless. Option C, Cloud Dataprep, is a data-wrangling UI for preparing data, not a streaming Beam execution engine. Option D, Cloud Composer, is a managed Apache Airflow workflow orchestrator for scheduling batch pipelines, not a streaming data processing service.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Dataproc
Why it's wrong here
Dataproc is a managed Hadoop and Spark cluster service, so it does not run Apache Beam pipelines natively and still requires cluster provisioning. It is tempting because it processes large datasets on Google Cloud, and it would be correct for migrating existing Spark or Hadoop jobs rather than serverless Beam streaming.
- ✓
Cloud Dataflow
Why this is correct
Cloud Dataflow is the fully managed, serverless runner for Apache Beam pipelines, providing the autoscaling the team requires. It executes the transform stage and writes results into BigQuery, so no compute infrastructure needs provisioning or management.
- ✗
Cloud Dataprep
Why it's wrong here
Cloud Dataprep is a serverless data-wrangling tool for visual preparation of batch datasets, not a streaming Beam runner. It cannot execute Apache Beam pipelines against unbounded Pub/Sub data. Dataprep would suit interactive cleansing of files in Cloud Storage before loading, not real-time transformation.
- ✗
Cloud Composer
Why it's wrong here
Cloud Composer orchestrates batch workflows via Airflow DAGs; it does not provide the autoscaling Apache Beam runner needed for streaming transformation. Dataflow performs that managed, serverless Beam execution. Composer would be correct for scheduling dependent batch pipelines, not continuous Pub/Sub-to-BigQuery streaming.
- ✓
Cloud Pub/Sub
Why this is correct
Cloud Pub/Sub supplies the fully managed, serverless messaging layer that ingests the incoming streaming events and buffers them durably, decoupling producers from the Dataflow pipeline that consumes and transforms the data before loading into BigQuery.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
Learn chapter
IAM Policies, Service Accounts, and Auditing
Key term
Out-of-band
Out-of-band refers to a separate, dedicated network path used for managing and configuring IT devices, distinct from the main data traffic path.
Key term
Pub/Sub
Pub/Sub is a messaging pattern where publishers send messages without knowing who receives them, and subscribers receive only the messages they care about.
About these practice questions
This PCA question is part of Courseiva's 807-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.