Courseiva
hardMultiple Select

PDE Practice Question: Your company is building a data processing system…

Your company is building a data processing system that ingests sensor data from millions of devices, processes it in near real-time to detect anomalies, and stores raw and processed data for long-term analytics. The system must meet a 99.9% uptime SLA and minimize data loss. Which THREE design choices are best? (Choose three.)

⚠ Common exam trap

Google Cloud often tests the misconception that a load balancer is needed to scale Dataflow workers, when in fact Dataflow auto-scales its own workers and uses Pub/Sub's pull subscriptions to distribute messages evenly across workers without a separate load balancer.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Cloud Pub/Sub as the ingestion layer with a dead-letter topic to capture unprocessed messages.

Option A is correct because Cloud Pub/Sub provides a highly available, globally distributed ingestion layer that decouples millions of devices from downstream processing, and a dead-letter topic captures messages that repeatedly fail processing so they are not silently lost, supporting the 99.9% uptime and minimal data-loss goals. Option C is correct because Dataflow with at-least-once processing guarantees ensures no message is dropped during failures, and performing deduplication downstream handles the duplicate records that at-least-once semantics can produce, which is the standard trade-off for minimizing data loss in near real-time pipelines. Option D is correct because Cloud Storage is the durable, low-cost archival store for raw sensor data, while BigQuery is the serverless analytics warehouse suited for long-term processed-data analytics at scale. Option B is not best because Bigtable is optimized for high-throughput random reads/writes of time-series or keyed data rather than long-term analytical querying, and Cloud Storage alone for processed data lacks BigQuery's SQL analytics capability. Option E is not appropriate because a global Cloud Load Balancer distributes HTTP(S)/TCP traffic to backends and is not used to front Dataflow workers, which are managed by the Dataflow service and scale internally.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Cloud Pub/Sub as the ingestion layer with a dead-letter topic to capture unprocessed messages.

    Why this is correct

    Cloud Pub/Sub decouples ingestion from processing, absorbing millions of device events while buffering against consumer failures, which supports the 99.9% uptime SLA. Its dead-letter topic captures messages that repeatedly fail processing, preventing silent data loss and satisfying the minimise-data-loss constraint.

  • ✗

    Store raw data in Cloud Bigtable and processed data in Cloud Storage.

    Why it's wrong here

    Bigtable is a low-latency NoSQL store for serving lookups, not a durable data lake for raw sensor data; Cloud Storage alone provides the cheap, durable retention long-term analytics needs. Bigtable is tempting when raw data must be queried at high throughput by row key.

  • ✓

    Use Dataflow with at-least-once processing guarantees and perform deduplication downstream.

    Why this is correct

    At-least-once delivery ensures no sensor reading is dropped during retries or worker restarts, directly serving the minimise-data-loss requirement. Because duplicates are inevitable, downstream deduplication—keyed on device ID and timestamp—restores exactly-once semantics for anomaly detection while preserving the 99.9% uptime SLA.

  • ✓

    Use Cloud Storage for raw data archival and BigQuery for processed analytics data.

    Why this is correct

    Cloud Storage provides durable, low-cost object storage for raw sensor archives, while BigQuery's columnar engine serves long-term analytics over processed data. This satisfies the stem's retention requirement, separating cheap immutable archival from query-optimised warehousing so neither workload compromises the other's performance or the 99.9% uptime SLA.

  • ✗

    Use a global Cloud Load Balancer in front of the Dataflow workers.

    Why it's wrong here

    Dataflow workers are a managed, autoscaled pool that Google load-balances internally; a global external load balancer in front of them adds no ingestion benefit and cannot distribute work across the pipeline. Global load balancing is tempting for routing user traffic to regional backends, not for Dataflow internals.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.