Courseiva
mediumMultiple Choice

PDE Practice Question: A retail company is building a recommendation…

A retail company is building a recommendation engine that requires processing customer clickstream data in near real-time. The data is ingested via Pub/Sub, and must be joined with a lookup table of product details (updated daily) before being used for model inference. Which design pattern should they use?

⚠ Common exam trap

Google Cloud often tests the distinction between streaming enrichment patterns that require external lookups (which add latency and cost) versus using side inputs for static or slowly-changing reference data, leading candidates to mistakenly choose a cache-based solution like Redis when the data is already available in Cloud Storage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a Dataflow pipeline that reads from Pub/Sub and uses a side input from a regularly refreshed PCollection loaded from Cloud Storage.

Dataflow can read streaming data from Pub/Sub and use a side input from a regularly refreshed PCollection loaded from Cloud Storage. This pattern allows the product lookup table (updated daily) to be periodically reloaded into the pipeline as a side input, enabling efficient, low-latency enrichment of each event without per-event external calls or batch delays.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enrich the stream by querying BigQuery for each event using a Cloud Function.

    Why it's wrong here

    Per-event BigQuery lookups add query latency and cost, breaking the near real-time requirement for clickstream enrichment. It is tempting because BigQuery holds the authoritative product table, and would be correct for low-throughput batch enrichment where freshness matters less than simplicity.

  • ✓

    Use a Dataflow pipeline that reads from Pub/Sub and uses a side input from a regularly refreshed PCollection loaded from Cloud Storage.

    Why this is correct

    A side input broadcasts the product lookup PCollection to every worker, letting each clickstream element be enriched in-flight without a per-element external lookup. Refreshing it daily from Cloud Storage matches the stem's daily update cadence while preserving near-real-time inference.

  • ✗

    Store product details in Cloud Memorystore (Redis) and have the streaming application look up each event.

    Why it's wrong here

    Redis lookups add a network hop and require cache invalidation when the daily product table changes, complicating correctness. It is tempting for low-latency enrichment, but the pattern calls for a side-input or broadcast state join in the streaming pipeline itself.

  • ✗

    Write events to BigQuery and use scheduled queries to join with the product table in batch.

    Why it's wrong here

    Scheduled batch queries introduce minutes-to-hours latency, so clickstream events are not enriched before model inference, violating near real-time processing. It is tempting because BigQuery joins are simple and the product table already lives there, and would suit daily batch scoring rather than streaming.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.