Courseiva
easyMultiple Choice

PDE Practice Question: Building a data lake on Cloud Storage for log…

A company is building a data lake on Cloud Storage for log analysis. Log files (CSV) arrive every 5 minutes from multiple sources. The files should be ingested into BigQuery for reporting within 15 minutes. Which approach best meets the requirements with minimal operational overhead?

⚠ Common exam trap

Google Cloud often tests the misconception that serverless options like Cloud Functions are only for simple tasks, but here they are the most efficient choice for near-real-time ingestion with minimal overhead, while Dataflow is overkill for this straightforward file-load pattern.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set up a Cloud Storage notification to trigger a Cloud Function that loads each file into BigQuery using the BigQuery API.

Cloud Storage notifications trigger a Cloud Function on each file upload, which then loads the file into BigQuery via the BigQuery API. This provides near-real-time ingestion (within seconds of file arrival) with minimal operational overhead, as there are no servers to manage and no scheduling needed. The 5-minute file arrival and 15-minute SLA are easily met without complex infrastructure.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Set up a Cloud Storage notification to trigger a Cloud Function that loads each file into BigQuery using the BigQuery API.

    Why this is correct

    Cloud Storage notifications triggering a Cloud Function gives event-driven, per-file ingestion, so each CSV lands in BigQuery within the 15-minute window. Serverless execution removes cluster or pipeline management, satisfying the minimal operational overhead constraint. The BigQuery API load job handles schema and append semantics directly, avoiding scheduled batch polling that could breach the latency requirement.

  • ✗

    Schedule a daily batch load from Cloud Storage to BigQuery using the BigQuery Data Transfer Service.

    Why it's wrong here

    A daily schedule delivers data up to 24 hours late, breaching the 15-minute requirement. The Data Transfer Service suits recurring bulk loads from Cloud Storage or external sources where latency of hours is acceptable, not five-minute log arrivals needing near-real-time reporting.

  • ✗

    Use Dataflow to read from Pub/Sub (ingested from Cloud Storage) and write to BigQuery.

    Why it's wrong here

    Dataflow with Pub/Sub adds a streaming pipeline to build and operate, contradicting minimal operational overhead when files already land in Cloud Storage every five minutes. It suits high-volume continuous event streams, not scheduled micro-batch file ingestion.

  • ✗

    Use BigQuery federated queries to query the CSV files directly from Cloud Storage.

    Why it's wrong here

    Federated queries read external CSV files at query time, so each report scans Cloud Storage rather than ingesting rows, adding latency and cost and giving no managed pipeline. They suit ad-hoc exploration of external data, not continuous 15-minute ingestion for reporting.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.