Courseiva

PDE Preparing and Using Data for Analysis Practice Question

You want to quickly estimate the number of distinct visitors to your website from a large BigQuery table. Which function provides an approximate count with low latency?

⚠ Common exam trap

PDE often tests the confusion between approximate and exact functions, leading candidates to choose COUNT(DISTINCT) for performance despite its high latency on large datasets.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

APPROX_COUNT_DISTINCT

APPROX_COUNT_DISTINCT is a built-in BigQuery function that returns an approximate count of distinct values using a HyperLogLog++ sketch. It is designed for large-scale data, providing fast, low-latency results with a small error rate (typically <1%). This makes it ideal for quickly estimating distinct visitors without the overhead of exact counting.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    APPROX_COUNT_DISTINCT

    Why this is correct

    APPROX_COUNT_DISTINCT uses HyperLogLog++ sketches to estimate cardinality within a small, bounded memory footprint, returning results in seconds rather than the minutes an exact COUNT(DISTINCT) would need on a large table. This directly satisfies the stem's low-latency requirement for estimating distinct website visitors.

  • ✗

    HyperLogLog++

    Why it's wrong here

    HyperLogLog++ is the approximate algorithm itself, not a BigQuery SQL function; invoking it directly is not how the engine exposes approximate counting. It is tempting because it underpins APPROX_COUNT_DISTINCT, which is the correct choice here, but naming the algorithm does not answer which function to call.

  • ✗

    COUNT(DISTINCT)

    Why it's wrong here

    COUNT(DISTINCT) computes exact distinct counts, forcing BigQuery to shuffle and merge every value, so latency and cost scale with table size. It is the right choice when exactness matters, such as billing or compliance reporting, but the stem explicitly asks for an approximation.

  • ✗

    APPROX_QUANTILES

    Why it's wrong here

    APPROX_QUANTILES returns approximate percentile boundaries from a value distribution, not a distinct-value estimate, so it cannot count unique visitors. It is tempting because it is an approximate aggregation, and would be correct when you need median or percentile latency values rather than cardinality.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.