PDE Preparing and Using Data for Analysis Practice Question
You want to quickly estimate the number of distinct visitors to your website from a large BigQuery table. Which function provides an approximate count with low latency?
⚠ Common exam trap
PDE often tests the confusion between approximate and exact functions, leading candidates to choose COUNT(DISTINCT) for performance despite its high latency on large datasets.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
APPROX_COUNT_DISTINCT
APPROX_COUNT_DISTINCT is a built-in BigQuery function that returns an approximate count of distinct values using a HyperLogLog++ sketch. It is designed for large-scale data, providing fast, low-latency results with a small error rate (typically <1%). This makes it ideal for quickly estimating distinct visitors without the overhead of exact counting.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
APPROX_COUNT_DISTINCT
Why this is correct
APPROX_COUNT_DISTINCT uses HyperLogLog++ sketches to estimate cardinality within a small, bounded memory footprint, returning results in seconds rather than the minutes an exact COUNT(DISTINCT) would need on a large table. This directly satisfies the stem's low-latency requirement for estimating distinct website visitors.
- ✗
HyperLogLog++
Why it's wrong here
HyperLogLog++ is the approximate algorithm itself, not a BigQuery SQL function; invoking it directly is not how the engine exposes approximate counting. It is tempting because it underpins APPROX_COUNT_DISTINCT, which is the correct choice here, but naming the algorithm does not answer which function to call.
- ✗
COUNT(DISTINCT)
Why it's wrong here
COUNT(DISTINCT) computes exact distinct counts, forcing BigQuery to shuffle and merge every value, so latency and cost scale with table size. It is the right choice when exactness matters, such as billing or compliance reporting, but the stem explicitly asks for an approximation.
- ✗
APPROX_QUANTILES
Why it's wrong here
APPROX_QUANTILES returns approximate percentile boundaries from a value distribution, not a distinct-value estimate, so it cannot count unique visitors. It is tempting because it is an approximate aggregation, and would be correct when you need median or percentile latency values rather than cardinality.
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.