Courseiva
Storing the Data →hardMultiple Choice

PDE Storing the Data Practice Question

A financial services company uses Cloud Bigtable to store trade data. They are experiencing hot-spotting on a single node, causing high latency. The row key format is [trade_id]#[timestamp]. Which row key design change would BEST distribute writes across tablets?

⚠ Common exam trap

A common trap in the Google Professional Data Engineer exam is confusing the order of key components. Simply reversing the order (e.g., putting the timestamp first) is often incorrectly thought to solve hot-spotting, but any monotonically increasing value at the start of the key will still cause a hotspot because Bigtable stores rows in lexicographic order.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a hashed prefix of the trade_id, e.g., [hash(trade_id)]#[trade_id]#[timestamp]

Adding a hashed prefix of the trade_id ensures that writes are evenly distributed across all Bigtable tablets. Bigtable partitions data by row key lexicographic order; without a hash, sequential trade IDs or timestamps cause all recent writes to land on a single tablet, creating a hotspot. The hash spreads the write load uniformly, regardless of the underlying key pattern.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a hashed prefix of the trade_id, e.g., [hash(trade_id)]#[trade_id]#[timestamp]

    Why this is correct

    Hashing the trade_id prefix scatters sequential writes across the entire key space, so consecutive trades land on different tablets rather than concentrating on one node. This directly resolves the hot-spotting constraint, since Cloud Bigtable distributes rows lexicographically by row key and a uniform hash prefix balances write load.

  • ✗

    Use a single row key of [timestamp]

    Why it's wrong here

    A timestamp-only key is itself monotonically increasing, so sequential writes still land on one tablet, worsening hot-spotting rather than spreading load. It is tempting because timestamps give natural range-scan ordering for time-series queries, but Bigtable distributes by key-range boundaries, so a low-cardinality leading value cannot scatter writes.

  • ✗

    Increase the number of Bigtable nodes to 20

    Why it's wrong here

    Adding nodes increases total cluster throughput but cannot redistribute a single hot tablet, since Bigtable splits load by row-key range, not node count. It is tempting because scaling is the usual remedy for latency, but the stem's root cause is key design, so extra nodes leave the hot-spotting pattern intact.

  • ✗

    Change row key to [timestamp]#[trade_id]

    Why it's wrong here

    Leading with timestamp keeps keys sequential, so consecutive writes still concentrate on the last tablet, leaving hot-spotting unresolved. It is tempting because timestamp-first ordering supports efficient time-range scans, but Bigtable splits by row-key range, so a monotonically increasing prefix defeats distribution regardless of the trailing trade_id.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.