Courseiva
Data Store Management →hardMultiple Select

DEA-C01 Data Store Management Practice Question

A company is migrating a legacy data warehouse to Amazon Redshift. They need to choose a distribution style to minimize data movement during joins. Which THREE factors should they consider?

⚠ Common exam trap

The trap here is that candidates may overthink irrelevant table properties like column count or data types, while the core considerations for minimizing data movement are table size, join frequency, and table role (fact vs. dimension).

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The size of the table (number of rows).

Option A is correct because table size (number of rows) is a primary driver of distribution style choice: small tables are typically assigned ALL distribution so they can be broadcast to every node, eliminating data movement during joins, while large tables need KEY or EVEN distribution. Option B is correct because the whole purpose of choosing a distribution style is to co-locate matching rows on the same slice; if a table is frequently joined on a specific column, distributing both tables on that join key keeps the join local and avoids network redistribution. Option D is correct because fact and dimension tables play different roles in a star schema: large fact tables are usually distributed on their most common join key (or EVEN), while smaller dimension tables are often set to ALL so they are replicated to every node and can be joined without movement. Option C is not a factor because the number of columns does not affect how rows are distributed across slices; it only affects storage width and I/O, not join data movement. Option E is not a factor because Redshift supports distribution keys of various data types, and the data type itself does not determine whether a join requires redistribution — what matters is whether the joined columns match and are used as distribution keys.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The size of the table (number of rows).

    Why this is correct

    Table size determines whether a table qualifies for ALL distribution, which broadcasts a small table to every node and eliminates join data movement entirely. For large tables, row count instead guides choosing DISTKEY on the frequent join column, since ALL replication becomes impractical. Both paths directly minimise the internode traffic the stem requires.

  • ✓

    The join frequency with other tables on specific columns.

    Why this is correct

    Join frequency on specific columns determines whether co-locating rows via KEY distribution outweighs the cost of redistribution at query time. Frequently joined columns justify matching distribution keys across tables, eliminating network traffic during joins. This directly satisfies the stem's constraint of minimising data movement, since infrequent joins rarely repay that co-location overhead.

  • ✗

    The number of columns in the table.

    Why it's wrong here

    Column count does not influence Redshift distribution; the planner distributes rows by key values or round-robin, so width is irrelevant to collocation and join data movement. It is tempting because wider tables do cost more to scan, which matters for columnar storage, not for distribution style selection.

  • ✓

    Whether the table is a fact or dimension table.

    Why this is correct

    Fact tables typically join to dimension tables on foreign keys, so distributing both on the same join key co-locates matching rows on the same slice, eliminating network redistribution during joins. Dimension tables are smaller and often replicated instead. Classifying by fact versus dimension therefore guides the distribution style choice that satisfies the minimal-data-movement constraint.

  • ✗

    The data type of the distribution key column.

    Why it's wrong here

    Distribution keys must be hashable, and Redshift supports common types, so the key's data type does not determine collocation or reduce join data movement. It is tempting because type mismatches can force casting in join predicates, which is a query-design concern rather than a distribution-style factor.

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.