Databricks-DE-Pro Data Modelling Practice Question
A logistics company wants to analyze shipment delays. The fact table `fact_shipments` has a `delay_minutes` measure. The team needs to slice delays by the reason for delay, which can be one of several predefined categories. Which dimension modeling approach is most suitable?
⚠ Common exam trap
The trap here is thinking that storing the reason as a string in the fact table is simpler and sufficient, but it leads to poor query performance and data quality issues.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a dimension table `dim_delay_reason` with a surrogate key and reference it from the fact table.
A dedicated dimension table for delay reasons provides a clean, efficient way to slice shipment delays. It supports consistent categorization and easy addition of new reasons. Storing as a string in the fact table or snowflaking adds performance and maintenance overhead. A junk dimension would obscure the ability to analyze by delay reason alone.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a dimension table `dim_delay_reason` with a surrogate key and reference it from the fact table.
Why this is correct
A dedicated dimension table for delay reasons allows efficient slicing and grouping. It ensures consistency and supports adding new reasons without altering the fact table. This is a standard star schema approach. In Databricks, it can be joined efficiently and optimized with Z-ORDER.
- ✗
Create a snowflake schema by normalizing delay reasons into multiple tables.
Why it's wrong here
Snowflaking delay reasons into multiple tables adds unnecessary joins and complexity. Delay reasons are typically a flat list, so normalization is not needed. This would degrade query performance and complicate the model. A single dimension table is sufficient.
- ✗
Use a junk dimension to combine delay reason with other low-cardinality attributes.
Why it's wrong here
A junk dimension is used to group multiple low-cardinality flags to reduce the number of dimensions. Here, delay reason is a single attribute that analysts want to slice by. Combining it with other attributes would make it harder to query independently. It is not the appropriate design for this requirement.
- ✗
Store the delay reason as a string column directly in the fact table.
Why it's wrong here
Storing the reason as a string in the fact table increases storage and reduces query performance due to string comparisons. It also risks data inconsistency if not standardized. While it simplifies ETL, it is not a good modeling practice for analytical queries that group by reason. A dimension table is preferred.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.