Databricks-DE-Pro · domain
Data Modelling
This domain covers dimensional and Lakehouse modeling on Databricks: Medallion (Bronze/Silver/Gold) layering, Delta Lake table design, and physical layout choices. Questions present a workload (filter columns, cardinality, streaming vs batch) and ask you to pick the right table structure, layer, or optimization to satisfy it.
Focused practice
Practice Data Modelling questions
Scored sessions drawing only from this domain — pick a length below.
What this domain covers
What to know about Data Modelling
A candidate must translate a described query workload into a concrete Delta Lake table design: correct Medallion layer, grain, partitioning, and clustering. The single most important thing is matching layout choices to the actual filter and join columns, not to intuition.
Choosing Bronze, Silver, or Gold layer responsibilities and data quality expectations
Designing Delta Lake fact and dimension tables with appropriate grain and keys
Applying OPTIMIZE, Z-ORDER, liquid clustering, and partitioning for query patterns
Using Delta Lake features like MERGE, time travel, and generated columns in models
Watch out for
Common Data Modelling exam traps
- ▸Partitioning on a high-cardinality column such as user_id or transaction_timestamp, which creates too many small files and hurts performance.
- ▸Treating the Gold layer as raw ingestion storage instead of curated, business-level aggregates and marts.
- ▸Ignoring data skipping and file layout, so queries filtering on event_date scan far more data than necessary.
Question index
All Data Modelling questions (22)
Click any question to see the full explanation, or start a practice session above.
Refer to the exhibit. The engineer wants to replace only one specific partition in the 'orders' table. What is the best method in Databricks?
Medium2Which of the following describes the purpose of the 'Gold' layer in a Lakehouse?
Medium3A company requires data to be physically deleted from the Bronze layer for GDPR compliance. What is the correct procedure to ensure complete removal?
Hard4A financial services firm maintains a Delta Lake table of account transactions that must support both current-state queries and full audit history of every change, including corrections that arrive days later. Regulators require the ability to query the table as it existed at any prior date. Which Delta Lake capability should the engineer rely on to satisfy the audit requirement?
Medium5What is the primary benefit of the Medallion architecture in a Databricks Lakehouse?
Easy6An engineer is building a Gold-layer star schema for a sales analytics workload. The business wants to analyze revenue by product, by store, and by promotion independently, and also drill down through a hierarchy of region to country to city. Which dimensional modeling structure best supports these requirements?
Easy7A Databricks workspace has a Delta table 'transactions' partitioned by 'txn_date'. Analysts frequently run queries that filter on 'txn_date' but also occasionally filter on 'account_id' alone. The table has 10 TB of data, and the team wants to improve performance for the 'account_id' queries without changing the partitioning scheme. Which Delta feature should they implement?
Medium8A retail company uses a Databricks Lakehouse with a star schema in the Gold layer. Their fact_sales table has billions of rows and is partitioned by sale_date. Analysts frequently run queries that filter on product_id and join to dim_product. Currently, queries scanning the entire fact table are slow. To improve performance for these queries, which approach is most appropriate?
Medium9A financial institution uses a Databricks Lakehouse with a Silver table transactions that is partitioned by transaction_date. The table is frequently queried with filters on transaction_date and account_id. The data engineering team notices that queries filtering on account_id are slow because they scan all partitions. They want to optimize the table to accelerate these queries without repartitioning. Which Delta Lake feature should they use?
Hard10A retail company is designing a Gold-layer dimension table in Delta Lake for its product catalog. The catalog changes slowly: a product's category is occasionally reclassified, but historical sales fact rows must continue to reflect the category that was valid at the time of each sale. The team wants to avoid duplicating the entire product row for every change. Which Delta Lake modeling technique should the engineer implement?
Medium11An engineer needs to optimize a massive table that is frequently joined with other large tables. Which strategy is most effective for performance?
Hard12Which design pattern is best suited for handling late-arriving data in a medallion architecture?
Medium13A logistics company wants to analyze shipment delays. The fact table `fact_shipments` has a `delay_minutes` measure. The team needs to slice delays by the reason for delay, which can be one of several predefined categories. Which dimension modeling approach is most suitable?
Medium14In the medallion architecture, which layer is primarily responsible for applying business logic and historical aggregations?
Easy15A data engineer is designing a Gold layer table for a retail company. The table must support efficient queries that filter on product_category (low cardinality) and sort by transaction_timestamp (high cardinality). The table is expected to grow to petabytes. Which Delta Lake table design should the engineer choose to optimize both filtering and sorting?
Medium16A data engineer is designing a Gold layer table that must support slowly changing dimension (SCD) Type 2 for a customer dimension. The source data arrives daily with updates to customer attributes. The engineer wants to implement this using Delta Lake. Which two features or techniques are essential for maintaining SCD Type 2? (Choose two.)
Hard17A retail company wants to analyze sales by product, store, and date. The data team is designing the Gold layer and needs to choose between a star schema and a snowflake schema. Which factor most strongly favors a star schema in a Databricks Lakehouse?
Easy18A healthcare company uses a Databricks Lakehouse. The Silver layer contains a table patient_visits that is updated with late-arriving data. The table is partitioned by visit_date. The data engineering team needs to efficiently merge new data that may include updates to existing records and inserts of new records. They want to minimize the impact on existing data and ensure ACID compliance. Which Delta Lake operation should they use?
Hard19A data engineer is building a Silver layer table that combines data from multiple Bronze tables. The engineer wants to ensure that the Silver table only contains the most recent version of each record based on a 'last_updated' timestamp. Which Delta Lake operation should be used to achieve this?
Easy20Refer to the exhibit. An engineer observes that queries filtering on 'customer_id' are running slowly despite Z-Ordering. What is the most likely cause?
Medium21A data engineering team is modeling a large Delta Lake fact table that stores clickstream events. Analysts frequently run queries that filter by event_date and then aggregate by user_id, and the table receives continuous appends plus occasional late-arriving corrections. The team wants to reduce bytes scanned and improve join performance. Which two design choices are most appropriate? (Choose two.)
Hard22A financial institution is building a Gold layer table that must support point-in-time queries to reconstruct account balances as of any past date. The source data includes transactions with effective dates and an audit log of changes. Which modeling technique is most appropriate?
MediumOther domains
All Databricks-DE-Pro exam domains
Frequently asked questions
- What does the Data Modelling domain cover on the Databricks-DE-Pro exam?
- A candidate must translate a described query workload into a concrete Delta Lake table design: correct Medallion layer, grain, partitioning, and clustering. The single most important thing is matching layout choices to the actual filter and join columns, not to intuition.
- How many questions are in this domain?
- This page lists all 22 Data Modelling questions in the Databricks-DE-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Modelling questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.