20+ practice questions focused on Data Modelling — one of the most tested topics on the Databricks Certified Data Engineer Professional exam. Each question includes a detailed explanation so you learn why the right answer is correct.
Start Data Modelling PracticeA data engineer is designing a Bronze-to-Silver pipeline for high-velocity IoT sensor data. The raw JSON logs arrive with varying schemas. Which approach best supports schema evolution while maintaining query performance in Silver?
Explanation: Using schema evolution in Delta Lake allows the Silver layer to absorb new fields without pipeline failure, while enforcing schema constraints ensures data quality. By leveraging Delta's schema evolution capabilities alongside compacting small files during the write process, the engineer ensures that downstream medallion layers remain performant. This design is critical in Databricks environments where source systems frequently update their payloads without notifying downstream consumers.
Which TWO of the following strategies best optimize a large-scale fact table in the Gold layer for point-in-time analytical queries?
Explanation: The explanation justifies partitioning by date, but Option A (partitioning) is marked as wrong. The correct options B and C are valid, but the explanation needs to be updated to focus on Z-Ordering and Star Schema design rather than partitioning.
Which THREE factors should be considered when choosing a partitioning strategy for a Delta table?
Explanation: Option A states to avoid low-cardinality columns, which is actually correct for partitioning (you should avoid high-cardinality columns like timestamps/IDs because they cause the small file problem). The common trap note incorrectly advises to 'always prioritize low-cardinality columns', which contradicts best practices and the explanation.
Refer to the exhibit. The pipeline is failing during a transformation. What is the most likely cause, and how should it be resolved?
Explanation: The failure is caused by a schema mismatch where the source data contains new fields not present in the target Delta table. Delta Lake enforces schema by default to prevent accidental data corruption. To resolve this, schema evolution must be enabled (e.g., using .option('mergeSchema', 'true') in Spark or the evolution mode in Delta Live Tables), allowing the target table to automatically update its schema to match the incoming data.
A retail company uses a Databricks Lakehouse. The dimension table `dim_product` has a high rate of updates, and the fact table `fact_sales` is huge. Analysts frequently run queries that join these tables and filter on product attributes. The team wants to optimize this pattern. Which two design choices are appropriate? (Choose two.)
Explanation: Using a star schema with surrogate keys ensures stable joins and simplifies dimension updates, while Z-ORDERing on join keys co-locates data to reduce shuffling during joins. Together, these choices optimize the frequent join and filter pattern in Databricks. Denormalizing or snowflaking adds complexity and does not address the performance need. Partitioning by category is not effective for join optimization.
+15 more Data Modelling questions available
Practice all Data Modelling questions1. Baseline your knowledge
Start with 10 questions to gauge your current understanding of Data Modelling. This tells you whether you need a concept refresher or just practice.
2. Review every explanation
For each question — right or wrong — read the full explanation. Understanding why an answer is correct is more valuable than knowing the answer itself.
3. Focus on exam traps
Data Modelling questions on the Databricks-DE-Pro frequently use trap wording. Look for subtle differences in answers that test your precision, not just general knowledge.
4. Reach 80% consistently
Do repeated sessions until you score 80%+ three times in a row. Then move to mixed-mode practice to test cross-topic recall under realistic conditions.
The exact number varies per candidate. Data Modelling is tested as part of the Databricks Certified Data Engineer Professional blueprint. Practicing with targeted Data Modelling questions ensures you can handle any format or difficulty that appears.
Yes. Courseiva provides free Databricks-DE-Pro practice questions across all exam topics and domains. The platform includes topic-based practice, mock exams, missed-question review, bookmarked questions, and readiness tracking — no account required.
Difficulty is subjective, but Data Modelling is a high-priority exam concept tested in multiple ways — direct recall, scenario analysis, and command-output interpretation. Consistent practice is the best way to build confidence.
Launch a full Data Modelling practice session with instant scoring and detailed explanations.
Start Data Modelling Practice →