20+ practice questions focused on Data Transformation and Modeling — one of the most tested topics on the Databricks Certified Data Engineer Associate exam. Each question includes a detailed explanation so you learn why the right answer is correct.
Start Data Transformation and Modeling PracticeA data engineer is processing streaming data using Delta Live Tables (DLT) in Python and needs to append incoming records to an existing Delta table without modifying historical records. Which declarative table decorator should be used?
Explanation: Delta Live Tables uses the @dlt.table decorator for both batch and streaming tables. Streaming behavior is defined by using the read_stream() function within the function body. There is no @dlt.append_flow decorator in the DLT Python API.
A data engineer is writing a PySpark script to transform a DataFrame and needs to compute rolling window statistics across ordered time-series events. Which Spark SQL function should be used in combination with window specifications to assign a unique sequential rank to rows within a partition without gaps?
Explanation: The question asks for a function that assigns a unique sequential rank without gaps. Both row_number() and dense_rank() do not have gaps, but row_number() assigns a *unique* sequential rank (1, 2, 3...) to every row even if values are identical, whereas dense_rank() assigns the same rank to ties. The stem specifically requests to 'assign a unique sequential rank to rows within a partition without gaps', which is the exact definition of row_number().
Refer to the exhibit. An engineer is attempting to perform a windowed aggregation on an incoming stream, but the pipeline fails with the provided error. What is the most likely cause of this error?
Explanation: The explanation mentions an 'event_time' column is missing, but the correct option (B) refers to the source data format lacking a timestamp field. Furthermore, the explanation incorrectly suggests that schema evolution prevents this, whereas the issue is a fundamental requirement for windowing operations in Structured Streaming.
A data engineer needs to perform an UPSERT operation on a Delta table. Which command is the correct way to achieve this functionality?
Explanation: In Delta Lake SQL, the 'WHEN MATCHED THEN UPDATE SET *' and 'WHEN NOT MATCHED THEN INSERT *' clauses do not support the '*' wildcard. The columns must be explicitly specified (e.g., 'SET target.col = source.col' or 'INSERT (col1, col2) VALUES (source.col1, source.col2)').
You are auditing a Databricks environment and notice that the Silver tables contain massive amounts of historical data that is no longer needed. Which TWO operations should be used to safely remove this data while keeping the Delta Lake ACID consistency?
Explanation: The DELETE command is used to remove specific rows from a table, which triggers a rewrite of the affected data files and creates a new version in the transaction log. VACUUM is then required to physically remove the old, unreferenced data files from storage to reclaim space. Together, they ensure data is removed while maintaining ACID consistency.
+15 more Data Transformation and Modeling questions available
Practice all Data Transformation and Modeling questions1. Baseline your knowledge
Start with 10 questions to gauge your current understanding of Data Transformation and Modeling. This tells you whether you need a concept refresher or just practice.
2. Review every explanation
For each question — right or wrong — read the full explanation. Understanding why an answer is correct is more valuable than knowing the answer itself.
3. Focus on exam traps
Data Transformation and Modeling questions on the Databricks-DE-Assoc frequently use trap wording. Look for subtle differences in answers that test your precision, not just general knowledge.
4. Reach 80% consistently
Do repeated sessions until you score 80%+ three times in a row. Then move to mixed-mode practice to test cross-topic recall under realistic conditions.
The exact number varies per candidate. Data Transformation and Modeling is tested as part of the Databricks Certified Data Engineer Associate blueprint. Practicing with targeted Data Transformation and Modeling questions ensures you can handle any format or difficulty that appears.
Yes. Courseiva provides free Databricks-DE-Assoc practice questions across all exam topics and domains. The platform includes topic-based practice, mock exams, missed-question review, bookmarked questions, and readiness tracking — no account required.
Difficulty is subjective, but Data Transformation and Modeling is a high-priority exam concept tested in multiple ways — direct recall, scenario analysis, and command-output interpretation. Consistent practice is the best way to build confidence.
Launch a full Data Transformation and Modeling practice session with instant scoring and detailed explanations.
Start Data Transformation and Modeling Practice →