Databricks-DE-Assoc · domain
Working with Lakeflow Jobs
This domain covers orchestrating data pipelines with Lakeflow Jobs in Databricks: notebook and task configuration, dependencies, retries, scheduling, and failure recovery. Questions present failure scenarios or multi-task DAGs and ask you to pick the correct setting, dependency behavior, or repair outcome, so you must know how task-level options interact.
Focused practice
Practice Working with Lakeflow Jobs questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Working with Lakeflow Jobs
Be able to configure a multi-task Lakeflow Job: set task dependencies, add task-level Retries for transient failures, and use Repair and Rerun to recover only failed portions. The key is knowing default dependency behavior: downstream tasks wait for all upstream parents to succeed.
Configuring task-level Retries and timeouts for notebook tasks hitting transient external API errors
Using 'Repair and Rerun' to rerun only failed and downstream tasks while preserving successful task results
Setting task dependencies so downstream tasks run only when all upstream parents succeed
Understanding how Lakeflow Jobs orchestrate multi-task workflows with run-if conditions and scheduling
Watch out for
Common Working with Lakeflow Jobs exam traps
- ▸Assuming Repair and Rerun reruns the entire job; it only reruns failed tasks and their downstream dependents, reusing prior successful task outputs.
- ▸Confusing task-level Retries with job-level retries or cluster restart policies, leading to wrong answers about where retry behavior is configured.
- ▸Believing downstream tasks run by default when a parent fails; default dependency requires all parents to succeed unless a run-if condition is set.
Question index
All Working with Lakeflow Jobs questions (30)
Click any question to see the full explanation, or start a practice session above.
Which TWO of the following statements are true regarding the behavior and capabilities of Databricks Jobs parameters and values?
Hard2A data engineer is configuring a Lakeflow Job that processes sensitive customer data. The job must notify the on-call team when a run fails and must also capture the run's output for auditing. Which TWO actions should the engineer take in the Lakeflow Jobs configuration? (Choose two.)
Medium3A data engineer has a Lakeflow Job with two tasks: Task1 and Task2. Task2 must run only if Task1 succeeds. The engineer also wants Task2 to be skipped if Task1 fails, but the overall job status should be marked as failed. Which configuration should the engineer use for the dependency between Task1 and Task2?
Hard4A data engineer needs to pass the execution date to a job task dynamically. Which feature should they use?
Medium5A data engineer needs to configure a Databricks Job containing multiple tasks where downstream tasks should only execute if all upstream parent tasks complete successfully. Which task dependency setting should be configured?
Medium6A data engineer is creating a Lakeflow Job that must run a Python script stored in DBFS. The engineer wants to ensure the script is executed with the correct dependencies and environment. Which task type should be used?
Easy7A data engineer is building a Lakeflow Job that must process a parameterized date range. The engineer wants to pass start_date and end_date values into a notebook task at runtime and have those values available as widget-like parameters inside the notebook. Which approach should the engineer use?
Medium8A data engineering team runs a nightly Lakeflow Job that ingests files from cloud storage, transforms them with a notebook, and then runs a SQL task. The team wants the SQL task to execute only after the notebook transform succeeds, but they do not want the SQL task to wait for a fixed delay. Which Lakeflow Jobs feature should they configure on the SQL task?
Medium9Which THREE of the following are benefits of using Delta Live Tables (DLT) for managing your data pipelines?
Hard10A data engineer has a Lakeflow Job with three tasks: bronze_ingest, silver_transform, and gold_aggregate. The silver_transform task must run only if bronze_ingest succeeds, and gold_aggregate must run only if silver_transform succeeds. The engineer also wants gold_aggregate to run even if silver_transform fails, so that partial results can be published. Which configuration should the engineer apply to gold_aggregate?
Hard11Which of the following is the primary benefit of using a 'Job Cluster' rather than an 'All-Purpose Cluster' for running scheduled data pipelines?
Easy12Refer to the exhibit. If 'task1' fails due to a timeout, what happens to 'task2'?
Hard13A data engineer is creating a Lakeflow Job that must run a notebook every weekday at 06:00 in the company's local time zone, which is America/New_York. The engineer configures a schedule trigger but the job runs at the wrong time. Which setting should the engineer verify first?
Easy14A data engineer has a Lakeflow Job that runs daily. They want to receive an email only when the job fails, not on every run. Which notification configuration should they set?
Easy15A data engineer wants to pass a file path from Task A to Task B in a Lakeflow Job. Task A is a notebook that computes the path, and Task B is a notebook that reads from that path. Which mechanism should the engineer use to share the value between tasks?
Medium16A data engineer has a Lakeflow Job with a linear dependency chain: Task A, then Task B, then Task C. Task B sometimes fails due to transient errors. The engineer wants Task C to run only if Task B succeeds, but also wants Task B to be retried automatically before considering the job failed. Which configuration should they use?
Hard17Refer to the exhibit. When is this job scheduled to run?
Hard18A data engineer is setting up a Lakeflow Job that runs a notebook task. The engineer needs the task to always execute even if the upstream task in the workflow fails. Which configuration should be applied to the dependent task's condition?
Medium19A data engineer is designing a Databricks Job workflow. Which TWO of the following are valid ways to trigger a Databricks Job?
Medium20A data engineer wants to ensure that a Databricks Job task only runs if the preceding task completes successfully, but needs to add a specific timeout threshold for this individual task. Where should this configuration be applied?
Medium21A data engineer configures a Lakeflow Job to run a notebook task on a job cluster. The notebook reads a parameter named run_date using the widget API. During a manual run, the engineer wants to supply a specific date without editing the notebook. The job also runs on a nightly schedule where the date should default to the current day. Which approach correctly supplies the parameter for both the manual and scheduled runs?
Hard22A data engineer has configured a Databricks Job with multiple dependent tasks forming a linear pipeline. Task A extracts data, Task B transforms it, and Task C loads it into a gold table. The pipeline runs daily. The team notices that if Task B fails due to an intermittent schema validation issue, the entire job run fails, but they want Task C to execute conditionally only if Task B succeeds, while alerting the on-call engineer immediately upon any failure. How should the task dependencies and conditional execution be configured?
Medium23A data engineer has a Lakeflow Job with a notebook task that occasionally fails due to transient network errors when reading from an external REST API. The engineer wants the task to automatically retry up to three times, but only for this specific task, without affecting other tasks in the job. What should the engineer do?
Medium24A data engineer is building a Lakeflow Job that must run a sequence of tasks across different compute types. The ingest task must run on a job cluster with a specific Spark configuration, the transform task must run as a Delta Live Tables pipeline, and the report task must run on a separate SQL warehouse. Which TWO statements about task-level compute configuration in Lakeflow Jobs are correct? (Choose two.)
Hard25A data engineer is building a Lakeflow Job with a task that runs a SQL notebook. The task must run only on weekdays and must be completed before 9 AM. The engineer wants to configure the schedule to meet these requirements. What should the engineer do?
Medium26What is the primary function of the 'Retries' setting in a Databricks Job task?
Medium27A data engineer needs to configure a Databricks Job containing three distinct tasks: ingest, transform, and report. The transform task must only execute if the ingest task completes successfully, but the report task should execute regardless of whether the transform task succeeds or fails. How should the task dependencies be configured?
Medium28A data engineer manages a Lakeflow Job that runs a long-running notebook task on a job cluster. The task occasionally fails due to transient cloud storage errors, and the engineer wants the task to retry automatically without failing the entire job on the first attempt. The engineer also wants to be alerted only if all retries are exhausted. Which configuration should the engineer apply?
Medium29When a Data Engineer uses a 'Repair and Rerun' functionality on a failed Databricks Job, what happens?
Medium30A data engineer maintains a Lakeflow Job with a scheduled trigger set to run every day at 08:00. The job's source table is refreshed by an upstream process that sometimes finishes later than expected, causing the job to process stale data. The engineer wants the job to start only after the upstream refresh completes, regardless of the clock time, while still preserving the existing 08:00 schedule as a fallback. Which trigger configuration should the engineer implement?
MediumOther domains
All Databricks-DE-Assoc exam domains
Frequently asked questions
- What does the Working with Lakeflow Jobs domain cover on the Databricks-DE-Assoc exam?
- Be able to configure a multi-task Lakeflow Job: set task dependencies, add task-level Retries for transient failures, and use Repair and Rerun to recover only failed portions. The key is knowing default dependency behavior: downstream tasks wait for all upstream parents to succeed.
- How many questions are in this domain?
- This page lists all 30 Working with Lakeflow Jobs questions in the Databricks-DE-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Working with Lakeflow Jobs questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.