Courseiva

CCNA Working Lakeflow Jobs Questions

30 questions · Working Lakeflow Jobs topic · All types, answers revealed

1
Multi-Selecthard

Which TWO of the following statements are true regarding the behavior and capabilities of Databricks Jobs parameters and values?

Select 2 answers
A.Task values can be used to pass small amounts of data, such as counts or status codes, from one task to a downstream task.
B.Job parameters defined at the workflow level can be overridden when triggering a run via the Databricks CLI or REST API.
C.Job parameters are automatically persisted in a Unity Catalog volume after the workflow completes successfully.
D.Notebook tasks cannot accept parameter values passed from the parent job orchestration configuration.
E.Task values support passing large dataframes containing millions of rows between dependent tasks.
AnswersA, B

The dbutils.jobs.taskValues API allows tasks to set key-value pairs that downstream tasks can retrieve. This facilitates dynamic branching and parameter passing based on runtime metrics generated during the execution of upstream workflow steps.

Why this answer

Databricks Jobs support parameterized runs, enabling dynamic workflows where values are passed at runtime or configured statically. Understanding how parameters propagate to tasks and how task values pass outputs between tasks is essential for building modular, reusable data engineering pipelines within the Lakeflow framework.

Exam trap

Candidates often think job-level parameters are immutable and cannot be altered at runtime, missing the flexibility of the Databricks API and CLI overrides.

2
Multi-Selectmedium

A data engineer is configuring a Lakeflow Job that processes sensitive customer data. The job must notify the on-call team when a run fails and must also capture the run's output for auditing. Which TWO actions should the engineer take in the Lakeflow Jobs configuration? (Choose two.)

Select 2 answers
A.Add an email notification for job failure that targets the on-call distribution list.
B.Configure a retry policy with three retries to ensure the on-call team is notified.
C.Add a task-level timeout so the job fails faster and generates an alert.
D.Enable the job's run history and export logs to a durable storage location for auditing.
E.Set the job's maximum concurrent runs to 1 to prevent overlapping audit records.
AnswersA, D

Email notifications configured for the failure event send an alert to specified recipients when the job run fails. Targeting the on-call distribution list ensures the team is paged without manual monitoring. This directly satisfies the requirement to notify the team on failure and is a native Lakeflow Jobs notification option.

Why this answer

Failure notifications deliver timely alerts to the on-call team, and preserving run history with exported logs provides the audit trail. Together these native Lakeflow Jobs capabilities meet both the alerting and auditing requirements without adding unnecessary concurrency or retry complexity that would not notify or record anything on its own.

Exam trap

The trap here is conflating reliability settings such as retries and concurrency limits with notification and audit capabilities, which are separate configuration areas.

3
MCQhard

A data engineer has a Lakeflow Job with two tasks: Task1 and Task2. Task2 must run only if Task1 succeeds. The engineer also wants Task2 to be skipped if Task1 fails, but the overall job status should be marked as failed. Which configuration should the engineer use for the dependency between Task1 and Task2?

A.Set Task2 to depend on Task1 with the 'All done' condition.
B.Set Task2 to depend on Task1 with the 'At least one succeeded' condition.
C.Set Task2 to depend on Task1 with the 'None failed' condition.
D.Set Task2 to depend on Task1 with the 'All succeeded' condition.
AnswerD

The 'All succeeded' condition means Task2 runs only if Task1 completes successfully. If Task1 fails, Task2 is skipped, and the job is marked as failed because a task failed. This matches the requirement: Task2 runs only on success, and the job fails if Task1 fails.

Why this answer

The 'All succeeded' condition ensures Task2 runs only when Task1 succeeds. If Task1 fails, Task2 is skipped, and the job is marked as failed because a task failed. This precisely meets the requirement of conditional execution and overall job failure.

Exam trap

The trap here is confusing 'All succeeded' with 'None failed' or 'All done', which have different semantics for skip and failure propagation.

4
MCQmedium

A data engineer needs to pass the execution date to a job task dynamically. Which feature should they use?

A.Hardcoding the date in the notebook.
B.Defining job parameters in the task configuration.
C.Creating a new cluster for each date.
D.Using global workspace variables.
AnswerB

Task-level parameters allow values to be passed to notebooks or scripts at execution time. These parameters can be referenced within the code using standard widget APIs or environment variable lookups, providing a clean and programmatic way to inject dynamic information into the task execution without altering the underlying logic.

Why this answer

Using Job Parameters allows engineers to inject dynamic values into tasks at runtime. This capability is crucial for backfilling data or running incremental pipelines where the processing logic depends on specific dates or data partitions. By parameterizing tasks, engineers create reusable jobs that can handle varying datasets without requiring manual code changes, thereby improving the maintainability and scalability of the entire data pipeline architecture.

Exam trap

Candidates often confuse task parameters with cluster-level environment variables or hardcoded notebook widgets, failing to use the native job task configuration feature.

5
MCQmedium

A data engineer needs to configure a Databricks Job containing multiple tasks where downstream tasks should only execute if all upstream parent tasks complete successfully. Which task dependency setting should be configured?

A.Configure a retry policy with exponential backoff on the parent task to guarantee eventual success.
B.Set up a global workspace webhook to trigger downstream tasks manually after the parent completes.
C.Add the upstream tasks as parents of the downstream task directly within the job configuration graph.
D.Enable concurrent run suppression on the job level to prevent overlapping workflow executions.
AnswerC

Declaring upstream tasks as parents of the downstream task in the job graph creates explicit dependencies, so Databricks triggers the downstream task only after every parent completes successfully. This directly enforces the all-upstream-success condition the scenario requires.

Why this answer

To ensure tasks execute conditionally based on the success of parent tasks, you configure task dependencies within the Databricks Jobs UI or JSON definition. By default, adding a parent task creates a strict dependency where downstream tasks trigger only upon successful completion. This mechanism orchestrates complex directed acyclic graphs for robust, reliable data pipelines.

Exam trap

Candidates sometimes try to manage task sequencing using external orchestration tools instead of native Databricks dependency graphs within the job configuration.

6
MCQeasy

A data engineer is creating a Lakeflow Job that must run a Python script stored in DBFS. The engineer wants to ensure the script is executed with the correct dependencies and environment. Which task type should be used?

A.JAR task
B.Notebook task
C.Pipeline task
D.Python script task
AnswerD

The Python script task type in Lakeflow Jobs allows you to specify a Python file stored in DBFS, Unity Catalog volumes, or cloud storage. It executes the script with the specified libraries and environment, making it ideal for running standalone Python scripts with dependencies. This directly matches the requirement.

Why this answer

Lakeflow Jobs offer a dedicated Python script task type that runs a Python file from DBFS or other storage locations. It supports specifying libraries and environment settings, ensuring the script runs with the correct dependencies. Notebook, JAR, and pipeline tasks serve different purposes and cannot directly execute a standalone Python script with its own environment.

Exam trap

The trap here is assuming a notebook task can run any Python file; notebook tasks specifically require a notebook format and do not execute standalone .py scripts.

7
MCQmedium

A data engineer is building a Lakeflow Job that must process a parameterized date range. The engineer wants to pass start_date and end_date values into a notebook task at runtime and have those values available as widget-like parameters inside the notebook. Which approach should the engineer use?

A.Define job parameters and reference them in the notebook using the task's base parameters, then read them with dbutils.widgets.get in the notebook.
B.Store the dates in a Delta table and have the notebook query that table at startup.
C.Use the notebook's %run magic to pass arguments from a parent notebook.
D.Set environment variables on the job cluster and read them with os.environ in the notebook.
AnswerA

Job parameters and task base parameters are passed to notebook tasks as widget values. The notebook can retrieve them with dbutils.widgets.get using the parameter key. This is the native way to parameterize a Lakeflow Job task and makes start_date and end_date available at runtime without hardcoding values.

Why this answer

Lakeflow Jobs parameters are surfaced to notebook tasks as widget values, so the notebook can call dbutils.widgets.get with the parameter name. Defining start_date and end_date as job or task base parameters makes them available per run and visible in the job UI, enabling dynamic date-range processing without hardcoding or external lookups.

Exam trap

The trap here is assuming that cluster environment variables or external tables are the primary way to pass runtime values, when Lakeflow Jobs exposes parameters to notebooks as widgets.

8
MCQmedium

A data engineering team runs a nightly Lakeflow Job that ingests files from cloud storage, transforms them with a notebook, and then runs a SQL task. The team wants the SQL task to execute only after the notebook transform succeeds, but they do not want the SQL task to wait for a fixed delay. Which Lakeflow Jobs feature should they configure on the SQL task?

A.Configure a time-based trigger on the SQL task using a cron expression that starts two minutes after the job run begins.
B.Set the SQL task's depends_on property to reference the notebook transform task.
C.Add a run_if condition to the SQL task that checks whether the transform task's output path exists.
D.Place the SQL task and the notebook transform task in separate jobs and chain them using a job-level trigger.
AnswerB

The depends_on property on a task creates an explicit dependency on another task in the same job. When the SQL task lists the notebook transform task in depends_on, the SQL task will not start until the transform task completes successfully. This enforces the required ordering without introducing a fixed delay or polling mechanism.

Why this answer

The depends_on property is the native Lakeflow Jobs mechanism for expressing task-level dependencies. By setting depends_on on the SQL task to point at the notebook transform task, the SQL task is scheduled only after the transform completes successfully. This avoids fixed delays, avoids polling, and keeps the pipeline deterministic and observable within a single job run.

Exam trap

The trap here is assuming that scheduling a task later in time or checking for an output file creates a dependency, when Lakeflow Jobs requires an explicit depends_on relationship to order tasks.

9
Multi-Selecthard

Which THREE of the following are benefits of using Delta Live Tables (DLT) for managing your data pipelines?

Select 3 answers
A.Automated data quality monitoring via expectations.
B.Manual management of cluster provisioning.
C.Simplified dependency management via declarative pipelines.
D.Automatic data lineage and observation.
E.Ability to use non-Delta file formats.
AnswersA, C, D

Expectations allow users to define data quality rules directly in the DLT pipeline code. DLT automatically tracks these metrics, allowing engineers to quarantine or fail records that don't meet requirements, which is a critical feature for building trust in data products within the lakehouse architecture environment.

Why this answer

Delta Live Tables simplifies the complexity of building reliable data pipelines by handling infrastructure orchestration, quality checks, and dependency management automatically. Understanding these benefits is essential for modern data engineering, as DLT reduces the burden of manual configuration, ensures data reliability through expectations, and provides declarative syntax that allows engineers to focus on business logic rather than low-level Spark optimization or cluster management tasks.

Exam trap

Candidates sometimes assume DLT requires manual Spark cluster tuning and custom orchestration, missing that it provides automated management and declarative syntax.

10
MCQhard

A data engineer has a Lakeflow Job with three tasks: bronze_ingest, silver_transform, and gold_aggregate. The silver_transform task must run only if bronze_ingest succeeds, and gold_aggregate must run only if silver_transform succeeds. The engineer also wants gold_aggregate to run even if silver_transform fails, so that partial results can be published. Which configuration should the engineer apply to gold_aggregate?

A.Remove depends_on from gold_aggregate and set run_if to ALL_DONE.
B.Set gold_aggregate's depends_on to include bronze_ingest and set run_if to AT_LEAST_ONE_SUCCESS.
C.Set gold_aggregate's depends_on to include silver_transform and set run_if to ALL_SUCCESS.
D.Set gold_aggregate's depends_on to include silver_transform and set run_if to ALL_DONE.
AnswerD

depends_on establishes the ordering so gold_aggregate waits for silver_transform. Setting run_if to ALL_DONE causes gold_aggregate to run regardless of whether silver_transform succeeded or failed, as long as upstream tasks have reached a terminal state. This combination satisfies both the ordering and the run-even-on-failure requirement.

Why this answer

Ordering and run conditions are separate controls. depends_on places gold_aggregate after silver_transform, and run_if set to ALL_DONE allows the task to execute once upstream tasks reach a terminal state, whether they succeeded or failed. This is the standard pattern for publishing partial results or running cleanup tasks after a failure.

Exam trap

The trap here is assuming that run_if alone controls ordering, when depends_on must also be present to define the dependency edge that run_if evaluates.

11
MCQeasy

Which of the following is the primary benefit of using a 'Job Cluster' rather than an 'All-Purpose Cluster' for running scheduled data pipelines?

A.It allows interactive development in notebooks.
B.It provides lower pricing for compute resources.
C.It supports faster library installation.
D.It caches data indefinitely for all users.
AnswerB

Databricks charges lower DBUs for job clusters because they are ephemeral, non-interactive resources. This pricing model encourages the use of automated workflows by lowering the barrier to entry, allowing engineers to scale their data processing tasks without the high overhead costs associated with keeping interactive clusters running constantly.

Why this answer

Job clusters are ephemeral and specifically designed for automated tasks, offering significant cost savings compared to all-purpose clusters. By terminating automatically after the job completes, they eliminate idle compute time costs. Understanding this distinction is essential for data engineers to balance performance requirements with budgetary constraints, as choosing the correct compute type directly impacts the overall operational efficiency and cost-effectiveness of the data platform.

Exam trap

Candidates mistakenly choose 'higher performance' or 'faster startup times' as the primary benefit, failing to realize that Job Clusters are specifically optimized for cost-efficiency in automated, non-interactive workloads.

12
MCQhard

Refer to the exhibit. If 'task1' fails due to a timeout, what happens to 'task2'?

A.It will run anyway because it is independent.
B.It will be skipped by the job scheduler.
C.It will pause until task1 is restarted.
D.It will run with a 'warning' status.
AnswerB

The scheduler evaluates the DAG status before executing each task. When a prerequisite task fails, the scheduler marks the dependent task as skipped. This prevents the execution of logic that relies on uncomputed results, maintaining the integrity of the data lineage and ensuring consistent output for downstream users.

Why this answer

In Databricks Jobs, dependencies are strictly evaluated based on the success of preceding tasks. If a task fails—regardless of whether it was a timeout or a code error—the dependent tasks in the directed acyclic graph (DAG) will not be triggered. This behavior ensures that downstream data quality is protected from incomplete or stale upstream data processing, which is a fundamental requirement for maintaining reliable data pipelines.

Exam trap

Candidates often assume that 'task2' will still run or that the job will retry automatically, ignoring the fact that downstream tasks in a DAG are strictly dependent on the success of upstream tasks.

13
MCQeasy

A data engineer is creating a Lakeflow Job that must run a notebook every weekday at 06:00 in the company's local time zone, which is America/New_York. The engineer configures a schedule trigger but the job runs at the wrong time. Which setting should the engineer verify first?

A.The job's time zone setting, which controls how the cron expression is interpreted.
B.The notebook's default language, because Python and SQL notebooks evaluate cron expressions differently.
C.The job's maximum concurrent runs setting, because concurrency delays the start time.
D.The cluster's Spark configuration, because executor time zones affect cron evaluation.
AnswerA

Lakeflow Jobs schedules use a cron expression plus a time zone. If the time zone is left at UTC while the team expects America/New_York, the job fires at the wrong local hour. Verifying and setting the schedule time zone to America/New_York ensures the cron expression is evaluated in the intended local time.

Why this answer

The schedule time zone determines how the cron expression is interpreted. If the job is configured with a cron for 06:00 but the time zone is UTC, the run occurs at 06:00 UTC, which is a different local time. Setting the schedule time zone to America/New_York aligns the trigger with the team's expectation and avoids manual conversion mistakes.

Exam trap

The trap here is assuming that cluster or notebook settings influence scheduling, when the schedule time zone is the setting that controls cron interpretation.

14
MCQeasy

A data engineer has a Lakeflow Job that runs daily. They want to receive an email only when the job fails, not on every run. Which notification configuration should they set?

A.Add an email notification with the condition 'On all events' for the job.
B.Add an email notification with the condition 'On failure' for the job.
C.Add an email notification with the condition 'On success' for the job.
D.Add an email notification with the condition 'On start' for the job.
AnswerB

Lakeflow Jobs allow notification settings with conditions such as 'On failure', 'On success', or 'On start'. Selecting 'On failure' ensures that email alerts are sent only when the job fails, which matches the requirement. This avoids noisy notifications on successful runs and focuses attention on issues that need remediation.

Why this answer

Notification conditions in Lakeflow Jobs determine when alerts are sent. To receive an email only when the job fails, the condition must be set to 'On failure'. Other conditions such as 'On success', 'On start', or 'On all events' would send notifications at times that do not match the failure-only requirement, leading to unnecessary alerts.

Exam trap

The trap here is selecting a broader notification condition like 'On all events' or confusing 'On start' with failure alerts.

15
MCQmedium

A data engineer wants to pass a file path from Task A to Task B in a Lakeflow Job. Task A is a notebook that computes the path, and Task B is a notebook that reads from that path. Which mechanism should the engineer use to share the value between tasks?

A.Use job parameters: Task A writes the path to a job parameter, and Task B reads it.
B.Use the job's run ID: Task A writes the path to a Delta table keyed by run ID, and Task B queries it.
C.Use task values: Task A sets a task value with dbutils.jobs.taskValues.set, and Task B retrieves it with dbutils.jobs.taskValues.get.
D.Use a global variable in the notebook: Task A sets a global variable, and Task B reads it.
AnswerC

Task values are designed for passing small values between tasks in a Lakeflow Job. Task A can set a task value using dbutils.jobs.taskValues.set, and Task B can retrieve it using dbutils.jobs.taskValues.get, specifying the originating task and key. This is the recommended way to share dynamic values like file paths within a job run.

Why this answer

Task values are the intended feature for passing small, dynamically computed values between tasks in a Lakeflow Job. Task A sets a value using dbutils.jobs.taskValues.set, and Task B retrieves it with dbutils.jobs.taskValues.get, referencing the task that set it. Job parameters are read-only, global variables are not shared across tasks, and using a Delta table is unnecessarily complex for this use case.

Exam trap

The trap here is assuming job parameters can be written by a task, when they are actually read-only inputs set at submission time.

16
MCQhard

A data engineer has a Lakeflow Job with a linear dependency chain: Task A, then Task B, then Task C. Task B sometimes fails due to transient errors. The engineer wants Task C to run only if Task B succeeds, but also wants Task B to be retried automatically before considering the job failed. Which configuration should they use?

A.Set Task B's Retries to 2 and configure Task C to depend on Task B with the 'At least one failed' condition.
B.Set Task B's Retries to 0 and configure Task C to depend on Task B with the 'At least one succeeded' condition.
C.Set Task B's Retries to 2 and configure Task C to depend on Task B with the default 'All succeeded' condition.
D.Set Task B's Retries to 2 and configure Task C to depend on Task B with the 'All done' condition.
AnswerC

Task B's Retries setting will automatically reattempt the task if it fails, up to the specified count. Task C's default dependency condition 'All succeeded' means it runs only if Task B ultimately succeeds after retries. If Task B fails after all retries, Task C will not run, and the job will be marked as failed. This matches the requirement exactly.

Why this answer

To automatically retry Task B on transient failures, set its Retries to a positive number. To ensure Task C runs only if Task B succeeds, use the default dependency condition 'All succeeded' for Task C. This combination provides retries for Task B and conditional execution for Task C.

Other conditions like 'All done' or 'At least one failed' would cause Task C to run in undesired situations.

Exam trap

The trap here is choosing a dependency condition that ignores success, such as 'All done', while focusing only on retries.

17
MCQhard

Refer to the exhibit. When is this job scheduled to run?

A.Every minute at 12 seconds.
B.At 12:00 PM every day.
C.Every 12 hours.
D.Every Monday at 12:00 AM.
AnswerB

In Quartz cron syntax, 0 0 12 * * ? corresponds to seconds=0, minutes=0, hours=12, day-of-month=any, month=any, day-of-week=any. This results in the job triggering precisely at noon daily, which is a standard pattern for daily ingestion tasks requiring a single window of execution each day.

Why this answer

Understanding Quartz cron syntax is fundamental for scheduling jobs in Databricks. The expression '0 0 12 * * ?' maps to 12:00 PM (noon) daily. Being able to interpret and define schedules correctly is vital for data engineers to ensure that data pipelines run at the exact times required by business stakeholders, ensuring that data is fresh and available for reporting during peak business hours without unnecessary delays.

Exam trap

Candidates often misinterpret the Quartz cron syntax, specifically confusing the position of the hour and minute fields, or failing to account for the standard 12-hour clock format used in the expression.

18
MCQmedium

A data engineer is setting up a Lakeflow Job that runs a notebook task. The engineer needs the task to always execute even if the upstream task in the workflow fails. Which configuration should be applied to the dependent task's condition?

A.Set the task's "Run if" condition to "All done".
B.Set the task's "Run if" condition to "All succeeded".
C.Set the task's "Run if" condition to "At least one failed".
D.Set the task's "Run if" condition to "At least one succeeded".
AnswerA

"All done" causes the task to execute after all upstream dependencies have reached a terminal state, whether they succeeded or failed. This guarantees the task runs regardless of upstream outcomes, which directly fulfills the requirement to always execute even when a preceding task fails. It is the correct choice for cleanup or notification tasks that must run unconditionally.

Why this answer

In Lakeflow Jobs, task dependencies use "Run if" conditions to control execution based on upstream task states. "All done" ensures the task runs after all dependencies finish, regardless of success or failure. This is the only condition that satisfies the requirement to always execute, making it suitable for tasks that must run unconditionally, such as cleanup or alerting steps.

Exam trap

The trap here is assuming that "At least one succeeded" guarantees execution even when all upstream tasks fail, but it does not run if no upstream task succeeds.

19
Multi-Selectmedium

A data engineer is designing a Databricks Job workflow. Which TWO of the following are valid ways to trigger a Databricks Job?

Select 2 answers
A.Using the 'Run now' button in the UI.
B.Directly editing the Python source code file.
C.Using the Databricks Jobs REST API.
D.Changing the cluster node type.
E.Modifying the job owner permissions.
AnswersA, C

The 'Run now' feature is the primary way to manually execute a job immediately for testing or ad-hoc data processing requirements. It bypasses the schedule and triggers an instance of the job using the existing configuration, making it indispensable for troubleshooting or re-running failed tasks.

Why this answer

Databricks Jobs offer multiple entry points for execution. Understanding these triggers is essential for integrating pipelines into broader CI/CD cycles or event-driven architectures. By mastering these methods, engineers can effectively decouple their data processing logic from the scheduling mechanism, allowing for both manual verification during development and robust automated execution in production environments through APIs or service principals.

Exam trap

Candidates often select manual UI actions only, or mistakenly include unsupported methods, missing that REST APIs and manual triggers are both valid.

20
MCQmedium

A data engineer wants to ensure that a Databricks Job task only runs if the preceding task completes successfully, but needs to add a specific timeout threshold for this individual task. Where should this configuration be applied?

A.At the job level settings.
B.Within the individual task definition.
C.Inside the cluster configuration JSON.
D.Using a global Workspace policy.
AnswerB

Task definitions contain specific configurations like timeout_seconds, retries, and depends_on arrays. Setting the timeout here ensures that only this specific task is terminated if it exceeds the threshold, while allowing downstream tasks to potentially trigger or fail gracefully based on the defined job dependency graph.

Why this answer

In Databricks Jobs, task-level configurations such as timeouts, retries, and dependencies are defined within the task definition itself. By specifying a timeout in the task settings, the engineer prevents runaway processes from consuming cluster resources indefinitely. Understanding this granularity is vital for cost management and workflow reliability, as it allows engineers to set different SLAs for ingestion versus transformation tasks within a single orchestrated pipeline.

Exam trap

Candidates often look for cluster-level settings or global workspace configurations, incorrectly assuming that timeouts and retries are managed outside the specific task definition.

21
MCQhard

A data engineer configures a Lakeflow Job to run a notebook task on a job cluster. The notebook reads a parameter named run_date using the widget API. During a manual run, the engineer wants to supply a specific date without editing the notebook. The job also runs on a nightly schedule where the date should default to the current day. Which approach correctly supplies the parameter for both the manual and scheduled runs?

A.Hard-code the date inside the notebook and create a second notebook for manual runs.
B.Configure the notebook to read the date from a Unity Catalog table that the engineer updates before manual runs.
C.Use a cluster environment variable to store the date and change it before each manual run.
D.Define a job parameter with a default value and pass it to the notebook task, then override it when starting a manual run.
AnswerD

Job parameters can be defined with default values and referenced by tasks, and a manual run can override them at submission time. The notebook reads the parameter through the widget API, so the same notebook works for both scheduled and manual runs. This satisfies the requirement to supply a specific date manually while defaulting to the current day on schedule.

Why this answer

Job parameters with default values let a task receive a value automatically on schedule, while a manual run can override the same parameter at submission. The notebook consumes it via the widget API, so no code changes are needed between run modes. This is the standard Lakeflow Jobs pattern for dynamic per-run inputs and avoids hard-coding or external stores.

Exam trap

The trap here is thinking that a scheduled run cannot use a default while a manual run overrides the same parameter; job parameters support exactly that pattern.

22
MCQmedium

A data engineer has configured a Databricks Job with multiple dependent tasks forming a linear pipeline. Task A extracts data, Task B transforms it, and Task C loads it into a gold table. The pipeline runs daily. The team notices that if Task B fails due to an intermittent schema validation issue, the entire job run fails, but they want Task C to execute conditionally only if Task B succeeds, while alerting the on-call engineer immediately upon any failure. How should the task dependencies and conditional execution be configured?

A.Set Task C to run on failure of Task B to capture the exception state before alerting.
B.Convert Task B and Task C into a single notebook task and handle exceptions internally with try-except blocks.
C.Ensure Task C has a dependency pointing exclusively to Task B, and configure a job-level email notification for failures.
D.Disable the task dependency between Task B and Task C and run them in parallel with different cluster configurations.
AnswerC

Establishing a direct parent-child dependency between Task B and Task C guarantees that the load step never runs prematurely. Adding job-level failure alerts ensures immediate notification to the engineering team without compromising pipeline DAG semantics.

Why this answer

Configuring explicit task dependencies via the UI or API ensures that downstream tasks only execute upon the successful completion of their parents. By linking Task C strictly to Task B's success, you prevent corrupting downstream layers with partial or failed upstream transformations. This workflow pattern is essential for maintaining reliable data engineering pipelines in production Lakeflow environments.

Exam trap

Candidates incorrectly assume that failing upstream tasks automatically trigger downstream clean-up tasks, failing to explicitly configure success-only dependencies.

23
MCQmedium

A data engineer has a Lakeflow Job with a notebook task that occasionally fails due to transient network errors when reading from an external REST API. The engineer wants the task to automatically retry up to three times, but only for this specific task, without affecting other tasks in the job. What should the engineer do?

A.Set the task's Retries to 3 in the task configuration.
B.Configure the task's Retries setting to 3 and add a retry condition using the task's timeout.
C.Use a notebook %run magic command to wrap the API call in a loop that retries three times.
D.Set the job-level maximum concurrent runs to 3 and enable retries.
AnswerA

In Lakeflow Jobs, each task has a Retries field that specifies how many times the task should be retried if it fails. Setting it to 3 on the specific task ensures only that task retries, leaving other tasks unaffected. This is the direct and supported method for per-task retry behavior.

Why this answer

The correct approach is to set the Retries field on the specific task to 3. This is a built-in feature of Lakeflow Jobs that applies per task, so other tasks remain unaffected. It is the simplest and most direct way to handle transient failures without modifying code or job-level settings.

Exam trap

The trap here is confusing task-level retries with job-level concurrency or timeout settings, which do not control automatic retries.

24
Multi-Selecthard

A data engineer is building a Lakeflow Job that must run a sequence of tasks across different compute types. The ingest task must run on a job cluster with a specific Spark configuration, the transform task must run as a Delta Live Tables pipeline, and the report task must run on a separate SQL warehouse. Which TWO statements about task-level compute configuration in Lakeflow Jobs are correct? (Choose two.)

Select 2 answers
A.A SQL warehouse task can only run if the job uses an all-purpose cluster for all other tasks.
B.Task-level compute configuration is only available for notebook tasks, not for DLT pipeline or SQL tasks.
C.A Delta Live Tables pipeline task and a SQL warehouse task can be combined in the same Lakeflow Job with dependencies between them.
D.All tasks in a Lakeflow Job must share the same compute; task-level compute is not supported.
E.Each task in a Lakeflow Job can be configured with its own compute, including job clusters, DLT pipelines, and SQL warehouses.
AnswersC, E

Lakeflow Jobs allow dependencies between tasks regardless of their compute type. A DLT pipeline task can run, and upon success, a dependent SQL warehouse task can execute. This is useful when transformation results need to be queried or reported on a SQL warehouse. The job orchestrator manages the dependency graph and triggers each task on its designated compute.

Why this answer

Lakeflow Jobs are designed to orchestrate heterogeneous tasks, and each task can specify its own compute. This means a job can include a notebook task on a job cluster, a Delta Live Tables pipeline task, and a SQL warehouse task, with dependencies that control execution order. The two correct statements reflect that per-task compute is supported and that tasks with different compute types can be combined with dependencies.

Exam trap

The trap here is assuming all tasks must share compute or that only notebooks can have custom compute, when Lakeflow Jobs support per-task compute across task types.

25
MCQmedium

A data engineer is building a Lakeflow Job with a task that runs a SQL notebook. The task must run only on weekdays and must be completed before 9 AM. The engineer wants to configure the schedule to meet these requirements. What should the engineer do?

A.Configure the job to run continuously with a 1-hour pause between runs and enable a weekday filter.
B.Use a cron expression with a schedule that runs at 8 AM Monday through Friday.
C.Set the job to trigger on a file arrival event and filter for weekdays.
D.Use a Databricks SQL alert to trigger the job when a condition is met.
AnswerB

Lakeflow Jobs support cron-based schedules. A cron expression like '0 0 8 ? * MON-FRI *' (Quartz format) runs at 8 AM on weekdays. This meets the requirement of running only on weekdays and completing before 9 AM, assuming the task finishes within an hour. This is the standard way to set time-based schedules.

Why this answer

The correct method is to use a cron schedule that runs at 8 AM Monday through Friday. Lakeflow Jobs natively support cron expressions, allowing precise control over run times and days. This satisfies the weekday-only and before-9-AM requirements without additional complexity.

Exam trap

The trap here is assuming that event-based triggers or continuous execution can enforce time-of-day and weekday constraints, which they cannot.

26
MCQmedium

What is the primary function of the 'Retries' setting in a Databricks Job task?

A.To increase the task execution speed.
B.To automatically recover from transient failures.
C.To allow tasks to run in parallel.
D.To change the cluster node type.
AnswerB

Retries are explicitly designed to recover from transient, non-deterministic errors. By automatically re-running a task after a failure, the job platform increases the overall reliability of the pipeline, ensuring that temporary external outages do not cause a complete workflow failure that would otherwise require manual intervention.

Why this answer

The 'Retries' setting is a robust mechanism for handling transient failures, such as network timeouts or temporary cloud provider issues, without manual intervention. By configuring retries, engineers increase the resilience of their pipelines. This is a critical best practice in production environments where external dependencies may be unstable, as it minimizes the need for on-call support and ensures that jobs eventually succeed despite minor, non-permanent infrastructure hiccups.

Exam trap

Candidates often confuse the retries setting with data quality error handling or schema auto-repair, missing that it specifically addresses transient infrastructure failures.

27
MCQmedium

A data engineer needs to configure a Databricks Job containing three distinct tasks: ingest, transform, and report. The transform task must only execute if the ingest task completes successfully, but the report task should execute regardless of whether the transform task succeeds or fails. How should the task dependencies be configured?

A.Configure transform to depend on ingest, and configure report to depend on transform.
B.Configure ingest to depend on transform, and configure report to depend on ingest.
C.Configure transform to depend on ingest, and do not include transform as a dependency for report.
D.Configure both transform and report to depend on ingest simultaneously.
AnswerC

Setting transform to depend on ingest ensures the data is loaded before transformation begins. Omitting transform from report's dependencies allows the reporting workflow to run independently of the transformation task's success or failure state.

Why this answer

Task dependencies in Databricks Jobs are managed by defining upstream parent tasks using the depends_on field. To make transform run after ingest, transform depends on ingest. To make report run regardless of transform's outcome, report must not depend on transform, or should be configured independently.

Correctly structuring task dependencies ensures downstream consumers do not ingest corrupt data while maintaining resilience for reporting tasks.

Exam trap

Candidates often assume that a failure in an upstream task automatically blocks all subsequent tasks, not realizing they can define independent dependency paths for reporting or monitoring tasks.

28
MCQmedium

A data engineer manages a Lakeflow Job that runs a long-running notebook task on a job cluster. The task occasionally fails due to transient cloud storage errors, and the engineer wants the task to retry automatically without failing the entire job on the first attempt. The engineer also wants to be alerted only if all retries are exhausted. Which configuration should the engineer apply?

A.Set the task's retry count to a value greater than zero and configure a notification for job failure.
B.Enable continuous mode on the job and configure a notification for run start.
C.Set the cluster's autoscaling to a higher maximum and add a notification for cluster termination.
D.Configure a job-level timeout and enable email notifications on every task start.
AnswerA

Task-level retries allow the task to re-execute automatically when it fails, and the job only reports failure after retries are exhausted. A job failure notification then fires only when the task ultimately fails, matching the alerting requirement. This combination handles transient errors gracefully while avoiding noisy alerts on each retry attempt.

Why this answer

Task-level retries re-execute a failed task automatically, so transient storage errors can be absorbed without failing the job immediately. The job is only marked failed after retries are exhausted, at which point a job failure notification alerts the team. This matches the requirement to retry silently and alert only on final failure, unlike cluster or continuous-mode settings that do not target task retries.

Exam trap

The trap here is assuming that a job-level timeout or continuous mode provides retries; only task-level retry settings re-execute a failed task within the same run.

29
MCQmedium

When a Data Engineer uses a 'Repair and Rerun' functionality on a failed Databricks Job, what happens?

A.It deletes all previous successful task logs.
B.It restarts from the failed task.
C.It resets all task statuses to 'Pending'.
D.It requires a new cluster to be provisioned.
AnswerB

The Repair and Rerun functionality intelligently resumes the workflow starting from the point of failure. It uses the output of previously successful tasks, ensuring that only the failed task and its downstream dependencies are executed, which optimizes compute usage and speeds up the time to recovery for pipelines.

Why this answer

Repair and Rerun is a critical feature for managing complex DAG workflows. It allows engineers to restart a job from the specific failed task, inheriting the successful results of prior tasks. This saves significant time and compute costs by avoiding redundant processing of already completed data, which is essential for maintaining efficient pipelines and preventing unnecessary data re-computation in production environments.

Exam trap

Candidates often assume 'Repair and Rerun' restarts the entire job from the very beginning, failing to realize it is designed specifically to resume only from the point of failure.

30
MCQmedium

A data engineer maintains a Lakeflow Job with a scheduled trigger set to run every day at 08:00. The job's source table is refreshed by an upstream process that sometimes finishes later than expected, causing the job to process stale data. The engineer wants the job to start only after the upstream refresh completes, regardless of the clock time, while still preserving the existing 08:00 schedule as a fallback. Which trigger configuration should the engineer implement?

A.Convert the scheduled trigger to a continuous trigger so the job restarts immediately after each run.
B.Add a table update trigger on the upstream table while keeping the scheduled trigger, so the job runs when the table refreshes or at 08:00.
C.Add a file arrival trigger that monitors the upstream table's storage location and remove the scheduled trigger.
D.Add a run-duration timeout to the scheduled trigger so the job fails if the upstream has not finished by 08:00.
AnswerB

Lakeflow Jobs support table update triggers that start a run when a specified table receives a commit, which directly signals upstream completion. Keeping the scheduled trigger provides the fallback at 08:00 in case no update is detected. This combination satisfies both the event-driven and time-based requirements precisely without extra orchestration code.

Why this answer

The requirement is a job that starts when upstream data is ready, yet still runs at 08:00 if that event has not occurred. A table update trigger detects upstream commits and starts the job promptly, while the retained scheduled trigger supplies the fallback. Together they deliver event-driven orchestration with a time-based safety net, which neither a pure schedule nor a polling file trigger provides.

Exam trap

The trap here is assuming that a time-based schedule can be replaced by a file arrival trigger, when the requirement actually calls for a table update trigger combined with the existing schedule.

Ready to test yourself?

Try a timed practice session using only Working Lakeflow Jobs questions.