Courseiva

CCNA Understanding the Databricks Platform Questions

40 questions · Understanding the Databricks Platform · All types, answers revealed

1
MCQhard

Refer to the exhibit. An analyst is reviewing a cluster configuration JSON. If the cluster is currently idle, what happens after 40 minutes of inactivity?

A.The cluster scales down to 2 workers.
B.The cluster remains idle at 2 workers.
C.The cluster terminates.
D.The cluster scales up to 8 workers.
AnswerC

The configuration sets 'autotermination_minutes' to 30. This parameter dictates that the cluster will shut down after 30 minutes of inactivity. Since 40 minutes is greater than 30, the cluster will have successfully terminated, stopping the compute costs associated with the idle nodes in the environment.

Why this answer

Databricks clusters with auto-termination enabled will shut down after the specified inactivity threshold is met. In this configuration, the 'autotermination_minutes' is set to 30. Therefore, once the cluster remains idle for 30 minutes, it will automatically terminate.

By the 40-minute mark, the cluster has already completed the termination sequence, ensuring no additional costs are incurred for idle resources beyond the defined grace period.

Exam trap

Candidates often assume the cluster stays active for the full duration or wait for a manual shutdown, forgetting that auto-termination initiates the shutdown process immediately upon reaching the inactivity threshold.

2
MCQmedium

A data analyst joined a new Databricks workspace and sees a catalog named 'sales_prod' containing schemas 'bronze', 'silver', and 'gold'. The analyst needs to query only the 'gold' schema tables and must not see or query the 'bronze' or 'silver' tables. Which Unity Catalog object should the workspace administrator grant to satisfy this least-privilege requirement?

A.USE CATALOG on sales_prod and USE SCHEMA plus SELECT on the gold schema
B.USE SCHEMA on bronze and silver only
C.SELECT on the sales_prod catalog
D.ALL PRIVILEGES on the sales_prod catalog
AnswerA

Unity Catalog privileges are hierarchical: USE CATALOG lets the analyst traverse the catalog, and USE SCHEMA plus SELECT on gold grants read access to that schema's tables only. Without USE SCHEMA or SELECT on bronze and silver, the analyst cannot see or query those tables, achieving least privilege exactly as required.

Why this answer

Unity Catalog enforces a privilege hierarchy where traversing a catalog requires USE CATALOG and reading a schema's tables requires USE SCHEMA and SELECT on that schema. Granting USE CATALOG on sales_prod and USE SCHEMA plus SELECT on gold gives read access to only the gold tables, while the absence of grants on bronze and silver keeps them hidden.

Exam trap

The trap here is assuming a single broad grant at the catalog level can be scoped down, when Unity Catalog privileges only widen access and never restrict it.

3
Multi-Selectmedium

A data analyst is new to a Databricks workspace and needs to work with data stored in Unity Catalog. The analyst wants to understand which statements about Unity Catalog are accurate. (Choose two.)

Select 2 answers
A.Unity Catalog access control can be applied at the catalog, schema, table, and column levels, including row and column filters.
B.Unity Catalog provides a centralized governance model with a three-level namespace of catalog, schema, and table or view.
C.Unity Catalog metadata is stored separately in each workspace, so grants must be repeated for every workspace.
D.Unity Catalog can only be queried from notebooks and not from Databricks SQL warehouses or dashboards.
E.Unity Catalog requires that all tables be stored as external tables in cloud object storage rather than managed tables.
AnswersA, B

Unity Catalog supports fine-grained privileges at multiple levels, and it also supports row filters and column masks attached to tables. This lets administrators grant broad access to a catalog or schema while restricting sensitive rows or columns for specific users or groups. The unified privilege model means analysts see only the data they are entitled to, regardless of the query tool they use.

Why this answer

Unity Catalog centralizes governance with a three-level namespace and supports privileges from the catalog level down to columns, including row filters and column masks. These two characteristics let an analyst reference data with fully qualified names and trust that access rules are enforced consistently. The other statements misstate storage requirements, metadata placement, or the range of supported query interfaces.

Exam trap

The trap here is mixing up Unity Catalog with the older workspace-local Hive metastore, leading to false beliefs that grants are per-workspace or that only notebooks can query governed data.

4
Multi-Selectmedium

A Databricks workspace administrator wants to optimize costs and manage resources effectively. Which TWO of the following capabilities allow for automatic cluster termination and scaling?

Select 2 answers
A.Auto-termination
B.Auto-scaling
C.Cluster Tagging
D.Delta Cache
E.Instance Pools
AnswersA, B

Auto-termination is a critical cost-saving feature that monitors cluster inactivity. When no jobs are running for a specified period, the cluster automatically terminates, ensuring that users are not billed for compute resources that are idle. This is essential for managing budgets effectively in a shared cloud environment.

Why this answer

Databricks provides built-in mechanisms to manage compute costs through Auto-termination and Auto-scaling. Auto-termination automatically shuts down clusters after a period of inactivity, preventing wasteful billing. Auto-scaling allows the cluster to resize itself based on the current workload requirements, adding or removing workers dynamically.

These features are fundamental for maintaining a cost-efficient data platform while ensuring sufficient performance for varying data volumes during analysis.

Exam trap

Candidates often select 'Auto-scaling' and 'Auto-stop' as synonyms, or confuse them with 'Spot Instances'. They fail to recognize these as two distinct cost-saving mechanisms for cluster lifecycle management.

5
MCQhard

A data analyst is querying a Unity Catalog managed table named sales in the analytics catalog and the marketing schema. A query fails with an error indicating the table cannot be found, even though the analyst can see the table in Catalog Explorer. The analyst is connected to a SQL Warehouse in the same workspace. Which action is most likely to resolve the access issue?

A.Restart the SQL Warehouse so it picks up the latest Unity Catalog metadata
B.Verify that the analyst has USE CATALOG on analytics, USE SCHEMA on marketing, and SELECT on the sales table
C.Ask a workspace administrator to enable table access control on the all-purpose cluster
D.Convert the managed table to an external table so the analyst can read the underlying files
AnswerB

In Unity Catalog, accessing a table requires a chain of privileges: USE CATALOG on the catalog, USE SCHEMA on the schema, and SELECT on the table. Seeing the table in Catalog Explorer does not guarantee these privileges are granted to the user. If any link in the chain is missing, the query fails even though the object is visible, making this the most likely resolution.

Why this answer

Unity Catalog enforces a privilege hierarchy: USE CATALOG on the catalog, USE SCHEMA on the schema, and SELECT on the table. Visibility in Catalog Explorer does not imply these grants. Restarting the warehouse, converting the table to external, or enabling legacy cluster table access control do not address missing Unity Catalog privileges.

Exam trap

The trap here is equating visibility in Catalog Explorer with query access, when Unity Catalog still requires the full USE CATALOG, USE SCHEMA, and SELECT chain.

6
MCQmedium

Which Databricks feature provides a detailed, lineage-based view of how data flows from source to destination?

A.Databricks SQL Audit Logs
B.Unity Catalog Data Lineage
C.Delta Live Tables Pipeline Monitoring
D.Cluster Event Logs
AnswerB

Unity Catalog automatically captures lineage information for queries, jobs, and tables. It provides a visual and programmatic way to explore how data is derived, identifying the source tables, intermediate transformations, and destination tables. This is critical for data governance, impact analysis, and maintaining high trust in the platform's analytical outputs.

Why this answer

Data lineage in Unity Catalog tracks the relationships between data assets, transformations, and end-users. This visibility is essential for impact analysis, debugging, and compliance. Understanding lineage helps analysts understand the origin of their data and the downstream effects of any changes they make, ensuring that business stakeholders always have confidence in the data's quality, provenance, and the transformation logic applied throughout the analytical pipeline.

Exam trap

Candidates confuse table history or audit logs with data lineage, overlooking the specific visual mapping capabilities provided by Unity Catalog.

7
MCQhard

A data analyst has a Databricks SQL dashboard that reads a table in the 'finance' catalog. The dashboard works for the analyst but shows an error for a colleague who has SELECT on the table. The colleague lacks USE CATALOG on 'finance' and USE SCHEMA on the containing schema. What is the most likely cause of the error the colleague sees?

A.The dashboard was created with the analyst's credentials embedded, so it always runs as the analyst
B.The colleague is connected to a different SQL warehouse that cannot reach the finance catalog
C.The table is a view that references tables the colleague cannot access
D.The colleague cannot traverse the catalog and schema, so Unity Catalog denies resolution of the table even though SELECT is granted
AnswerD

Unity Catalog requires USE CATALOG on the parent catalog and USE SCHEMA on the parent schema in addition to SELECT on the table. Without those traversal privileges, the colleague cannot resolve the table's fully qualified name, so the dashboard query fails despite the SELECT grant, which explains the error exactly.

Why this answer

Unity Catalog resolves objects through their full three-part name, which requires USE CATALOG on the catalog and USE SCHEMA on the schema before SELECT on the table is even evaluated. The colleague holds SELECT but lacks the traversal privileges, so name resolution fails and the dashboard query errors until those grants are added.

Exam trap

The trap here is assuming SELECT on a table is sufficient, when Unity Catalog also requires traversal privileges on the enclosing catalog and schema.

8
MCQmedium

A data analyst is building a dashboard in Databricks SQL and needs to give viewers the ability to change the date range and a region filter without editing the underlying query. The dashboard should refresh results based on the viewer's selections. Which feature should the analyst use?

A.A materialized view that pre-aggregates the data by date and region
B.Dashboard parameters that are referenced in the dataset queries
C.A Delta Live Tables pipeline that refreshes the dashboard dataset hourly
D.A notebook widget embedded in the dashboard as a separate visualization
AnswerB

Dashboard parameters let viewers supply values, such as a date range or region, that are substituted into the dataset queries at run time. This is exactly the interactive filtering behavior required, and it does not require viewers to edit SQL. Parameters can be exposed as widgets on the dashboard, so stakeholders can change the date range and region and see updated results.

Why this answer

Dashboard parameters are the built-in way to let viewers supply values such as a date range or region that are injected into dataset queries. They provide interactive controls without editing SQL and refresh results on selection. Materialized views, Delta Live Tables, and notebook widgets address performance, pipeline refresh, or notebook interactivity, not dashboard viewer filtering.

Exam trap

The trap here is confusing a performance feature like a materialized view with an interactive filtering feature like dashboard parameters.

9
Multi-Selectmedium

A data analyst is working in a Databricks workspace and needs to ensure that their notebook code is version-controlled and collaborative. Which TWO actions should they take?

Select 2 answers
A.Use the built-in Databricks Repos feature to connect to a Git provider.
B.Copy and paste code into a local text file.
C.Enable 'Collaboration' settings on the notebook to allow multiple users to edit simultaneously.
D.Create a new notebook file for every code change.
E.Download the notebook to their local machine every hour.
AnswersA, C

Databricks Repos provides seamless integration with Git providers like GitHub, GitLab, and Azure DevOps. This allows users to perform standard Git operations like pull, push, and commit directly from the notebook interface. This native integration is the recommended path for managing source code versions within the Databricks ecosystem for data projects.

Why this answer

Integrating version control and collaboration is essential for professional data engineering workflows. By linking notebooks to Git providers, analysts ensure that changes are tracked, auditable, and easily revertible. This standard practice prevents accidental data loss and fosters teamwork, allowing multiple contributors to manage code repositories effectively within the Databricks environment while maintaining a single source of truth for all analytical project documentation and complex transformation logic.

Exam trap

Candidates frequently select manual file exports or third-party version control plugins, overlooking Databricks Repos and native notebook collaboration settings as the built-in enterprise solutions.

10
MCQmedium

A data analyst is troubleshooting a performance issue in a notebook. The query runs slowly when processing a large table. Which approach should the analyst take to improve performance?

A.Increase the number of users connected to the workspace.
B.Run the 'OPTIMIZE' command on the Delta table.
C.Delete the table and re-create it without Delta Lake.
D.Force the notebook to run on a single-node cluster.
AnswerB

The OPTIMIZE command compacts small files into larger, more efficient files, significantly improving read performance. It is a standard procedure for data analysts to maintain high performance in Delta Lake tables. This simple action can drastically reduce I/O overhead for analytical queries scanning large volumes of data.

Why this answer

Performance tuning in Databricks often involves examining how data is stored and accessed. Identifying bottlenecks like small files or lack of partitioning is key. By using built-in optimization commands, analysts can improve query speed significantly.

This process is a core skill for any Databricks analyst, as it directly impacts the cost of compute resources and the user experience when working with massive, real-world datasets.

Exam trap

Candidates often confuse the 'OPTIMIZE' command for compaction with 'VACUUM' for old file removal or cluster-level scaling options, missing that small files directly degrade query performance on Delta tables.

11
MCQeasy

Which component of the Databricks Data Intelligence Platform allows users to discover, govern, and share data across the entire organization?

A.Databricks SQL
B.Unity Catalog
C.Delta Live Tables
D.Compute Clusters
AnswerB

Unity Catalog is the primary governance and discovery component in Databricks. It enables administrators to manage access control lists, perform auditing, and document data assets across multiple workspaces. It serves as the single source of truth for metadata, facilitating secure collaboration and compliance across the entire organizational data landscape.

Why this answer

Unity Catalog is the centralized governance solution for data, analytics, and AI on the Databricks platform. It provides a single interface to manage permissions, track data lineage, and discover assets. Mastery of Unity Catalog is essential for any data analyst as it ensures that data access is secure, compliant, and transparent, effectively bridging the gap between raw storage and meaningful, governed business insights for all users.

Exam trap

Candidates often mistake workspace folders or cloud storage buckets for platform-wide governance tools, ignoring Unity Catalog's role as the central cataloging mechanism.

12
MCQeasy

What is the primary function of the 'Databricks Assistant' within the notebook environment?

A.To automatically execute all notebooks in a production pipeline.
B.To assist in writing, debugging, and explaining notebook code.
C.To manage user permissions across the entire workspace.
D.To store and back up notebook versions automatically.
AnswerB

The Databricks Assistant is an AI-integrated tool that provides suggestions, debugging help, and code explanations directly within the notebook. It enhances developer productivity by accelerating the coding process and providing immediate support for syntax-related issues, making it a powerful tool for analysts at all experience levels to write efficient code.

Why this answer

The Databricks Assistant is an AI-powered coding companion designed to help users write, debug, and optimize code faster. By leveraging context-aware suggestions, it helps analysts become more efficient. Understanding how to use the Assistant is beneficial for streamlining development workflows, as it assists with complex syntax, troubleshooting common errors, and generating boilerplate code, allowing the analyst to focus on the core logic of their data analysis tasks.

Exam trap

Candidates often mistakenly believe the Assistant is for data visualization or automated dashboard generation, failing to recognize its core role as a coding companion for writing, debugging, and explaining code.

13
Multi-Selectmedium

Which TWO of the following are true regarding Databricks SQL Warehouses?

Select 2 answers
A.They must be manually started every time a user runs a query.
B.They are specifically optimized for SQL-based analytical queries.
C.They support the same multi-language notebook features as all-purpose clusters.
D.They can automatically scale to handle concurrent query demand.
E.They are only available for administrative users.
AnswersB, D

SQL Warehouses are built specifically to provide high-performance SQL execution, serving as the compute engine for Databricks SQL. They include optimizations such as query result caching and specialized query planning, making them vastly superior to general-purpose clusters for structured SQL queries and interactive BI reporting tasks.

Why this answer

SQL Warehouses are specialized compute resources optimized for SQL workloads. Understanding the difference between these and general-purpose clusters is crucial for analysts, as warehouses offer features like serverless startup and auto-stop, which significantly optimize cost and performance. These warehouses enable the platform's BI and dashboarding capabilities, making them a cornerstone of the data analyst's toolkit for delivering consistent, reliable reports to business stakeholders.

Exam trap

Candidates often mistakenly believe SQL Warehouses are general-purpose compute resources, failing to understand they are specialized for SQL analytics and lack the flexibility of general-purpose clusters for Python or Scala code.

14
MCQmedium

A data analyst needs to perform ad-hoc SQL queries on a large dataset while ensuring the compute resources automatically terminate when idle to minimize costs. Which compute resource is most appropriate for this task?

A.Job Compute
B.SQL Warehouse
C.All-Purpose Compute
D.Delta Live Tables Pipeline
AnswerB

SQL Warehouses offer managed, auto-scaling compute resources tailored for SQL analytics. They provide serverless or classic scaling options with built-in auto-stop features that terminate idle resources, ensuring cost-efficiency. This aligns perfectly with the analyst's requirement to perform ad-hoc SQL queries while maintaining strict control over compute costs during periods of inactivity.

Why this answer

SQL Warehouses are designed specifically for SQL-based workloads, supporting BI tools and ad-hoc analysis through the SQL editor. They feature auto-stop functionality, which shuts down the cluster after a specified period of inactivity, directly addressing the cost-optimization requirement. Understanding the distinction between SQL Warehouses and All-Purpose clusters is vital for analysts, as Warehouses provide optimized performance for SQL queries and efficient resource management for collaborative reporting environments.

Exam trap

Candidates frequently confuse All-Purpose clusters with SQL Warehouses, failing to recognize that SQL Warehouses are specifically optimized for ad-hoc SQL queries and auto-stop cost management.

15
MCQhard

Refer to the exhibit. A data analyst attempts to run a query in the SQL editor but receives the error displayed. Which action should the administrator take to resolve this?

A.Restart the SQL Warehouse to refresh the user's session token.
B.Grant the user the 'CAN_USE' permission on the specific SQL Warehouse.
C.Upgrade the user's workspace role to 'Workspace Admin'.
D.Increase the maximum number of clusters in the SQL Warehouse configuration.
AnswerB

The 'CAN_USE' permission is explicitly required for users to execute queries on a SQL Warehouse. By granting this permission in the workspace settings, the administrator enables the user's identity to authenticate with the compute resource, resolving the 403 Forbidden error and allowing the analyst to proceed with their work.

Why this answer

The error indicates a lack of the necessary access control privilege for the specific SQL Warehouse. In Databricks, permissions are managed via Access Control Lists (ACLs). The analyst requires the 'CAN_USE' permission to attach to and execute queries on a warehouse.

Administrators must update the warehouse's permission settings in the Databricks workspace to grant this access level to the user or their associated group.

Exam trap

Candidates often check table-level permissions or cluster instance types instead of inspecting the specific 'CAN_USE' access control list for the SQL Warehouse.

16
MCQhard

An analyst needs to combine data from a Delta table in the 'dev' catalog with a CSV file uploaded to a volume. Which feature allows the analyst to manage both in a single query?

A.Unity Catalog Volumes
B.External locations
C.Databricks File System (DBFS) mounting
D.Manual file upload via a web browser
AnswerA

Volumes allow for the management of unstructured data (files) within the Unity Catalog's governed space. By using volumes, an analyst can access these files via SQL and join them with Delta tables, providing a unified and secure approach to handling both structured and semi-structured data within Databricks.

Why this answer

Unity Catalog Volumes enable analysts to manage non-tabular data, such as CSV or JSON files, alongside traditional Delta tables. This unification simplifies the data integration process. By treating files in volumes as first-class objects within the catalog, analysts can use SQL to query and join disparate data sources, reducing the complexity of data pipelines and making it easier to perform multi-source analyses within a secure, governed environment.

Exam trap

Candidates often choose traditional DBFS or external cloud storage paths instead of Unity Catalog Volumes, forgetting that volumes are specifically designed to govern non-tabular files in a single unified namespace.

17
MCQmedium

A data analyst needs to run a scheduled SQL query every morning and deliver the result to a finance team as a CSV file in cloud storage. The query logic is already tested in a Databricks SQL query. Which Databricks capability should the analyst use to automate this delivery?

A.A Databricks SQL query schedule with a destination
B.A Databricks SQL alert
C.A Delta Live Tables pipeline
D.A notebook-scoped widget parameter
AnswerA

Scheduling a saved SQL query lets it run at defined times, and configuring a destination delivers results to cloud storage, email, or a dashboard. This directly matches the need to run daily and deliver CSV output to finance. It is the native Databricks SQL automation mechanism for recurring query results.

Why this answer

Databricks SQL supports scheduling saved queries so they execute on a recurring cadence, and a schedule can include a destination such as cloud storage or email. That combination automates both execution and delivery of CSV results, exactly matching the finance team's requirement. Alerts, widgets, and Delta Live Tables address monitoring, interactivity, and pipeline ETL rather than scheduled result delivery.

Exam trap

The trap here is confusing monitoring features such as alerts with delivery features, when only a scheduled query with a destination both runs on a cadence and sends the file.

18
MCQeasy

An analyst opens the Databricks SQL editor and wants to browse the tables available in the workspace's Unity Catalog metastore before writing a query. Which workspace object should the analyst use to navigate catalogs, schemas, and tables?

A.The Cluster configuration page
B.The Repos folder in the workspace sidebar
C.The Jobs run history UI
D.The Data Explorer in the Databricks workspace
AnswerD

The Data Explorer is the workspace UI for browsing Unity Catalog objects such as catalogs, schemas, tables, and volumes. It lets an analyst inspect metadata, permissions, and sample data without writing SQL first. In this scenario, it is the intended navigation surface for discovering available tables before composing a query in the SQL editor.

Why this answer

The Data Explorer is the built-in workspace UI for browsing Unity Catalog objects, letting analysts navigate catalogs and schemas, inspect table metadata, and preview data. Before writing SQL, an analyst can confirm table names and columns there. The other listed surfaces handle Git integration, compute configuration, or job monitoring, none of which expose the metastore hierarchy needed for table discovery.

Exam trap

The trap here is assuming any workspace sidebar item that lists files or objects can browse Unity Catalog tables, when only the Data Explorer presents the catalog-schema-table hierarchy.

19
MCQeasy

A data analyst at a retail company needs a serverless, fully managed SQL environment in Databricks to run BI queries against a gold Delta table. The team has no interest in managing clusters, and the queries must start instantly without a warm-up period. Which Databricks SQL warehouse type should the analyst select?

A.Serverless SQL warehouse
B.All-purpose cluster with the SQL dialect enabled
C.Pro SQL warehouse with Photon enabled
D.Classic SQL warehouse with a fixed cluster size
AnswerA

Serverless SQL warehouses run in the Databricks account's compute plane, so Databricks manages the infrastructure and the warehouse starts in seconds with no cluster configuration. This matches the requirement for a fully managed, instant-start environment for BI queries on a gold Delta table, removing all capacity planning from the analyst's team.

Why this answer

The requirement is a fully managed SQL endpoint that starts immediately and requires no cluster sizing, which is exactly what a serverless SQL warehouse provides. It runs on Databricks-managed compute, scales automatically for BI workloads, and eliminates warm-up and capacity planning. Classic and Pro warehouses and all-purpose clusters all involve customer-managed compute.

Exam trap

The trap here is assuming any SQL warehouse is serverless, when classic and Pro warehouses still run on compute provisioned inside the customer's own cloud account.

20
MCQmedium

An analyst wants to include a chart from a Databricks SQL query in a presentation and also allow stakeholders to explore the same query interactively later. Which two-part approach best fits this need?

A.Create a separate notebook for each stakeholder and email the notebook files
B.Download the chart as an image for the presentation, and share the saved query or dashboard for interactive exploration
C.Take a screenshot of the SQL editor and paste it into the slides
D.Copy the query text into the presentation slides and ask stakeholders to run it themselves
AnswerB

Exporting the visualization as an image satisfies the static presentation need, while sharing the saved query or dashboard gives stakeholders a live, interactive view. This combination covers both requirements without duplicating logic. It is the practical Databricks SQL workflow for mixing static and interactive sharing.

Why this answer

The scenario needs both a static visual for slides and a live, explorable view for stakeholders. Exporting the chart as an image handles the presentation, and sharing the saved query or dashboard provides governed interactivity. Copying code, distributing notebooks, or screenshotting the editor fails to produce a proper chart and does not support shared exploration.

Exam trap

The trap here is assuming one artifact can serve both a static presentation and live exploration, when the practical split is an exported image plus a shared query or dashboard.

21
MCQmedium

A data analyst needs to grant a colleague the ability to run an existing Databricks SQL query and view its results, but the colleague must not be able to edit the query text or change its schedule. The query is saved in the workspace. Which permission level should the analyst assign on the query object?

A.CAN EDIT
B.CAN VIEW
C.CAN RUN
D.CAN MANAGE
AnswerC

CAN RUN grants the ability to execute the saved query and see its results without allowing edits to the query text or its schedule. This precisely matches the requirement: the colleague can run the query and view output, but cannot modify it. It is the least-privilege permission that still enables the needed action.

Why this answer

CAN RUN is the least-privilege permission that lets a user execute a saved query and view results while preventing edits to the query text or schedule. CAN VIEW does not grant run capability, while CAN EDIT and CAN MANAGE both allow modifications that the scenario explicitly forbids.

Exam trap

The trap here is choosing CAN VIEW for read-only access when the colleague also needs to execute the query, which requires CAN RUN.

22
Multi-Selectmedium

A data analyst is new to a Databricks workspace and needs to understand which compute options are available for running SQL and notebooks. Which TWO of the following statements accurately describe Databricks compute in this context? (Choose two.)

Select 2 answers
A.SQL warehouses are optimized for running SQL and serving dashboards and scheduled queries.
B.All-purpose clusters are required to run any query in the Databricks SQL editor.
C.SQL warehouses and all-purpose clusters share the same configuration settings and scaling behavior.
D.A job cluster is created for a specific job run and terminated when the run completes.
E.Serverless compute for notebooks cannot be used with Unity Catalog.
AnswersA, D

SQL warehouses are compute resources specifically tuned for SQL workloads, including BI dashboards, Databricks SQL queries, and scheduled reports. They separate storage from compute and can auto-stop when idle. For an analyst focused on SQL and visualization, a SQL warehouse is the appropriate compute choice, making this statement accurate.

Why this answer

SQL warehouses are purpose-built for SQL, dashboards, and scheduled queries, while job clusters are ephemeral compute created for a specific run and terminated afterward. Together these describe two real compute behaviors an analyst should know. The other statements misstate requirements or capabilities, such as claiming all-purpose clusters are needed for the SQL editor or that serverless cannot use Unity Catalog.

Exam trap

The trap here is assuming there is a single compute type for all Databricks work, when SQL warehouses, job clusters, and all-purpose clusters serve distinct purposes.

23
Multi-Selectmedium

A data analyst is preparing to publish a Databricks SQL dashboard for a team of business users. The analyst wants the dashboard to load quickly and to remain usable as the underlying Delta table grows. Which two practices should the analyst follow? (Choose two.)

Select 2 answers
A.Create the dashboard datasets against a SQL Warehouse that is sized appropriately and has auto-stop configured
B.Store the dashboard datasets as CSV files in a Unity Catalog volume for faster reads
C.Pre-aggregate large fact tables into summary tables or materialized views used by the dashboard datasets
D.Disable the SQL Warehouse auto-stop so the dashboard always has warm compute
E.Grant every viewer CAN MANAGE on the dashboard so they can tune queries themselves
AnswersA, C

Sizing the SQL Warehouse appropriately ensures the dashboard queries have enough compute to return results quickly, and auto-stop controls cost when the dashboard is idle. This directly supports fast loading and sustainable operation as usage grows. The warehouse is the compute that serves dashboard queries, so choosing and configuring it correctly is a core best practice.

Why this answer

Fast, scalable dashboards rely on appropriately sized SQL Warehouse compute with auto-stop and on reducing the data scanned through pre-aggregation such as summary tables or materialized views. Storing datasets as CSV, disabling auto-stop, and granting broad manage permissions do not improve performance and introduce cost or governance problems.

Exam trap

The trap here is thinking that keeping compute always on or moving data to CSV improves dashboard performance, when sizing, auto-stop, and pre-aggregation are the real levers.

24
MCQeasy

Which component of the Databricks platform acts as the central entry point for users to manage their data, notebooks, and experiments?

A.The Control Plane
B.The Databricks Workspace
C.The Data Plane
D.The Unity Catalog
AnswerB

The Databricks Workspace is the user-facing environment for creating and managing data assets. It serves as the primary hub where analysts develop notebooks, manage permissions, and track experiments, offering a cohesive experience for team-based data analysis and collaborative data science projects within an organization.

Why this answer

The Workspace is the primary user interface in Databricks where data analysts interact with the platform. It provides a unified environment for writing code in notebooks, organizing folders, managing libraries, and viewing experiments. Understanding the Workspace is essential for navigating the platform, as it organizes all collaborative assets and provides the tools necessary for data exploration, cleaning, and visualization across the entire data lifecycle.

Exam trap

Candidates frequently confuse the 'Workspace' with the 'SQL Warehouse' or 'Compute' clusters. They mistake the compute engine for the interface where users actually organize and manage their data assets.

25
Multi-Selectmedium

Which THREE features are provided by Unity Catalog to enhance data governance?

Select 3 answers
A.Centralized access control using standard SQL.
B.Automated data lineage capture.
C.Direct management of physical cloud infrastructure servers.
D.A unified namespace for data assets.
E.Automatic generation of Power BI dashboards.
AnswersA, B, D

Unity Catalog enables administrators to manage permissions using standard SQL statements like GRANT and REVOKE. This unified approach makes security management consistent across different workspaces, reducing the complexity and manual overhead associated with managing user rights in a distributed data environment, while ensuring granular control over sensitive data assets.

Why this answer

Unity Catalog is the centralized governance layer for Databricks. It provides three critical capabilities: a unified namespace for data across workspaces, centralized access control using standard SQL GRANT/REVOKE syntax, and automated data lineage tracking. These features ensure that data is secure, discoverable, and auditable across an entire organization.

For analysts, this simplifies the process of finding data and ensures that security policies are consistently applied, regardless of the workspace or tool being used.

Exam trap

Candidates often select features like 'data ingestion' or 'query optimization,' which are not core governance features of Unity Catalog, leading to incorrect selections in multi-choice questions.

26
MCQhard

Refer to the exhibit. Given the provided JSON configuration for a Databricks cluster, what is the primary use case for this resource?

A.Running automated production ETL jobs.
B.Interactive data analysis and notebook development.
C.Long-running streaming data ingestion.
D.Batch processing of large ML models.
AnswerB

The all-purpose cluster type is designed for interactive development in notebooks. It allows users to start, stop, and restart clusters to run ad-hoc queries and perform data exploration. This configuration is standard for analytical tasks where developers need a responsive environment to test code and visualize findings in real-time.

Why this answer

Refer to the exhibit. The configuration shows an all-purpose cluster with autoscaling enabled. All-purpose clusters are primarily used for interactive development and data exploration within notebooks.

Because they consume more DBU resources compared to job clusters, setting an idle termination limit is a cost-optimization best practice. Understanding cluster types is fundamental for managing platform costs and ensuring that resources are allocated appropriately based on the specific requirements of the workload being executed.

Exam trap

Candidates frequently confuse all-purpose clusters with job clusters, assuming the configuration is for production pipelines when the presence of autoscaling and idle termination clearly points to interactive, exploratory development use cases.

27
MCQmedium

Refer to the exhibit. What is the most appropriate action to resolve this access issue?

A.Change the cluster configuration to use a different Spark version.
B.Request the workspace administrator to grant appropriate privileges in Unity Catalog.
C.Move the data to a local file in DBFS.
D.Re-create the table in a different schema.
AnswerB

Unity Catalog uses a centralized grant-based permission model. The administrator must explicitly grant the 'USE CATALOG' privilege to the user for the 'sales_data' catalog. This is the correct procedure for resolving access errors, ensuring that security remains intact while allowing the analyst to perform their required data operations.

Why this answer

Refer to the exhibit. This error indicates a failure in the Unity Catalog permission model. When a user lacks the necessary 'USE' or 'SELECT' grants on a catalog, they are blocked from accessing any data within it.

Understanding how Unity Catalog manages access at the catalog, schema, and table levels is vital for analysts to troubleshoot access issues and comply with organizational security policies.

Exam trap

Candidates often suggest rewriting the SQL query or recreating the table, missing that the underlying root cause is a missing Unity Catalog privilege grant.

28
MCQmedium

A data analyst is working in a notebook and notices that the query results are inconsistent compared to an earlier run, despite no code changes. What is the most likely cause?

A.The notebook is not using a SQL Warehouse.
B.The underlying data files are being modified by a concurrent process.
C.The cluster has automatically terminated and restarted.
D.The notebook needs more memory to process the dataset.
AnswerB

If the data is not stored as a Delta table, or if the process bypasses transaction logs, concurrent writes will cause inconsistent reads. Delta Lake solves this by providing snapshot isolation, which guarantees that once a query starts, it reads a consistent version of the data, regardless of concurrent modifications.

Why this answer

Inconsistent results often stem from the underlying data being updated concurrently without proper versioning or isolation. Databricks provides ACID transactions via Delta Lake. If the analyst is querying raw files instead of Delta tables, or if the table is being updated by another process, the results might vary.

Understanding how Delta Lake handles snapshot isolation is crucial for ensuring that analytical reports remain consistent and reproducible over time.

Exam trap

Candidates often blame the notebook environment or caching, ignoring that concurrent writes to the underlying data source are the most frequent cause of inconsistent results.

29
MCQmedium

A data analyst needs to share a notebook with a colleague who should be able to run the code but not modify it. Which permission level should the analyst grant to the colleague?

A.Can Manage
B.Can Run
C.Can Edit
D.Can View
AnswerB

The 'Can Run' permission allows a user to attach the notebook to a cluster and execute its cells. It specifically prevents the user from modifying the notebook's code or changing its settings, making it the appropriate choice for analysts who need to run reports without altering the source data processing logic.

Why this answer

The 'Can Run' permission is designed for scenarios where users need to execute code within a notebook without altering the underlying logic. In the Databricks access control model, this permission ensures that the notebook's integrity remains intact while still allowing collaborative execution. It is a critical security practice to follow the principle of least privilege, ensuring users have only the necessary permissions for their specific analytical tasks.

Exam trap

Candidates often select 'Can Edit' or 'Can Manage' out of habit, ignoring the principle of least privilege. They fail to realize 'Can Run' is sufficient for executing code without altering logic.

30
Multi-Selecthard

Which THREE of the following are benefits of using Delta Lake over standard Parquet files in Databricks?

Select 3 answers
A.ACID transaction support
B.Automatic data compression to non-standard formats
C.Time travel capabilities
D.Schema enforcement and evolution
E.The ability to run queries without a compute engine
AnswersA, C, D

Delta Lake provides Atomicity, Consistency, Isolation, and Durability (ACID) guarantees. This ensures that concurrent reads and writes are handled safely, preventing data corruption and partial writes. This is a fundamental requirement for reliable data warehousing on top of cloud object storage, ensuring users always see consistent data states.

Why this answer

Delta Lake introduces ACID transactions, time travel, and schema enforcement to the data lake, transforming it into a lakehouse. These features are critical for maintaining data integrity in complex production environments. Understanding these benefits allows analysts to make informed decisions about storage formats, ensuring that the data platform remains reliable, scalable, and capable of handling complex analytical requirements without risking data corruption or inconsistency.

Exam trap

Candidates often include features like 'data compression' or 'partitioning' as benefits of Delta Lake, failing to select the core architectural advantages that differentiate it from standard Parquet files.

31
Multi-Selectmedium

Which TWO of the following statements accurately describe the relationship between Databricks SQL Warehouses and Delta Lake?

Select 2 answers
A.SQL Warehouses must always be co-located in the same cloud region as the Delta Lake storage.
B.SQL Warehouses utilize Delta Lake's transaction log to ensure consistent query results.
C.Delta Lake is a compute engine that replaces the need for SQL Warehouses.
D.SQL Warehouses can query Delta tables using optimized Parquet files.
E.SQL Warehouses require all data to be imported into a proprietary internal database format.
AnswersB, D

SQL Warehouses leverage Delta Lake's transaction log to provide snapshot isolation. This ensures that when an analyst queries a table, they see a consistent version of the data, even if other processes are concurrently writing to the same table, thereby maintaining data integrity during complex ad-hoc analytical operations.

Why this answer

Databricks SQL Warehouses provide the compute layer to execute ANSI-compliant SQL against data stored in Delta Lake format. Delta Lake serves as the underlying storage layer, providing ACID transactions and schema enforcement. Analysts must understand this separation, as it allows compute resources to scale independently from storage.

This architecture ensures that SQL Warehouses can efficiently query massive volumes of data while leveraging the reliability and performance optimizations inherent in the Delta Lake storage structure.

Exam trap

Candidates often assume SQL Warehouses store their own data, failing to realize they are strictly compute engines that query data residing in external Delta Lake storage via the transaction log.

32
MCQeasy

A data analyst at a retail company has been asked to build interactive sales reports that refresh automatically each morning and can be shared with regional managers. The analyst wants to use Databricks SQL to create queries, schedule them, and publish visualizations without writing notebook code. Which Databricks component should the analyst use to accomplish this?

A.Databricks Jobs
B.Databricks Marketplace
C.Databricks Repos
D.Databricks SQL
AnswerD

Databricks SQL is the built-in service for SQL-centric analytics. It provides the SQL editor, query scheduling, dashboards, and alerting, and it runs queries on SQL warehouses. It lets an analyst build and share reports without notebook code, and the scheduled refresh delivers the required daily updates to regional managers, so this component directly satisfies every stated requirement.

Why this answer

Databricks SQL is purpose-built for SQL analysts: it includes a query editor, visualization tools, dashboards, alerts, and scheduled refreshes running on SQL warehouses. Because the analyst needs to author queries, schedule them, and publish shareable visualizations without notebook code, this is the only listed component that covers the full workflow end to end.

Exam trap

The trap here is assuming that any scheduling-capable component, such as Jobs, can substitute for the SQL analytics experience, when only Databricks SQL provides the editor, dashboards, and scheduled query refresh together.

33
MCQmedium

A data analyst needs to query a table that contains sensitive customer financial data. Company policy requires that analysts see only masked values for account numbers in query results, and the masking must apply regardless of which tool or user queries the table. Which Databricks capability should be used to enforce this consistently at the data layer?

A.A view that filters rows before returning results
B.Column masks applied to the table
C.Workspace-level notebook permissions
D.Cluster access mode configuration
AnswerB

Column masks are table-level security features that evaluate a masking expression each time the column is read, returning masked values for unauthorized users. Because the mask is attached to the table itself, it applies no matter which client, notebook, or dashboard issues the query, satisfying the requirement that masking be enforced consistently at the data layer for all users.

Why this answer

Column masks attach a masking expression directly to a table column, so the transformation is evaluated at query time for every read, independent of the client or user. This data-layer enforcement is exactly what the policy demands: unauthorized analysts receive masked account numbers whether they query via a notebook, the SQL editor, or a BI tool, and authorized users can still see real values.

Exam trap

The trap here is confusing access control with value masking: permissions and views restrict who or what rows are returned, but only a column mask rewrites the actual values for unauthorized readers on every query path.

34
MCQeasy

Which Databricks component is the primary interface for collaborative, interactive data analysis and visualization?

A.Databricks SQL Warehouse
B.Databricks Notebooks
C.Delta Live Tables
D.Unity Catalog
AnswerB

Notebooks are the primary interface for collaborative data science and analysis. They allow multiple users to edit code simultaneously, document findings with Markdown, and generate charts directly from query results. This collaborative capability is essential for data analysts who need to build and share reproducible analytical pipelines and reports.

Why this answer

Databricks Notebooks are the core collaborative environment where users write code, execute queries, and create visualizations. They support polyglot programming (SQL, Python, R, Scala) and are integrated directly with the platform's compute resources. For analysts, notebooks provide a persistent workspace to document workflows, share insights with stakeholders, and execute complex data transformations, making them the most important tool for iterative data analysis and team collaboration within the Databricks environment.

Exam trap

Candidates mistakenly select SQL Warehouses or Jobs when asked for the primary interface where code is interactively written, visualized, and shared.

35
MCQeasy

What is the primary purpose of the Databricks Catalog Explorer?

A.To manage the deployment and monitoring of automated machine learning models.
B.To browse and manage metadata for catalogs, schemas, tables, and volumes.
C.To monitor the health and logs of active Spark clusters.
D.To develop and execute Python scripts in interactive notebooks.
AnswerB

Catalog Explorer provides a visual interface to explore the Unity Catalog hierarchy, including catalogs, schemas, tables, and volumes. It enables users to inspect table schemas, view data samples, check permissions, and track lineage, which is the core function of the tool for data analysts navigating the data lakehouse.

Why this answer

Catalog Explorer is a unified interface within Databricks for managing and exploring the data ecosystem. It allows users to browse schemas, tables, and views, manage data permissions, and view data lineage. For a data analyst, it is the central hub for understanding data assets, verifying table metadata, and ensuring that the necessary access rights are in place before starting analysis, making it an essential tool for data discovery and governance.

Exam trap

Candidates frequently confuse Catalog Explorer with a data transformation tool or a query editor, forgetting that its primary purpose is metadata management, discovery, and governance.

36
MCQhard

Refer to the exhibit. An analyst receives this error when attempting to query a table. Which step should they take to fix this?

A.Increase the SQL Warehouse size to handle larger metadata lookups.
B.Check the Catalog Explorer to confirm the correct fully qualified table name.
C.Switch the notebook language to Python to access the table.
D.Request Workspace Admin privileges to access all system tables.
AnswerB

Catalog Explorer allows users to navigate the data hierarchy and identify the exact path to a table. By finding the table in the explorer, the analyst can determine its actual catalog, schema, and name, ensuring they use the correct fully qualified path in their query to resolve the reference error.

Why this answer

The error indicates a Namespace resolution issue. In Unity Catalog, objects must be referenced using the three-level namespace: catalog.schema.table. The error shows that the system is looking in 'main.default', but the table might be in a different catalog or schema.

The analyst needs to verify the correct path using the Catalog Explorer to ensure they are referencing the qualified table name correctly in their SQL query.

Exam trap

Candidates often try to debug the code syntax or the table schema, ignoring the most common cause of 'table not found' errors in Unity Catalog: using an incomplete or incorrect three-level namespace.

37
MCQmedium

When sharing an analysis with stakeholders, what is the best practice for ensuring they can view the results without needing access to the underlying raw data?

A.Grant the stakeholders 'Can Manage' permissions on the notebook.
B.Export the dashboard as a static PDF document.
C.Use Databricks Dashboards to share visual insights while keeping data access restricted.
D.Upload the underlying data files to a public cloud storage bucket.
AnswerC

Databricks Dashboards provide a controlled interface for stakeholders to consume insights. By sharing the dashboard, the analyst enables access to the visualizations while the underlying data security is managed by Unity Catalog. This ensures that users see only what they are authorized to see, maintaining the integrity of data governance policies.

Why this answer

To share results without exposing raw data, analysts should use Databricks Dashboards or materialized views. These tools allow for the creation of curated visualizations that present calculated results while keeping the raw data access restricted. This adheres to security best practices, ensuring that stakeholders see only the insights they need, while the underlying tables remain governed by strict Unity Catalog access policies that protect sensitive information from unauthorized exposure.

Exam trap

Candidates often suggest granting read access to raw data or exporting CSVs, failing to realize that Databricks Dashboards allow sharing insights securely without exposing the underlying storage layer.

38
MCQhard

An analyst runs a notebook against an all-purpose cluster and notices that the first cell, which reads a large Delta table, takes several minutes while subsequent similar queries finish quickly. The analyst wants to understand why this happens and ensure the same quick response on later runs. Which explanation best describes the underlying behavior?

A.The first query compiled a query plan that Databricks permanently stored in the table metadata
B.The first query used a different execution engine, and later queries switched to Photon
C.The Delta table was automatically optimized after the first read, rewriting all files
D.The cluster had to start and initialize, and Delta caching kept data in memory for later queries
AnswerD

When an all-purpose cluster starts, it must provision resources and initialize the Spark and Databricks runtime, which delays the first query. As the first read scans the Delta table, Databricks caches the underlying files on the cluster's local SSD, so subsequent queries against the same data skip remote storage reads and complete much faster. This explains both the initial delay and the later speedup.

Why this answer

The slow first cell reflects two overlapping costs: the all-purpose cluster must start and initialize before any query runs, and the initial scan reads Delta files from remote storage. Once that scan completes, Databricks caches the data on local SSD, so later queries that touch the same files avoid remote reads. Warm compute plus caching explains the dramatic speedup on subsequent runs.

Exam trap

The trap here is attributing the speedup to table optimization or an engine change, when the real cause is cluster startup latency combined with automatic disk caching of Delta data during the first read.

39
MCQhard

A data analyst has a notebook that reads a Delta table and produces summary statistics. The analyst wants colleagues to see the latest results in a dashboard without rerunning the notebook manually each time. Which Databricks capability should the analyst use to keep the dashboard current?

A.Add a markdown cell describing how to run the notebook
B.Increase the cluster size used by the notebook
C.Convert the notebook into a Databricks SQL dashboard and schedule the underlying query
D.Export the notebook results to a static PDF for distribution
AnswerC

A Databricks SQL dashboard is backed by saved queries that can be scheduled to refresh, so viewers see updated results without manual notebook runs. Migrating the summary logic into a query and scheduling it keeps the dashboard current automatically. This matches the goal of sharing live results while removing the need for manual execution.

Why this answer

Databricks SQL dashboards are built on saved queries that can be scheduled to refresh, so viewers see current results without anyone rerunning a notebook. Recreating the summary logic as a query and scheduling it automates currency. Documentation, larger clusters, and static exports do not refresh a dashboard, so they fail to meet the requirement of up-to-date shared results.

Exam trap

The trap here is treating a notebook visualization as automatically shared and refreshed, when dashboards require scheduled queries to stay current without manual runs.

40
MCQhard

A data analyst needs to ensure that sensitive information in a table is not visible to unauthorized users. Which Unity Catalog feature is the most efficient way to achieve this at the row level?

A.Create a separate view for each user group.
B.Apply Row Filters and Column Masks in Unity Catalog.
C.Use hard-coded SQL 'WHERE' clauses in every notebook.
D.Move sensitive data to a completely separate, private workspace.
AnswerB

Row filters and column masks are built-in Unity Catalog features that enforce fine-grained access control dynamically. They allow administrators to define security rules that automatically filter rows or mask sensitive column values based on the user's role or attributes, providing a robust, scalable, and centralized method for protecting sensitive data assets.

Why this answer

Row-level security (RLS) and column-level security are fundamental features of Unity Catalog that enable fine-grained access control. By applying these policies, organizations can ensure compliance with privacy regulations while maintaining data usability. Understanding how to define these filters allows analysts to secure data effectively without creating multiple copies of the same dataset for different user groups, which simplifies maintenance and ensures consistency.

Exam trap

Candidates often suggest creating separate filtered views or duplicate tables for different user groups, which violates governance best practices and creates maintenance overhead.

Ready to test yourself?

Try a timed practice session using only Understanding the Databricks Platform questions.