Courseiva

CCNA Importing Data Questions

23 questions · Importing Data · All types, answers revealed

1
MCQeasy

A data analyst needs to quickly upload a small CSV file from their local machine into Databricks to explore its contents. They want to avoid writing code and prefer a graphical interface. Which Databricks feature should they use?

A.The Add Data UI
B.The Spark DataFrame API
C.The Databricks CLI
D.The DBFS CLI
AnswerA

The Add Data UI in Databricks provides a point-and-click interface for uploading files from a local machine. It supports CSV, JSON, and other formats, automatically infers schema, and creates a table. This is ideal for quick, one-time uploads without coding. It also allows previewing and editing column types before creating the table, making it suitable for analysts who prefer a visual approach.

Why this answer

The Add Data UI is specifically designed for uploading local files through a graphical interface, automatically creating tables and inferring schema. It requires no coding and is perfect for quick exploration. Other options involve command-line or programmatic interfaces, which are not as user-friendly for ad-hoc uploads.

Exam trap

The trap here is confusing the Add Data UI with command-line tools like the Databricks CLI or DBFS CLI, which also allow file uploads but require coding or terminal commands.

2
MCQmedium

A data analyst needs to ingest a large volume of CSV files from an external S3 bucket into a Delta table. Which method provides the most efficient, fault-tolerant, and incremental loading approach in Databricks?

A.Use the COPY INTO command with a static file path.
B.Manually read files into a DataFrame and use write.mode('append').
C.Implement Auto Loader using cloudFiles source.
D.Use the Databricks SQL 'Import' wizard via the UI.
AnswerC

Auto Loader uses cloud-native notifications and directory listing to detect new files efficiently. By managing its own checkpointing, it ensures exactly-once semantics and handles schema evolution automatically. It is the industry standard for production-grade ingestion pipelines in Databricks, providing superior performance and reliability compared to manual file-based ingestion methods.

Why this answer

Auto Loader is specifically designed to process files as they arrive in cloud storage. It maintains state information in a checkpoint location, ensuring that only new or modified files are ingested during subsequent runs. This pattern is critical for production pipelines where data volume scales over time, as it avoids full scans of the source directory, significantly reducing latency and compute costs while ensuring idempotent processing for reliability.

Exam trap

Candidates often choose COPY INTO for continuous, large-scale ingestion, failing to recognize that Auto Loader is the superior, stateful, and fault-tolerant choice for cloud file streams.

3
MCQhard

A data analyst is using the Databricks SQL Connector for Python to query a large Delta table. The query returns millions of rows, and the analyst wants to process the results in batches to avoid memory issues. Which method should the analyst use to fetch rows in chunks?

A.cursor.fetchone()
B.cursor.fetchall()
C.cursor.fetchmany(size)
D.cursor.execute() with a LIMIT clause
AnswerC

fetchmany(size) returns the next set of rows up to the specified size, enabling batch processing. This method is ideal for large result sets because it allows the analyst to iterate over chunks, reducing memory overhead. It is part of the Python DB API 2.0 specification and supported by the Databricks SQL Connector.

Why this answer

The fetchmany(size) method allows fetching a specified number of rows at a time, enabling batch processing of large result sets. This is the standard way to avoid loading all rows into memory. Other methods either fetch all rows, one row, or require manual query modifications.

Exam trap

The trap here is assuming that fetchall() is acceptable for large datasets or that fetchone() is efficient, when actually fetchmany() is designed for batch processing.

4
Multi-Selecthard

A data analyst is setting up a recurring ingestion job that must load only new files from a cloud storage directory into a Delta table, must track which files have already been processed, and must fail clearly if the source schema changes unexpectedly. Which TWO capabilities should the analyst rely on? (Choose two.)

Select 2 answers
A.Mounting the storage location and using spark.read on the directory with overwrite mode.
B.CREATE TABLE AS SELECT on the source directory rebuilds the table from all files on every run.
C.A scheduled notebook that lists the directory and compares filenames to a manually maintained table.
D.Auto Loader with a checkpoint location tracks processed files and supports schema enforcement or evolution policies.
E.COPY INTO maintains a load history so previously ingested file paths are skipped on subsequent runs.
AnswersD, E

Auto Loader uses a checkpoint directory to record which files it has consumed, enabling incremental processing of new arrivals. Its schema handling can be configured to fail on unexpected changes or to evolve, which matches the requirement to fail clearly when the source schema shifts. This makes it suitable for the recurring job described.

Why this answer

Both COPY INTO and Auto Loader maintain ingestion state so only new files are processed, and both can be configured to surface unexpected schema changes rather than silently accepting them. COPY INTO uses table-level load history, while Auto Loader uses a checkpoint directory with an explicit schema policy. Either satisfies the incremental and schema-safety requirements.

Exam trap

The trap here is reaching for a manual notebook or full-table rebuild when the managed ingestion tools already provide file tracking and schema policy handling.

5
Multi-Selectmedium

Which THREE of the following are valid sources for importing data into a Databricks Delta table? (Choose three)

Select 3 answers
A.S3, ADLS, and GCS cloud object storage.
B.Local files stored on the Databricks Driver node.
C.External relational databases via JDBC/ODBC connectors.
D.Direct memory buffers from the user's browser.
E.Streaming services like Apache Kafka or Amazon Kinesis.
AnswersA, C, E

These cloud object storage services are the native foundation for data lakes in Databricks. They are the most common sources for large-scale data ingestion and are fully supported by Auto Loader, COPY INTO, and standard Spark read APIs, making them the primary choice for enterprise data warehousing strategies.

Why this answer

Databricks provides flexible connectivity to ingest data from diverse sources into the Lakehouse. Whether the data is in cloud storage, a relational database, or a streaming message broker, the platform offers optimized connectors to load this information into Delta format. Mastering these integration points allows analysts to architect robust data pipelines that can pull information from virtually any enterprise system for unified analysis and reporting within a single environment.

Exam trap

Candidates often mistakenly include internal Databricks features like 'Delta Live Tables' or 'Unity Catalog' as data sources, failing to distinguish between data ingestion sources and governance tools.

6
MCQeasy

A data analyst has a 5 MB pipe-delimited text file with a header row on their local laptop and needs to load it into a Databricks Unity Catalog table for a one-time analysis. They want the fastest path with the least configuration. Which approach should they use?

A.Use the Add Data UI to upload the file and create the table through the guided wizard.
B.Run COPY INTO against the local filesystem path of the downloaded file.
C.Configure an Auto Loader stream pointing at a cloud storage location to ingest the file.
D.Create an external table over the local file so the data stays in place.
AnswerA

The Add Data UI is designed for exactly this case: small local files uploaded through the browser, with a wizard that previews the data, lets the analyst set the delimiter and header options, and creates a Unity Catalog table without writing code. It is the fastest low-configuration path for a one-time small-file load.

Why this answer

For a small local file that needs to become a table once, the Add Data UI is the intended workflow: it uploads the file, previews it, lets the analyst confirm delimiter and header handling, and creates the table. The other approaches assume cloud storage paths or streaming infrastructure that a laptop file simply does not have.

Exam trap

The trap here is assuming that any ingestion method can read a file sitting on a local laptop, when most Databricks ingestion paths require cloud storage paths.

7
MCQeasy

A data analyst needs to perform a one-time upload of a 15 MB compressed Parquet file from a local laptop into a Unity Catalog volume in a Databricks workspace. The analyst has the Databricks SQL editor open and no cluster running. Which approach should the analyst take to place the file into the volume path?

A.Navigate to the volume in Catalog Explorer, use the Upload to volume action, and select the local Parquet file to transfer it into the volume path.
B.Run COPY INTO from the SQL editor, referencing the local file path as the source and the volume path as the target.
C.Use dbfs cp from a terminal with the Databricks CLI to copy the file into /Volumes/catalog/schema/volume/path.
D.Use the Catalog Explorer's Create table UI to upload the file, selecting the target volume as the destination and the Parquet format.
AnswerA

Catalog Explorer exposes an Upload to volume action on Unity Catalog volumes that transfers local files directly into the volume directory without requiring a running cluster. This is the supported, low-friction path for small one-time uploads and places the file exactly where downstream SQL or notebook code can read it by its volume path.

Why this answer

Unity Catalog volumes are designed to hold non-tabular files, and Catalog Explorer provides an Upload to volume action that pushes a local file directly into the chosen volume directory. This avoids spinning up a cluster and does not require the SQL warehouse, making it the appropriate tool for a small, one-time upload from a laptop into a governed storage location.

Exam trap

The trap here is reaching for cluster-based or DBFS-based tooling such as COPY INTO or dbfs cp when a simple UI upload into a Unity Catalog volume is both supported and sufficient.

8
MCQeasy

An analyst is preparing to upload a small local CSV file to a Databricks workspace. Which tool is most appropriate for a quick, one-time upload without requiring infrastructure setup?

A.Auto Loader
B.Databricks Add Data UI
C.COPY INTO command
D.Databricks Connect
AnswerB

The Add Data UI allows users to drag and drop files from a local computer directly into the workspace. It creates a managed table or file path in DBFS, making it the most efficient and straightforward method for a small, one-time upload requirement by a data analyst.

Why this answer

The 'Add Data' UI in Databricks provides a user-friendly interface to upload small files directly from a local machine to DBFS. This approach is ideal for exploratory analysis or small datasets where setting up external storage or complex ingestion pipelines is unnecessary. Understanding this tool allows analysts to quickly pivot from local analysis to platform-based workflows without needing deep knowledge of cloud storage configurations.

Exam trap

Candidates often overcomplicate the solution by suggesting external cloud storage ingestion or CLI tools, failing to recognize the simplicity of the 'Add Data' UI for one-time, small-scale local file uploads.

9
Multi-Selectmedium

Which TWO of the following scenarios are valid use cases for using the COPY INTO command? (Choose two)

Select 2 answers
A.Ingesting data that arrives in files at unpredictable, irregular intervals.
B.Implementing a continuous streaming pipeline with sub-second latency.
C.Loading CSV files into Delta tables while performing basic transformation.
D.Performing complex, multi-stage ELT pipeline orchestration.
E.Ingesting historical data from a database using JDBC connectors.
AnswersA, C

COPY INTO is idempotent, meaning if you run the same command multiple times on the same files, it will not duplicate the data. This makes it perfect for irregular batch uploads where you just want to load everything that has landed since the last successful execution without building streams.

Why this answer

COPY INTO is a powerful, idempotent command for loading data into Delta tables. It is best suited for scenarios where you need to ingest files from cloud storage periodically without setting up complex streaming infrastructure. Understanding when to use COPY INTO versus Auto Loader helps analysts choose the right balance between simplicity and advanced streaming capabilities for their specific data engineering requirements.

Exam trap

Candidates often try to use COPY INTO for complex, real-time streaming transformations, failing to recognize it is optimized for simpler, idempotent file ingestion tasks.

10
MCQmedium

A data analyst runs COPY INTO to load new CSV files from an S3 path into a Delta table every night. One night the job completes successfully but the row count does not change, even though new files were placed in the source path. The files have the same names as files loaded previously. What is the most likely cause?

A.The CSV files were written with a different compression codec, so COPY INTO could not parse them.
B.COPY INTO skips files whose paths were already loaded unless the load is forced or the files are modified.
C.The Delta table's transaction log has reached its retention limit and stopped accepting new commits.
D.Auto Loader had already claimed the files, so COPY INTO is blocked from reading them.
AnswerB

COPY INTO tracks which file paths it has already ingested and by default does not reload them, even if the underlying object was replaced with new content. Because the new files reused previous names, their paths matched already-processed entries, so the command silently skipped them. This idempotent behavior is intentional but surprises analysts who overwrite files.

Why this answer

COPY INTO is idempotent by design: it records the paths it has loaded and skips them on subsequent runs unless the files are new or the load is explicitly forced. Because the new files reused prior names, their paths matched the load history, so no rows were added despite a successful job status.

Exam trap

The trap here is assuming that replacing a file's contents under the same path will cause COPY INTO to reload it, when path-based tracking means it is skipped.

11
MCQhard

An analyst is loading Parquet files from cloud storage into a Delta table using COPY INTO. The source directory contains files with several different schemas because upstream teams added columns over time. The analyst wants the table to accept the union of all columns without manual intervention. Which COPY INTO behavior should the analyst configure?

A.Enable overwrite mode so each run replaces the table with the latest file's schema.
B.Create the target table with a manually defined superset schema containing every possible column.
C.Use a permissive mode that silently drops columns not present in the target table schema.
D.Set mergeSchema to true in the COPY INTO options so new columns are added to the target table.
AnswerD

COPY INTO supports a mergeSchema option that automatically adds columns present in the source files but missing from the target table. This is exactly the union-of-columns behavior the analyst wants, and it avoids manual ALTER TABLE steps. Without it, files containing extra columns would be rejected or would fail schema matching.

Why this answer

The mergeSchema option in COPY INTO adds any columns found in the source files to the target table, producing the union of columns across heterogeneous files. This matches the requirement to accept new columns without manual ALTER TABLE statements, while retaining all previously loaded data.

Exam trap

The trap here is confusing overwrite behavior with schema merging, when overwrite replaces data and mergeSchema extends the table definition.

12
MCQmedium

Refer to the exhibit. An analyst is attempting to read a CSV file using Spark. Why is the code failing?

A.The cluster is running on a single-node configuration.
B.The file path provided does not exist in the mounted location.
C.Spark does not support reading CSV files directly from DBFS.
D.The user lacks sufficient memory to read the file.
AnswerB

The AnalysisException clearly states that the path does not exist. This indicates that either the mount point /mnt/data is not correctly established or the specific file sales_2023.csv is missing or misspelled in the underlying storage. The analyst must verify the mount point and the file existence.

Why this answer

This error occurs because the specified path in DBFS does not resolve to an actual file. This often happens if the mount point was not created correctly or if the file name was mistyped. Debugging file paths is a fundamental skill for analysts, as it involves verifying storage connectivity and ensuring the file system abstraction layer correctly points to the underlying cloud storage container where the data resides.

Exam trap

Candidates frequently assume the error is related to file permissions or Spark syntax, ignoring the most common issue: a simple typo or incorrect path string in the file system mount point.

13
MCQmedium

A data analyst must load a Parquet file from an S3 bucket into a Databricks DataFrame, but the bucket is in a different AWS account and requires temporary credentials. The analyst has an AWS access key ID, secret access key, and session token. Which code snippet correctly configures Spark to read the file securely?

A.spark.conf.set("fs.s3.access.key", access_key) spark.conf.set("fs.s3.secret.key", secret_key) spark.conf.set("fs.s3.session.token", session_token) df = spark.read.parquet("s3://bucket/path/file.parquet")
B.spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3://bucket/path/file.parquet")
C.spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) df = spark.read.parquet("s3a://bucket/path/file.parquet")
D.spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3a://bucket/path/file.parquet")
AnswerD

This is correct because S3A requires explicit configuration of access key, secret key, and session token for temporary credentials. Setting these via spark.conf.set ensures the Spark session can authenticate to the cross-account S3 bucket. The s3a:// scheme is appropriate for Hadoop-based access, and this approach avoids hardcoding credentials in code, aligning with Databricks security best practices for external data access.

Why this answer

To read from S3 with temporary credentials, you must set the access key, secret key, and session token using the fs.s3a.* configuration keys and use the s3a:// URI scheme. This ensures Spark authenticates correctly to the cross-account bucket. Omitting the session token or using the wrong scheme will cause failures, so all three elements are essential for a successful read.

Exam trap

The trap here is assuming that access key and secret key alone are sufficient for temporary credentials, forgetting that the session token is also required for S3A authentication.

14
MCQeasy

A junior data analyst attempts to load a large JSON file into a Databricks DataFrame using spark.read.json(path), but the resulting DataFrame contains numerous null values across critical columns. Upon inspection, the raw JSON records have inconsistent nesting structures and missing attributes. What is the most effective way for the analyst to inspect the inferred schema before transforming the data?

A.Execute the display() function on the raw file path string directly without creating a DataFrame object first.
B.Run the schema property or printSchema() method on the DataFrame to review the hierarchical structure and data types inferred by Spark.
C.Query the system information tables in the Hive metastore using a standard SQL DESCRIBE DATABASE statement.
D.Export the JSON file into Microsoft Excel locally to manually count the frequency of null values in each column.
AnswerB

printSchema() displays the inferred schema, including nested struct fields and data types, revealing why inconsistent nesting produced nulls. Reviewing this hierarchy lets the analyst correct the read options or define an explicit schema before transformation.

Why this answer

Using the printSchema() method on the loaded DataFrame prints out the hierarchical data types and inferred structure determined by Spark during the read operation. Data analysts rely heavily on this method to quickly identify schema evolution issues, unexpected nulls, and nested arrays before building downstream analytical queries and dashboards.

Exam trap

Candidates often try to use 'display()' or 'show()' to inspect schemas. While these show data content, they do not provide the structural metadata or hierarchical types needed to debug schema issues.

15
MCQmedium

A data analyst needs to load a daily batch of 500 CSV files from an S3 bucket into a Delta table. The files are consistently formatted, and the analyst wants to ensure that files already processed are not re-imported in subsequent runs. Which approach is most efficient for this idempotency requirement?

A.Using an INSERT INTO statement to append data from a temporary view.
B.Using the MERGE INTO command to check for existing records based on a key.
C.Using the COPY INTO command to load the files from the S3 bucket.
D.Creating a standard SELECT query on the cloud location and appending results.
AnswerC

COPY INTO provides a declarative way to load data while maintaining an internal record of processed files to ensure idempotency. This command simplifies the ingestion process by handling schema mapping and file discovery automatically, allowing analysts to run the same script repeatedly without risking data duplication in the target Delta table.

Why this answer

COPY INTO is the ideal SQL command for idempotent batch loading from cloud object storage. It automatically tracks which files have been processed, preventing duplicate ingestion without requiring complex state management by the user. While Auto Loader is also idempotent, COPY INTO is often preferred for simpler batch SQL workflows where low-latency streaming is not a primary requirement for the analyst.

Exam trap

Candidates often suggest manual filtering or 'spark.read' with path lists. They ignore the built-in idempotency of 'COPY INTO', which natively tracks file state to prevent duplicate data ingestion.

16
MCQeasy

When importing data into Databricks using the 'Add Data' UI, what is the default file format for the created table if the source is a CSV file?

A.CSV
B.Parquet
C.Delta
D.JSON
AnswerC

Delta Lake is the default table format in Databricks. When you use the UI to import a CSV file, Databricks performs an internal transformation to load that data into a Delta table, enabling all the benefits of the Lakehouse architecture, including ACID guarantees and optimized performance for SQL queries.

Why this answer

Databricks defaults to Delta Lake for all managed table creation because it provides ACID transactions, scalability, and time travel. Even when importing simple CSV files, Databricks automatically converts them into the Delta format to ensure that the data is queryable with high performance and reliability. Recognizing this default behavior is essential for understanding how Databricks standardizes data storage behind the scenes for all analytical workloads.

Exam trap

Candidates often assume the uploaded file retains its original CSV format, failing to realize that Databricks automatically converts all ingested data into Delta format to ensure ACID compliance and performance.

17
MCQmedium

An analyst is loading a CSV file where one column contains values like 007, 012, and 003. After creating a table through the Add Data UI, those values appear as 7, 12, and 3. The analyst needs the leading zeros preserved for downstream reporting. What should the analyst do?

A.Specify the column as a string type in the schema settings before creating the table.
B.Apply a formatting function after loading that pads the numbers back to three characters.
C.Enable schema evolution so the column can switch types after the first load.
D.Set the file's delimiter to a character that prevents numeric interpretation of the column.
AnswerA

Leading zeros are only meaningful in a string representation. If the column is inferred or declared as an integer, the parser strips the zeros because 007 and 7 are numerically identical. Overriding the inferred type to string preserves the original characters. The Add Data UI and schema editor both allow this type override before the table is created.

Why this answer

Leading zeros carry meaning only in text form, so the column must be declared as a string before the table is created. Letting the reader infer an integer type silently drops the zeros, and no post-load fix reliably restores them. Overriding the type in the schema settings preserves the original values.

Exam trap

The trap here is trusting automatic type inference, which correctly identifies a numeric pattern but destroys leading zeros that matter for identifiers.

18
MCQmedium

When importing data into Databricks, what is the primary benefit of using Delta Lake over standard Parquet files?

A.Delta Lake files require less storage space than Parquet.
B.Delta Lake supports ACID transactions and schema enforcement.
C.Delta Lake is faster for reading large tables than Parquet.
D.Delta Lake is natively supported by all external BI tools.
AnswerB

Delta Lake provides ACID transactions, which allow multiple readers and writers to interact with the data without corruption. It also enforces schema constraints, ensuring that new data matches the table structure. This makes it far more robust than standard Parquet, which lacks these native management features for distributed data.

Why this answer

Delta Lake adds a transaction log (the _delta_log folder) to Parquet files, enabling ACID transactions and time travel. This prevents data corruption during concurrent writes and allows users to query previous versions of data. This is a foundational concept for data reliability, as it ensures that analytical queries are always executed against a consistent and version-controlled state, which is impossible with standard, non-transactional Parquet files in a distributed system.

Exam trap

Candidates frequently confuse Delta Lake with simple storage formats like Parquet or Avro, overlooking that Delta's primary value proposition is the transaction log enabling ACID properties and concurrency.

19
Multi-Selecthard

A data analyst is using the Databricks Add Data UI to import a CSV file from cloud storage into a Delta table. The analyst wants to ensure the import process handles the data correctly. Which two actions can the analyst perform directly in the Add Data UI? (Choose two.)

Select 2 answers
A.Set up a scheduled job to automatically ingest new files from the same location.
B.Enable schema evolution for future files.
C.Define a custom schema using DDL statements.
D.Specify a custom delimiter for the CSV file.
E.Preview the data and manually adjust column types before creating the table.
AnswersD, E

The Add Data UI allows the analyst to specify a custom delimiter if the CSV file does not use a comma. This is useful for files that use semicolons, tabs, or other separators. The UI provides an option to set the delimiter during the import process.

Why this answer

The Add Data UI allows analysts to preview data, adjust column types, and specify a custom delimiter during import. It does not support scheduling, DDL schema definition, or schema evolution, which are features of other tools like Auto Loader or Jobs.

Exam trap

The trap here is assuming the Add Data UI can handle ongoing ingestion or complex schema definitions, when it is limited to interactive, one-time imports with basic adjustments.

20
Multi-Selecthard

Which THREE actions occur when using Auto Loader with schema evolution enabled? (Choose three)

Select 3 answers
A.New columns found in the source files are automatically added to the target table.
B.Existing columns with data type conflicts cause the pipeline to stop immediately.
C.The schema is inferred from the metadata of the cloud storage provider.
D.A 'rescued data' column is created to store data that does not match the inferred schema.
E.The inferred schema is saved to a checkpoint location to maintain consistency across restarts.
AnswersA, D, E

Auto Loader detects new columns in source data and can automatically update the target table's schema. This prevents pipeline failure when source data expands, allowing the table to evolve alongside the data, which is a key advantage of the Delta Lake schema evolution capabilities provided during ingestion.

Why this answer

Schema evolution is a powerful feature of Auto Loader that allows pipelines to adapt to changes in upstream data without requiring manual code updates. By understanding how the schema is managed, detected, and reconciled, analysts can build flexible ingestion systems that reduce maintenance overhead. This is critical in modern data environments where upstream sources frequently add or rename columns, requiring the platform to dynamically update downstream tables to accommodate the new structure.

Exam trap

Candidates mistakenly believe Auto Loader automatically deletes unparseable files or alters original source files, when it actually captures corrupted data in a rescued data column.

21
MCQmedium

An analyst is using the Spark DataFrame API to read a large JSON dataset. The dataset contains nested fields that are causing schema inference to fail. Which approach best resolves this?

A.Use the option 'mergeSchema' set to true.
B.Define a schema explicitly using StructType.
C.Convert the JSON to CSV before reading it.
D.Increase the number of partitions during read.
AnswerB

Explicitly defining the schema using StructType ensures that Spark knows exactly how to map the nested JSON fields to the DataFrame columns. This eliminates the need for Spark to perform a full file scan for inference, preventing errors on complex datasets and ensuring consistent data types for downstream analytical tasks.

Why this answer

When dealing with complex, nested JSON, relying on automated schema inference often fails or results in poor performance. Manually defining a schema using StructType allows the analyst to explicitly control data types, including handling nested structures. This approach is more robust for production-grade pipelines where data consistency is required, as it prevents schema mismatch errors during ingestion and improves overall query performance and type safety within the Spark environment.

Exam trap

Candidates often think schema inference handles nested JSON automatically, but complex or deep nesting typically causes inference failures or incorrect data types, requiring an explicit StructType definition.

22
MCQmedium

A data analyst needs to ingest a daily batch of comma-separated value (CSV) files stored in cloud object storage into a Databricks Delta table. The files contain a header row, use a semicolon (;) as the delimiter, and occasionally include embedded newlines within quoted fields. Which approach ensures the parser correctly handles the multi-line fields and missing values?

A.Use spark.read.table() with default parameters because Spark automatically infers delimiters and multiline configurations for all incoming CSV files.
B.Apply the spark.read.format("csv").option("delimiter", ";").option("header", "true").option("multiLine", "true").load(path) command.
C.Convert the CSV files into Parquet format using a local Python script outside Databricks before uploading them to the workspace file system.
D.Execute a standard SQL COPY INTO command without specifying any options, as Delta tables automatically resolve non-standard delimiters.
AnswerB

Setting the delimiter option to a semicolon, enabling header parsing, and setting multiLine to true allows Spark to correctly interpret complex rows. This combination handles formatting variations common in enterprise data dumps without dropping records.

Why this answer

Configuring the Spark CSV reader with explicit options like multiLine set to true ensures that records containing embedded newlines are not split prematurely across rows. Data analysts frequently encounter messy raw external datasets where standard single-line parsing fails, leading to corrupted row counts and schema mismatches. Explicitly defining formatting options guarantees reliable ingestion before downstream business intelligence reporting takes place.

Exam trap

Candidates frequently overlook the 'multiLine' option. They assume standard CSV parsing handles embedded newlines automatically, which leads to corrupted data rows and schema errors during the ingestion process.

23
Multi-Selecthard

A user wants to import a local 50MB CSV file into a Databricks workspace for a quick ad-hoc analysis. Which TWO methods are available within the Databricks UI to accomplish this task directly?

Select 2 answers
A.Using the 'Create or modify table from file upload' feature in the 'Add' menu.
B.Using the Databricks File System (DBFS) CLI to push the file to the workspace.
C.Using the SQL Editor to copy and paste the raw text into a new table.
D.Uploading the file to a Unity Catalog Volume using the Catalog Explorer.
E.Using the Workspace Sidebar 'Import' button to select the local CSV file.
AnswersA, D

The 'Create or modify table from file upload' interface allows users to drag and drop files directly into the browser. This tool automatically infers the schema, provides a data preview, and handles the conversion of the local CSV into a managed Delta table within the selected schema and catalog.

Why this answer

Databricks provides several entry points for local file uploads to facilitate rapid analysis. The 'Create or modify table from file upload' UI is the primary method for non-programmatic ingestion, while Unity Catalog Volumes offer a modern way to manage non-tabular data. Both options allow analysts to bypass complex cloud storage configurations when dealing with small, local datasets for immediate exploratory data science tasks.

Exam trap

Candidates often assume they must use 'dbutils.fs.put' or write complex code to upload files. They overlook the simple, built-in UI features available directly within the Databricks Catalog Explorer.

Ready to test yourself?

Try a timed practice session using only Importing Data questions.