Courseiva

CCNA Data Ingestion and Acquisition Questions

19 questions · Data Ingestion and Acquisition · All types, answers revealed

1
MCQmedium

What is the primary advantage of using Delta Lake as the sink for your data ingestion pipelines compared to raw Parquet files?

A.Delta Lake offers faster read performance for all queries.
B.Delta Lake provides ACID transactions and schema enforcement.
C.Delta Lake is required for all streaming sources.
D.Delta Lake automatically compresses data by 90% more than Parquet.
AnswerB

ACID transactions ensure that all writes are atomic, preventing partial updates to tables. Schema enforcement ensures that only data matching the defined schema is written, preventing the 'data swamp' problem. Together, these features provide the reliability required for modern data lakehouse architectures, which is not natively available with raw Parquet.

Why this answer

The primary advantage of Delta Lake over raw Parquet is its support for ACID transactions and schema enforcement. These features prevent data corruption during concurrent writes and ensure that only clean, well-structured data enters the lake. This reliability is fundamental for enterprise data pipelines, as it eliminates common issues like partial writes, corrupted files, and downstream failures caused by unexpected schema changes, which are difficult to manage with raw Parquet.

Exam trap

Candidates frequently choose performance-related features like compression or speed as the primary benefit, ignoring that ACID transactions and schema enforcement are the foundational architectural advantages of Delta Lake.

2
MCQhard

You are ingesting data from multiple source systems with varying file formats (JSON, CSV, Parquet) into a centralized Bronze landing zone. Which architecture pattern is the most scalable for maintaining this ingestion layer?

A.A single monolithic notebook with a massive if/else block for each source.
B.A separate ingestion pipeline or notebook for each source system.
C.Ingest everything into a single raw table before applying transformations.
D.Use a third-party tool exclusively and bypass Databricks for ingestion.
AnswerB

Isolating each source in its own pipeline allows for independent configuration, scheduling, and error handling. This modularity is essential for scalability, as it minimizes the blast radius of any individual pipeline failure and allows teams to manage and optimize ingestion logic based on the specific requirements of each source system.

Why this answer

A modular ingestion architecture that decouples source-specific logic from the common landing layer is the industry standard. By using parameterized notebooks or DLT pipelines for each source, you isolate potential failures, enable independent scaling, and simplify the management of schema mappings. This pattern ensures that changes in one source system do not impact the ingestion pipelines of others, creating a highly resilient and maintainable data acquisition ecosystem.

Exam trap

Candidates often suggest a 'monolithic' pipeline to handle all formats, failing to recognize that isolating source-specific logic is necessary for scalability and maintenance.

3
MCQmedium

A data engineer is building a streaming ingestion pipeline from Apache Kafka to a Delta table. The pipeline must perform deduplication on a unique event_id field and handle late-arriving data. The engineer wants to use Structured Streaming with a watermark of 10 minutes. Which of the following approaches correctly implements deduplication and watermarking?

A.Use dropDuplicates with event_id and the event timestamp column without setting a watermark.
B.Use a windowed aggregation on event_id and timestamp, then filter out duplicates using row_number.
C.Use foreachBatch to manually deduplicate by querying the Delta table for existing event_id values.
D.Use dropDuplicates with event_id and set the watermark on the event timestamp column.
AnswerD

This approach correctly uses dropDuplicates on the unique event_id to remove duplicates within the watermark window. The watermark on the event timestamp column allows the stream to handle late data by tracking how late data can arrive. This is the standard pattern for deduplication in Structured Streaming with watermarking, ensuring exactly-once semantics when combined with Delta Lake.

Why this answer

The correct approach is to use dropDuplicates on the unique event_id and set a watermark on the event timestamp column. This combination allows the stream to deduplicate events within the watermark window, effectively handling late-arriving data while bounding state. It is the recommended pattern for streaming deduplication in Databricks Structured Streaming.

Exam trap

The trap here is assuming that dropDuplicates alone can handle late data without a watermark, but without a watermark, state grows indefinitely and late data may be dropped incorrectly.

4
MCQmedium

A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the ingestion must handle late-arriving data and produce correct aggregations. The engineer wants to ensure that watermarks are applied correctly. Which approach should be used?

A.Use the Kafka timestamp column provided by the Kafka source, and apply a watermark on that column.
B.Use the current_timestamp() function to generate a processing-time column and apply a watermark on that column.
C.Parse the timestamp field from the Kafka value payload, cast it to a timestamp, and apply a watermark on that column.
D.Set the Kafka source option 'startingOffsets' to 'earliest' and rely on the default watermark.
AnswerC

To correctly handle late-arriving data based on event time, the watermark must be applied on the event-time column derived from the payload. Parsing the timestamp field, casting it to a timestamp type, and then applying a watermark on that column allows Structured Streaming to track event time and drop or update late data according to the watermark threshold. This is the standard approach for event-time processing with Kafka and Delta Lake.

Why this answer

For correct event-time processing with Kafka and Delta Lake, the event timestamp must come from the message payload, not from Kafka metadata or processing time. Parsing that field and applying a watermark on it enables Structured Streaming to manage late-arriving data accurately. This ensures that aggregations and stateful operations reflect the true event time and that late data is handled according to the defined watermark.

Exam trap

The trap here is assuming that the Kafka timestamp or processing time can be used for watermarks, but event-time watermarks must be based on the actual event timestamp from the payload.

5
Multi-Selecthard

A data engineer is building a streaming ingestion pipeline using Databricks Auto Loader to ingest JSON files from cloud storage into a Delta Bronze table. The pipeline must handle schema evolution without failing and must minimize the number of files that require reprocessing when the schema changes. The engineer wants to configure Auto Loader appropriately. Which two configuration settings should be used to achieve these requirements? (Choose two.)

Select 2 answers
A.Set cloudFiles.schemaLocation to a dedicated directory in cloud storage.
B.Set cloudFiles.inferColumnTypes to true.
C.Set cloudFiles.useNotifications to true.
D.Set cloudFiles.schemaEvolutionMode to 'addNewColumns'.
E.Set cloudFiles.format to 'parquet'.
AnswersA, D

The cloudFiles.schemaLocation specifies where Auto Loader stores the inferred schema and metadata about schema evolution. By providing a persistent location, Auto Loader can track schema changes over time and avoid reprocessing files that were already ingested with an older schema. This is essential for minimizing reprocessing and ensuring consistent schema management across pipeline restarts.

Why this answer

To handle schema evolution without failing and minimize reprocessing, Auto Loader must be configured to add new columns dynamically and to store schema metadata persistently. Setting the schema evolution mode to addNewColumns ensures new fields are incorporated, while specifying a schema location allows Auto Loader to track schema changes and avoid reprocessing already-ingested files. Together, these settings provide robust schema evolution with efficient incremental processing.

Exam trap

The trap here is assuming that enabling file notification or type inference alone will handle schema evolution and reduce reprocessing, when in fact schema evolution mode and schema location are the critical settings.

6
MCQmedium

When ingesting data using Auto Loader, what is the purpose of the 'cloudFiles.schemaLocation' parameter?

A.It specifies the target directory where the processed Delta table data is stored.
B.It stores metadata about the inferred schema and tracks evolution.
C.It defines the temporary directory used for shuffling large datasets during joins.
D.It is used to cache the raw JSON files before they are parsed.
AnswerB

The schema location is where Auto Loader saves the inferred schema and tracks historical changes. This allows the process to maintain state regarding the data structure, ensuring that subsequent batches are processed correctly even as the source schema changes over time across multiple runs or restarts.

Why this answer

The schema location is vital for Auto Loader because it stores the inferred schema and tracks schema evolution history. By persisting this information in a managed location, Auto Loader can resume ingestion after a job restart without re-inferring the schema from scratch. This ensures consistency and prevents potential ingestion failures caused by schema drift when processing new files in a long-running stream.

Exam trap

Candidates frequently mistake this parameter for the data storage location itself, confusing the metadata/schema tracking path with the actual raw data destination path used by Auto Loader.

7
MCQhard

A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the messages can arrive out of order by up to 10 minutes. The engineer wants to perform time-windowed aggregations on the ingested data while minimizing state store overhead. Which approach should be used to handle the out-of-order data correctly?

A.Use withWatermark on the event-time column extracted from the Kafka value and set the watermark to 10 minutes.
B.Configure the Kafka source with 'startingOffsets' set to 'latest' and enable auto-commit of offsets.
C.Repartition the stream by the Kafka partition ID and sort within each partition using a custom foreachBatch function.
D.Set the Kafka consumer property 'isolation.level' to 'read_committed' to ensure only committed messages are processed.
AnswerA

Watermarking allows the stream to handle late data by defining a threshold beyond which late events are dropped. By extracting the event-time column from the Kafka value and applying a 10-minute watermark, the aggregation can correctly include events that arrive within that window. This minimizes state store overhead because the watermark instructs the engine to clean up old state. It is the standard approach for out-of-order data in Structured Streaming.

Why this answer

To handle out-of-order data in Structured Streaming, you must use event-time processing with watermarks. Extracting the event-time column from the Kafka value and applying a watermark of 10 minutes allows the engine to wait for late data up to that threshold and then clean up state. This minimizes state store overhead and ensures correct time-windowed aggregations.

Other options do not address event-time semantics or state management.

Exam trap

The trap here is confusing Kafka consumer properties or offset management with event-time processing, leading to solutions that do not actually handle late-arriving data or manage state.

8
MCQmedium

A data engineer is ingesting streaming data from Apache Kafka into a Delta table using Databricks Structured Streaming. The engineer wants to ensure exactly-once processing and handle late-arriving data. Which combination of features should the engineer use?

A.Set the Kafka consumer group ID to a unique value and use foreachBatch to write to Delta.
B.Enable checkpointing and use watermarks with a time-based window.
C.Configure the Delta table with mergeSchema enabled and use append mode for writes.
D.Use the Kafka source with startingOffsets set to earliest and enable auto.offset.reset to latest.
AnswerB

Checkpointing ensures exactly-once processing by storing the offset and state information reliably, allowing recovery without reprocessing. Watermarks define how late data can arrive and still be processed, enabling the engine to drop overly late data and manage state. Together, they provide exactly-once semantics and handle late data in windowed aggregations, which is essential for reliable streaming ingestion from Kafka.

Why this answer

Exactly-once processing in Structured Streaming requires reliable checkpointing to track offsets and state. Watermarks are used to handle late-arriving data by defining a threshold beyond which data is dropped, and they enable state cleanup for windowed operations. Combining checkpointing with watermarks provides both exactly-once semantics and late data handling, making it the correct approach for ingesting Kafka data into Delta.

Exam trap

The trap here is thinking that consumer group settings or write modes alone can ensure exactly-once; checkpointing and watermarks are fundamental for stateful stream processing.

9
MCQeasy

A data engineer is using Databricks Auto Loader to ingest CSV files into a Delta table. The engineer notices that some files have a different delimiter (semicolon instead of comma). Which option should be used to handle this variation?

A.Set the cloudFiles.schemaEvolutionMode to 'addNewColumns' to handle delimiter changes.
B.Use the cloudFiles.format option with a custom delimiter per file.
C.Set the delimiter option to ';' for all files.
D.Preprocess the files to standardize the delimiter before ingestion.
AnswerD

Preprocessing the files to use a consistent delimiter (e.g., comma) is the most reliable approach. Auto Loader expects a uniform format; by standardizing delimiters, you ensure correct parsing. This can be done with a separate job or using a Databricks notebook to rewrite files. It avoids ingestion failures and data corruption.

Why this answer

The correct approach is to preprocess the files to standardize the delimiter before ingestion. Auto Loader does not support per-file delimiter detection; it requires a consistent format. By converting all files to a common delimiter, you ensure that Auto Loader parses them correctly and ingests data without errors.

This is a common preprocessing step in data pipelines.

Exam trap

The trap here is assuming Auto Loader can automatically detect and handle different delimiters, but it requires a consistent delimiter across all files.

10
MCQeasy

A data engineer is using Auto Loader to ingest JSON files from cloud storage into a Delta table. The files contain a nested field 'address' with subfields 'city' and 'zip'. The engineer wants to flatten the nested structure during ingestion so that 'city' and 'zip' become top-level columns in the Bronze table. Which Auto Loader feature should be used to achieve this?

A.Apply a select transformation with col('address.city') and col('address.zip') after reading the stream.
B.Set cloudFiles.flattenNested to true in the Auto Loader options.
C.Use the cloudFiles.schemaEvolutionMode set to 'addNewColumns' to automatically flatten nested fields.
D.Set cloudFiles.schemaHints to specify the nested fields as top-level columns.
AnswerA

Auto Loader does not have a built-in flattening option; flattening is achieved through DataFrame transformations. After reading the stream with Auto Loader, you can use select or withColumn to extract nested fields into top-level columns. For example, selecting col('address.city').alias('city') and col('address.zip').alias('zip') will produce the desired flat schema. This is the standard approach for flattening nested data during ingestion.

Why this answer

Auto Loader ingests data with its original nested structure. To flatten nested fields into top-level columns, you must apply DataFrame transformations after reading the stream. Using select or withColumn to extract subfields like address.city and address.zip is the correct method.

Auto Loader does not have a built-in flattening feature, so transformations are required.

Exam trap

The trap here is assuming that Auto Loader has a configuration option to flatten nested data automatically, when in fact flattening must be done via DataFrame operations after ingestion.

11
MCQmedium

Which approach is most appropriate for ingesting data from a JDBC source into Delta Lake where the source table has no 'updated_at' or 'version' column for incremental loading?

A.Use the 'partitionColumn' parameter with a random UUID.
B.Perform a full overwrite of the Delta table for every load.
C.Enable streaming ingestion using the JDBC source readStream API.
D.Use the 'fetchSize' parameter to optimize the load.
AnswerB

Since there is no mechanism to identify changed data, performing a full overwrite ensures the target table always matches the source. This is the standard pattern for handling tables without watermark columns. You should balance the frequency of the load with the size of the table to manage compute costs.

Why this answer

When a source table lacks a watermark column, you cannot use standard incremental loading techniques. The most robust approach is a full overwrite of the destination table, ensuring that the target remains a faithful copy of the source. While this can be resource-intensive for large tables, it is the only way to ensure data integrity without primary keys or timestamps to track changes.

Exam trap

Candidates often suggest using MERGE or incremental ingestion logic even when no watermark column exists, forgetting that these techniques require a reliable way to identify new or modified records.

12
MCQmedium

Your organization is ingesting sensitive PII data. You need to ensure that personal identifiers are masked during the ingestion process before they are stored in the Bronze layer of your Medallion architecture. What is the best practice for this?

A.Mask the data using Delta Lake column masking after the data reaches the Silver layer.
B.Apply masking logic within the initial streaming ingestion transformation.
C.Configure Unity Catalog to mask columns only for specific users.
D.Use a post-ingestion job to delete PII columns.
AnswerB

Applying masking logic during the initial ingestion transformation (e.g., using withColumn and sha2 or similar functions) ensures that PII is protected before it is ever committed to storage. This maintains data privacy from the moment the data enters the ecosystem, preventing sensitive information from ever reaching the persistent Bronze table.

Why this answer

Implementing masking at the ingestion layer using Delta Live Tables (DLT) expectations or standard Spark transformations ensures that sensitive data is never written in plaintext to the Bronze layer. Applying transformations during the 'Acquisition' phase is a critical security practice, ensuring that governance requirements are met before data becomes available to analysts, thereby reducing the risk of accidental exposure and maintaining data privacy compliance across the pipeline.

Exam trap

Students mistakenly think PII data should be masked downstream in the Gold layer for final reporting, overlooking the critical compliance requirement to secure raw data early.

13
MCQhard

Refer to the exhibit. You are using Auto Loader to ingest data with evolving schemas. After running the job for a week, you realize that new columns added to the source JSON are not being captured in the destination table. What must you add to the configuration?

A.Add 'cloudFiles.schemaEvolutionMode': 'rescue'.
B.Add 'cloudFiles.schemaEvolutionMode': 'addCol'.
C.Add 'cloudFiles.maxFiles': '1000'.
D.Add 'cloudFiles.allowOverwrites': 'true'.
AnswerB

Setting schema evolution mode to 'addCol' enables the ingestion process to detect new columns in the source data and automatically add them to the target Delta table. This ensures the table structure remains synchronized with the incoming data stream, preventing the loss of new attributes arriving in source files.

Why this answer

By default, Auto Loader schema inference only detects the schema during the initial load. To capture schema evolution, you must explicitly enable 'cloudFiles.schemaEvolutionMode'. Without this parameter, Auto Loader ignores new fields to protect downstream consumers from breaking changes.

Enabling 'addCol' allows the schema to expand dynamically, ensuring the target Delta table reflects the structure of the incoming data files as they arrive.

Exam trap

Candidates tend to look for manual ALTER TABLE commands or checkpoint resets, forgetting that Auto Loader needs a specific configuration property for schema evolution.

14
MCQmedium

Which THREE of the following are essential components of an effective ingestion monitoring strategy in Databricks?

A.Monitoring the total number of files in the cloud storage bucket.
B.Tracking the 'numInputRows' metric in the streaming query progress.
C.Setting up alerts on failed expectations in DLT.
D.Using DLT event logs to analyze pipeline execution details.
E.Regularly restarting the cluster to clear cache.
AnswerB, C, D

This metric tells you how many records are being processed in each micro-batch. It is essential for detecting data spikes, identifying potential ingestion lags, and verifying that the volume of data flowing through the pipeline aligns with the expected source throughput, which helps in capacity planning and performance tuning.

Why this answer

Monitoring ingestion requires visibility into both infrastructure health and data quality. By tracking metrics like micro-batch latency, file processing rates, and record-level validation through expectations, you gain a holistic view of the pipeline. These components allow engineers to identify bottlenecks, respond to data quality drops, and ensure SLAs are met consistently, which is critical for maintaining reliable downstream analytics in a data lakehouse architecture.

Exam trap

Candidates often select generic infrastructure metrics like CPU usage or memory, missing that ingestion monitoring specifically requires data-level metrics like record processing rates and quality expectations.

15
MCQmedium

A data engineer is ingesting data from an Azure SQL Database into a Delta Lake table using the JDBC connector in a Databricks notebook. The source table contains millions of rows, and the engineer wants to optimize the ingestion by reading the data in parallel. The source table has a numeric primary key column named 'id' that is evenly distributed. Which approach should the engineer use to achieve parallel reads?

A.Set the 'fetchsize' option to a large value to increase the number of rows retrieved per round trip.
B.Use the 'query' option with a custom SQL statement that includes a 'WHERE' clause with modulo arithmetic to split the data.
C.Specify the 'partitionColumn', 'lowerBound', 'upperBound', and 'numPartitions' options in the JDBC read configuration.
D.Enable 'spark.sql.adaptive.enabled' to automatically parallelize the JDBC read.
AnswerC

The JDBC connector supports parallel reads by partitioning the data based on a numeric column. By specifying partitionColumn (e.g., 'id'), along with lowerBound, upperBound, and numPartitions, Spark divides the query into multiple partitions that can be read concurrently. This significantly speeds up ingestion for large tables. The column must be numeric and evenly distributed, which 'id' satisfies. This is the correct approach for parallel JDBC reads.

Why this answer

To read a large JDBC table in parallel, you must configure the JDBC connector with partitioning options: partitionColumn, lowerBound, upperBound, and numPartitions. The partitionColumn must be numeric, and the bounds define the range for partitioning. Spark then creates multiple tasks to read the data concurrently.

This is the standard method for parallelizing JDBC reads in Databricks.

Exam trap

The trap here is thinking that increasing fetchsize or enabling adaptive query execution will parallelize the JDBC read, when in fact only explicit partitioning options create multiple concurrent tasks.

16
MCQmedium

You are designing an ingestion pipeline that must handle massive bursts of data at irregular intervals. Which feature should you prioritize to ensure the ingestion process remains cost-effective?

A.Use a fixed-size cluster with a large number of nodes.
B.Use auto-scaling clusters with a minimum of zero workers.
C.Run the pipeline continuously on a single-node cluster.
D.Increase the 'spark.sql.shuffle.partitions' to 10000.
AnswerB

Auto-scaling is the primary tool for managing bursty workloads. Setting the minimum to zero allows the cluster to shut down completely when no ingestion tasks are pending, effectively reducing costs to near zero. When data arrives, the cluster scales up automatically, ensuring throughput requirements are met during high-traffic periods.

Why this answer

Using Photon-accelerated clusters with auto-scaling is the most effective way to handle bursty workloads. By configuring the cluster to scale out when the queue of files is large and scale down during idle periods, you maximize resource utilization. This approach ensures you have enough compute to meet latency SLAs during bursts while minimizing costs by scaling to zero or minimal size when no data is being ingested, optimizing for both performance and budget.

Exam trap

Candidates often suggest fixed-size clusters to save money, ignoring that auto-scaling is essential for cost-effectively managing the irregular, bursty nature of the described workload.

17
MCQmedium

When ingesting data from a Kafka topic into Delta Lake, what is the best way to handle out-of-order data arriving in the stream?

A.Disable all aggregations to avoid processing errors.
B.Implement a watermark on the event time column.
C.Increase the 'spark.sql.shuffle.partitions' to 5000.
D.Use the 'Trigger.Once' execution mode.
AnswerB

Watermarking allows you to specify the maximum threshold for data lateness. Records arriving within this window are processed correctly, while data arriving later is dropped. This mechanism provides a mathematically sound way to balance correctness and system memory usage, effectively handling out-of-order data streams in a distributed environment.

Why this answer

In streaming systems, data can arrive delayed. Using Watermarking allows the system to define a threshold for how long it will wait for late data. By keeping state for the specified duration, the engine can correctly aggregate or join records that arrived out of order.

This is a standard and essential technique in Spark Structured Streaming to ensure the accuracy of time-windowed operations in high-throughput environments.

Exam trap

Candidates often propose using windowing functions or sorting the entire dataframe, which are inefficient and do not correctly handle the state management required for streaming late-arriving data.

18
MCQmedium

A data engineer is configuring an Auto Loader stream to ingest JSON files from an S3 bucket into a Bronze Delta table. The source bucket contains both .json and .json.gz files, and the engineer wants to ensure that only .json files are processed. Which parameter should be set to achieve this?

A.cloudFiles.schemaLocation
B.cloudFiles.includeExistingFiles
C.cloudFiles.format
D.cloudFiles.pathGlobFilter
AnswerD

The cloudFiles.pathGlobFilter parameter allows you to specify a glob pattern to filter files based on their path or extension. By setting it to '*.json', Auto Loader will only process files ending with .json, excluding .json.gz files. This is the correct way to selectively ingest files by extension in Auto Loader, ensuring only the desired files are read.

Why this answer

Auto Loader provides the cloudFiles.pathGlobFilter option to filter files using glob patterns. Setting it to '*.json' ensures that only files with the .json extension are processed, while other files like .json.gz are ignored. This is the intended mechanism for file selection based on path patterns in Auto Loader, making it the correct choice for this scenario.

Exam trap

The trap here is confusing the parameter that specifies the data format with the one that filters files by extension; cloudFiles.format defines parsing, not selection.

19
MCQhard

Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data ingestion over standard Structured Streaming pipelines?

A.DLT supports significantly higher throughput than Structured Streaming.
B.Declarative pipeline management and automated dependency handling.
C.Built-in data quality monitoring with Expectations.
D.DLT is the only way to read from cloud object storage.
E.DLT supports non-Delta storage formats for all outputs.
AnswerB, C

DLT allows you to define the pipeline in a declarative way, where the system automatically manages the creation and execution of the Directed Acyclic Graph (DAG) of dependencies. This eliminates the manual effort of coordinating complex streams, reducing operational overhead and the likelihood of human error in pipeline configuration.

Why this answer

DLT simplifies the operational complexity of data pipelines by providing automatic infrastructure management and built-in quality controls. It abstracts the configuration required for managing checkpoints, scaling, and handling schema drift, allowing engineers to focus on defining transformations. The 'Expectations' framework and declarative pipeline management provide superior observability and data quality enforcement compared to manual, imperative coding in standard Spark Structured Streaming.

Exam trap

Test-takers often confuse basic streaming features with DLT enhancements, overlooking declarative management and built-in quality expectations unique to DLT.

Ready to test yourself?

Try a timed practice session using only Data Ingestion and Acquisition questions.