Courseiva
Importing Data →mediumMultiple Choice

Databricks-DA-Assoc Importing Data Practice Question

A data analyst needs to load a daily batch of 500 CSV files from an S3 bucket into a Delta table. The files are consistently formatted, and the analyst wants to ensure that files already processed are not re-imported in subsequent runs. Which approach is most efficient for this idempotency requirement?

⚠ Common exam trap

Candidates often suggest manual filtering or 'spark.read' with path lists. They ignore the built-in idempotency of 'COPY INTO', which natively tracks file state to prevent duplicate data ingestion.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Using the COPY INTO command to load the files from the S3 bucket.

COPY INTO is the ideal SQL command for idempotent batch loading from cloud object storage. It automatically tracks which files have been processed, preventing duplicate ingestion without requiring complex state management by the user. While Auto Loader is also idempotent, COPY INTO is often preferred for simpler batch SQL workflows where low-latency streaming is not a primary requirement for the analyst.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Using an INSERT INTO statement to append data from a temporary view.

    Why it's wrong here

    INSERT INTO requires manual logic to filter out previously loaded files, making it prone to errors and duplicates. It does not natively track file metadata or ingestion state, which leads to significant overhead for the data analyst when dealing with daily recurring batch transfers from external cloud storage locations.

  • ✗

    Using the MERGE INTO command to check for existing records based on a key.

    Why it's wrong here

    The MERGE INTO command is typically used for upserting data based on a specific key rather than initial ingestion from raw files. While it can prevent duplicates based on data values, it is significantly more computationally expensive and complex to configure for simple idempotent file loading compared to purpose-built ingestion commands.

  • ✓

    Using the COPY INTO command to load the files from the S3 bucket.

    Why this is correct

    COPY INTO provides a declarative way to load data while maintaining an internal record of processed files to ensure idempotency. This command simplifies the ingestion process by handling schema mapping and file discovery automatically, allowing analysts to run the same script repeatedly without risking data duplication in the target Delta table.

  • ✗

    Creating a standard SELECT query on the cloud location and appending results.

    Why it's wrong here

    Creating a temporary view on the raw files and using a standard SELECT statement provides no mechanism for tracking state. Every time the query is executed, all files in the directory will be re-processed and appended, leading to massive data redundancy and increased storage costs within the Databricks Lakehouse environment.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.