Databricks-DA-Assoc Importing Data Practice Question
A data analyst needs to load a daily batch of 500 CSV files from an S3 bucket into a Delta table. The files are consistently formatted, and the analyst wants to ensure that files already processed are not re-imported in subsequent runs. Which approach is most efficient for this idempotency requirement?
⚠ Common exam trap
Candidates often suggest manual filtering or 'spark.read' with path lists. They ignore the built-in idempotency of 'COPY INTO', which natively tracks file state to prevent duplicate data ingestion.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Using the COPY INTO command to load the files from the S3 bucket.
COPY INTO is the ideal SQL command for idempotent batch loading from cloud object storage. It automatically tracks which files have been processed, preventing duplicate ingestion without requiring complex state management by the user. While Auto Loader is also idempotent, COPY INTO is often preferred for simpler batch SQL workflows where low-latency streaming is not a primary requirement for the analyst.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Using an INSERT INTO statement to append data from a temporary view.
Why it's wrong here
INSERT INTO requires manual logic to filter out previously loaded files, making it prone to errors and duplicates. It does not natively track file metadata or ingestion state, which leads to significant overhead for the data analyst when dealing with daily recurring batch transfers from external cloud storage locations.
- ✗
Using the MERGE INTO command to check for existing records based on a key.
Why it's wrong here
The MERGE INTO command is typically used for upserting data based on a specific key rather than initial ingestion from raw files. While it can prevent duplicates based on data values, it is significantly more computationally expensive and complex to configure for simple idempotent file loading compared to purpose-built ingestion commands.
- ✓
Using the COPY INTO command to load the files from the S3 bucket.
Why this is correct
COPY INTO provides a declarative way to load data while maintaining an internal record of processed files to ensure idempotency. This command simplifies the ingestion process by handling schema mapping and file discovery automatically, allowing analysts to run the same script repeatedly without risking data duplication in the target Delta table.
- ✗
Creating a standard SELECT query on the cloud location and appending results.
Why it's wrong here
Creating a temporary view on the raw files and using a standard SELECT statement provides no mechanism for tracking state. Every time the query is executed, all files in the directory will be re-processed and appended, leading to massive data redundancy and increased storage costs within the Databricks Lakehouse environment.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
About these practice questions
One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.