Databricks-DA-Assoc Importing Data Practice Question
A data analyst must load a Parquet file from an S3 bucket into a Databricks DataFrame, but the bucket is in a different AWS account and requires temporary credentials. The analyst has an AWS access key ID, secret access key, and session token. Which code snippet correctly configures Spark to read the file securely?
⚠ Common exam trap
The trap here is assuming that access key and secret key alone are sufficient for temporary credentials, forgetting that the session token is also required for S3A authentication.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3a://bucket/path/file.parquet")
To read from S3 with temporary credentials, you must set the access key, secret key, and session token using the fs.s3a.* configuration keys and use the s3a:// URI scheme. This ensures Spark authenticates correctly to the cross-account bucket. Omitting the session token or using the wrong scheme will cause failures, so all three elements are essential for a successful read.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
spark.conf.set("fs.s3.access.key", access_key) spark.conf.set("fs.s3.secret.key", secret_key) spark.conf.set("fs.s3.session.token", session_token) df = spark.read.parquet("s3://bucket/path/file.parquet")
Why it's wrong here
The configuration keys here use fs.s3 instead of fs.s3a, which is incorrect for the S3A filesystem. Databricks expects fs.s3a.* for S3A access. Additionally, the s3:// scheme may not be supported without additional configuration. This mismatch will cause authentication failures or unsupported filesystem errors, making the read operation unsuccessful.
- ✗
spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3://bucket/path/file.parquet")
Why it's wrong here
While the configuration keys are correct for S3A, the path uses s3:// instead of s3a://. Databricks typically requires the s3a:// scheme to leverage the S3A filesystem. Using s3:// may result in an error or fallback to a different implementation that does not honor the configured credentials, leading to access denial. The scheme must match the configured filesystem.
- ✗
spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) df = spark.read.parquet("s3a://bucket/path/file.parquet")
Why it's wrong here
This option omits the session token, which is mandatory when using temporary AWS credentials. Without the token, authentication to S3 will fail with a 403 error. While access key and secret key are necessary, they are insufficient for temporary credentials. The scenario explicitly mentions a session token, so ignoring it will prevent successful data access.
- ✓
spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3a://bucket/path/file.parquet")
Why this is correct
This is correct because S3A requires explicit configuration of access key, secret key, and session token for temporary credentials. Setting these via spark.conf.set ensures the Spark session can authenticate to the cross-account S3 bucket. The s3a:// scheme is appropriate for Hadoop-based access, and this approach avoids hardcoding credentials in code, aligning with Databricks security best practices for external data access.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
About these practice questions
This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.