Courseiva
Importing Data →mediumMultiple Choice

Databricks-DA-Assoc Importing Data Practice Question

A data analyst must load a Parquet file from an S3 bucket into a Databricks DataFrame, but the bucket is in a different AWS account and requires temporary credentials. The analyst has an AWS access key ID, secret access key, and session token. Which code snippet correctly configures Spark to read the file securely?

⚠ Common exam trap

The trap here is assuming that access key and secret key alone are sufficient for temporary credentials, forgetting that the session token is also required for S3A authentication.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3a://bucket/path/file.parquet")

To read from S3 with temporary credentials, you must set the access key, secret key, and session token using the fs.s3a.* configuration keys and use the s3a:// URI scheme. This ensures Spark authenticates correctly to the cross-account bucket. Omitting the session token or using the wrong scheme will cause failures, so all three elements are essential for a successful read.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    spark.conf.set("fs.s3.access.key", access_key) spark.conf.set("fs.s3.secret.key", secret_key) spark.conf.set("fs.s3.session.token", session_token) df = spark.read.parquet("s3://bucket/path/file.parquet")

    Why it's wrong here

    The configuration keys here use fs.s3 instead of fs.s3a, which is incorrect for the S3A filesystem. Databricks expects fs.s3a.* for S3A access. Additionally, the s3:// scheme may not be supported without additional configuration. This mismatch will cause authentication failures or unsupported filesystem errors, making the read operation unsuccessful.

  • ✗

    spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3://bucket/path/file.parquet")

    Why it's wrong here

    While the configuration keys are correct for S3A, the path uses s3:// instead of s3a://. Databricks typically requires the s3a:// scheme to leverage the S3A filesystem. Using s3:// may result in an error or fallback to a different implementation that does not honor the configured credentials, leading to access denial. The scheme must match the configured filesystem.

  • ✗

    spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) df = spark.read.parquet("s3a://bucket/path/file.parquet")

    Why it's wrong here

    This option omits the session token, which is mandatory when using temporary AWS credentials. Without the token, authentication to S3 will fail with a 403 error. While access key and secret key are necessary, they are insufficient for temporary credentials. The scenario explicitly mentions a session token, so ignoring it will prevent successful data access.

  • ✓

    spark.conf.set("fs.s3a.access.key", access_key) spark.conf.set("fs.s3a.secret.key", secret_key) spark.conf.set("fs.s3a.session.token", session_token) df = spark.read.parquet("s3a://bucket/path/file.parquet")

    Why this is correct

    This is correct because S3A requires explicit configuration of access key, secret key, and session token for temporary credentials. Setting these via spark.conf.set ensures the Spark session can authenticate to the cross-account S3 bucket. The s3a:// scheme is appropriate for Hadoop-based access, and this approach avoids hardcoding credentials in code, aligning with Databricks security best practices for external data access.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.