Courseiva
Data Preparation →mediumMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

A GenAI engineer must build a fine-tuning dataset from 40 TB of raw JSONL conversation logs stored in cloud object storage. The logs are immutable and only ever read once, in full, during preprocessing. Which storage and access configuration should be used to minimize cost while keeping the data readable by Spark on Databricks?

⚠ Common exam trap

The trap here is assuming that ingesting raw logs into Delta is always the right first step, when single-pass immutable reads are cheaper and better governed when left in place.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Register the cloud path as an external location in Unity Catalog and read the JSONL files directly with Spark.

Because the logs are immutable and read once in full, the lowest-cost approach is to leave them where they are and read them in place under Unity Catalog governance. Copying or converting the data creates a redundant full-size copy and extra compute for no reuse benefit, while legacy mounts sacrifice governance without reducing storage or access cost.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the JSONL to Parquet with an external tool, upload the Parquet back, then read it from the bucket.

    Why it's wrong here

    Rewriting JSONL to Parquet outside Databricks duplicates the full 40 TB and requires compute and egress outside the governed platform. Because the data is consumed exactly once, columnar compression yields no repeated-scan savings, so the conversion cost is pure overhead and the resulting files are still unmanaged by Unity Catalog.

  • ✗

    Run COPY INTO to ingest the JSONL into a managed Delta table, then read the Delta table for preprocessing.

    Why it's wrong here

    COPY INTO into a managed Delta table rewrites all 40 TB into Delta Parquet plus transaction log, doubling storage and adding write cost for data that is read only once. Since the logs are immutable and single-pass, the extra persisted copy delivers no reuse benefit, so this option increases rather than minimizes cost.

  • ✓

    Register the cloud path as an external location in Unity Catalog and read the JSONL files directly with Spark.

    Why this is correct

    Reading the immutable JSONL directly from its object-storage path through a Unity Catalog external location avoids copying 40 TB and avoids paying for a second persisted copy. Unity Catalog governs access while Spark reads the files in place, which matches the single-pass, read-only access pattern and minimizes both storage and egress cost.

  • ✗

    Mount the object storage bucket to DBFS with a legacy mount and read files through the /mnt path.

    Why it's wrong here

    Legacy DBFS mounts are discouraged and do not provide Unity Catalog governance, so fine-grained access control and lineage on the training data are lost. The mount also adds no cost benefit over an external location; it merely changes the path, so it fails the governance requirement without improving the read-only economics.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.