Courseiva
Model Development →easyMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

A data scientist is preparing a dataset for training a model on Databricks. They want to split the data into training and testing sets, and they need to ensure that the split is reproducible across different runs. Which PySpark method should they use to split the DataFrame with a fixed random seed?

⚠ Common exam trap

It's easy for candidates to confuse sampling with splitting, when randomSplit is specifically built to return multiple disjoint DataFrames with a seed for reproducibility.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

df.randomSplit([0.8, 0.2], seed=42)

randomSplit is the PySpark method designed to split a DataFrame into multiple parts based on weights, and it accepts a seed for reproducibility. This allows the data scientist to obtain consistent training and testing sets across runs. Other methods like sample, repartition, or cache do not provide a reproducible train-test split.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    df.randomSplit([0.8, 0.2], seed=42)

    Why this is correct

    randomSplit splits a DataFrame into multiple DataFrames according to the provided weights, and the seed parameter ensures reproducibility. Using seed=42 guarantees that the same split is generated each time the code runs, which is essential for consistent model evaluation. This method is the standard PySpark approach for train-test splitting with a fixed random seed.

  • ✗

    df.repartition(2)

    Why it's wrong here

    repartition changes the number of partitions in the DataFrame but does not split rows into training and testing sets. It redistributes data across partitions, which affects parallelism but not data splitting. It has no seed parameter for reproducibility of splits. Using it here would not achieve the desired train-test division, and it could cause data shuffling without providing a test set.

  • ✗

    df.sample(fraction=0.8, seed=42)

    Why it's wrong here

    sample returns a random subset of the DataFrame but does not split it into complementary parts. It cannot produce both training and testing sets in one call, and the complement is not automatically generated. While it uses a seed, it is not designed for train-test splitting. The scientist would need additional logic to obtain the test set, making it less suitable than randomSplit.

  • ✗

    df.cache()

    Why it's wrong here

    cache persists the DataFrame in memory to speed up subsequent operations but does not split data. It is an optimization technique, not a splitting method. It has no concept of random seeds or train-test division. Using cache would not produce separate training and testing sets, so it fails to meet the requirement of reproducible splitting.

About these practice questions

Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.