Courseiva
Develop data processing →easyMultiple Choice

DP-203 Develop data processing Practice Question

You are developing an Azure Databricks notebook that processes JSON files stored in Azure Data Lake Storage Gen2. You need to read the files into a DataFrame and automatically infer the schema. Which code should you use?

⚠ Common exam trap

Many candidates confuse the inferSchema option, which is used with the CSV reader, with schema inference for JSON, which happens by default without any option.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

spark.read.json("abfss://container@storage.dfs.core.windows.net/path")

The JSON data source in Spark automatically infers the schema when no schema is specified. Using spark.read.json with an abfss:// path reads the files from ADLS Gen2 and returns a DataFrame with inferred columns. The other options either misuse the API, use the wrong file format reader, or return unparsed text, so they do not meet the requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    spark.read.format("json").schema("infer").load("abfss://container@storage.dfs.core.windows.net/path")

    Why it's wrong here

    The schema method does not accept the string "infer"; it expects a StructType or a DDL-formatted string defining the schema. To infer the schema, you simply omit the schema call, as inference is the default behavior for JSON. This code would raise an error or misinterpret the schema definition, so it is not a valid way to infer the schema.

  • ✗

    spark.read.text("abfss://container@storage.dfs.core.windows.net/path")

    Why it's wrong here

    The text reader returns a DataFrame with a single string column named value, one row per line of the file. It does not parse JSON or infer any schema beyond that single column. This would not provide the structured columns expected from JSON data and would require additional parsing steps, failing the requirement to infer the schema automatically.

  • ✗

    spark.read.option("inferSchema", "true").csv("abfss://container@storage.dfs.core.windows.net/path")

    Why it's wrong here

    This uses the CSV reader with inferSchema, which is appropriate for delimited text files, not JSON. JSON is semi-structured and requires the JSON data source to parse nested objects and arrays. Reading JSON with the CSV reader would produce incorrect results, such as a single column containing raw JSON strings, and would not properly infer nested schema elements.

  • ✓

    spark.read.json("abfss://container@storage.dfs.core.windows.net/path")

    Why this is correct

    The spark.read.json method reads JSON files and infers the schema by default when no schema is provided. Using the abfss:// URI accesses ADLS Gen2 through the Azure Blob File System driver, which is the recommended protocol in Databricks. This single call returns a DataFrame with columns derived from the JSON structure, satisfying the requirement to infer the schema automatically.

About these practice questions

This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.