You are developing an Azure Databricks notebook that processes JSON files stored in Azure Data Lake Storage Gen2. You need to read the files into a DataFrame and automatically infer the schema. Which code should you use?
The spark.read.json method reads JSON files and infers the schema by default when no schema is provided. Using the abfss:// URI accesses ADLS Gen2 through the Azure Blob File System driver, which is the recommended protocol in Databricks. This single call returns a DataFrame with columns derived from the JSON structure, satisfying the requirement to infer the schema automatically.
Why this answer
The JSON data source in Spark automatically infers the schema when no schema is specified. Using spark.read.json with an abfss:// path reads the files from ADLS Gen2 and returns a DataFrame with inferred columns. The other options either misuse the API, use the wrong file format reader, or return unparsed text, so they do not meet the requirement.
Exam trap
The trap here is confusing the inferSchema option, which is used with the CSV reader, with schema inference for JSON, which happens by default without any option.