You are analyzing a large dataset using Pandas API on Spark. You have a Pandas-on-Spark DataFrame `psdf` that was created from a Spark DataFrame with multiple partitions. You call `psdf.head(10)` to quickly inspect the data. What does this operation return?
This is correct. In Pandas API on Spark, `head(n)` collects the first n rows from the distributed DataFrame and returns them as a local pandas DataFrame. This allows you to inspect a small sample without triggering a full distributed computation. The operation is efficient because it only processes the necessary partitions and brings a limited amount of data to the driver.
Why this answer
The `head(n)` method in Pandas API on Spark returns a local pandas DataFrame containing the first n rows. It is intended for quick inspection of data, and it collects only the necessary rows to the driver. This behavior matches the pandas API, where `head` returns a DataFrame, but here it triggers a small action to bring data locally.
Exam trap
The trap here is assuming that `head` returns a distributed DataFrame, but it actually returns a local pandas DataFrame for immediate inspection.