Question 1mediummultiple choice
Read the full Pandas API on Spark explanation →Databricks-Spark-Assoc Pandas API on Spark • Complete Question Bank
Complete Databricks-Spark-Assoc Pandas API on Spark question bank — all 0 questions with answers and detailed explanations.
Refer to the exhibit. [ERROR LOG] org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 1.0 failed 4 times, last failure: ResultTask(0, 4) failed: java.lang.UnsupportedOperationException: 'to_pandas' is not supported on a DataFrame with a very large size. [END OF EXHIBIT]
A data engineer is using Pandas API on Spark and needs to perform a join between two Pandas-on-Spark DataFrames `psdf1` and `psdf2` on a common column `id`. They write the following code:
```python result = psdf1.merge(psdf2, on='id', how='inner') ```
Which statement best describes the execution and potential issue with this operation?