PL-300 Prepare the data Practice Question
You are connecting to a large CSV file (10 GB) stored in Azure Blob Storage. You need to load the data into Power BI with optimal performance. Which THREE practices should you follow? (Choose three.)
⚠ Common exam trap
Many exam-takers think importing all columns and hiding them is harmless, but Power BI still loads the full data into memory, wasting resources; the correct approach is to remove unnecessary columns and rows early in Power Query.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Parquet format instead of CSV if possible.
Option A is correct because Parquet is a columnar, compressed format that Power BI can read far more efficiently than CSV, reducing file size and scan time for a 10 GB dataset. Option C is correct because splitting the large CSV into multiple smaller files in the same folder enables Power BI's parallel loading and folder-combine pattern, improving throughput versus reading one monolithic file. Option D is correct because applying row filters in Power Query early reduces the volume of data pulled into the model, lowering memory and refresh cost. Option B is wrong because importing all columns then hiding them still loads every column into the model, wasting memory and bandwidth. Option E is wrong because the on-premises data gateway is only needed for on-premises data sources, not for a native cloud source like Azure Blob Storage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Parquet format instead of CSV if possible.
Why this is correct
Parquet is a columnar storage format with built-in compression and predicate pushdown, meaning Power Query can read only the specific columns and row groups required instead of scanning the entire 10 GB text file. This dramatically reduces the amount of I/O and memory consumed during the Data Refresh, and also preserves the schema and data types natively, eliminating costly type-inference and parsing steps. By shifting to Parquet, you effectively move the heavy lifting to the storage engine, which is far more efficient for analytical workloads.
- ✗
Import all columns and use Power BI to hide unused ones.
Why it's wrong here
Hiding columns from the Fields pane only modifies the report's metadata—the hidden columns are still fully loaded into the in-memory Tabular model, consuming RAM and enlarging the .pbit and .pbix files. During refresh, all rows and columns of your 10 GB CSV are parsed, compressed, and stored even if you never create a visual from them, and hidden columns can still be referenced by DAX measures, relationships, or even security rules. A best practice is to remove unused columns in Power Query before loading them into the model, not to import everything and rely on hiding as a cosmetic cleanup.
- ✓
Split the large CSV file into multiple smaller files in the same folder.
Why this is correct
Power Query's folder-import functionality reads each file individually and in sequence, so processing a 10 GB monolith peaks your memory usage at roughly the size of that one file; splitting it into, say, 100 MB chunks keeps peak memory proportional to the largest component, not the whole dataset. Each file is independently processed by Power Query, which allows for easier parallelization and simpler handling of partial, incremental refreshes, while the same transform steps automatically apply to every file. However, ensure all files share the same schema or your combine step may break, and be aware that you still read all rows if the rows aren't partitioned by date or some other filter.
- ✓
Filter rows in Power Query to remove unnecessary data early in the transformation.
Why this is correct
By applying row-level filters in the very first steps of your Power Query—before any join, pivot, or expansion—you shrink the number of rows that all subsequent transformations process, directly reducing both the refresh duration and the memory footprint of the query running against your 10 GB CSV. Since flat CSV files offer no server-side query folding, the filtering happens client-side after the raw text is streamed into the engine, but doing it immediately after Source means the engine never materializes an oversized copy of the data in its buffer. This is especially true when splitting files (or using Folder) because filtering there prunes rows as each file enters the pipeline, keeping the engine's working set lean.
- ✗
Use the on-premises data gateway to connect to Azure Blob Storage.
Why it's wrong here
The on-premises data gateway is exclusively designed to bridge connections from Power Query to data sources that live inside your private network, such as SQL Server or a local file share; Azure Blob Storage is a public cloud endpoint accessed over HTTPS, which the Power BI service can connect to directly from its own virtual networks. Forcing that direct connection through a gateway adds unnecessary latency, requires an always-on VM, and introduces a single point of failure, and also violates the principle of least complexity for cloud-to-cloud connections. The gateway is simply the wrong tool here; you only need it for cloud sources that reside in a private cloud or are firewalled off.
Visual reference
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PL-300 question from scratch — 524 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PL-300 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PL-300 exam.