Courseiva
Prepare the data →hardMultiple Select

PL-300 Prepare the data Practice Question

You are connecting to a large CSV file (10 GB) stored in Azure Blob Storage. You need to load the data into Power BI with optimal performance. Which THREE practices should you follow? (Choose three.)

⚠ Common exam trap

Many exam-takers think importing all columns and hiding them is harmless, but Power BI still loads the full data into memory, wasting resources; the correct approach is to remove unnecessary columns and rows early in Power Query.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Parquet format instead of CSV if possible.

Option A is correct because Parquet is a columnar, compressed format that Power BI can read far more efficiently than CSV, reducing file size and scan time for a 10 GB dataset. Option C is correct because splitting the large CSV into multiple smaller files in the same folder enables Power BI's parallel loading and folder-combine pattern, improving throughput versus reading one monolithic file. Option D is correct because applying row filters in Power Query early reduces the volume of data pulled into the model, lowering memory and refresh cost. Option B is wrong because importing all columns then hiding them still loads every column into the model, wasting memory and bandwidth. Option E is wrong because the on-premises data gateway is only needed for on-premises data sources, not for a native cloud source like Azure Blob Storage.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Parquet format instead of CSV if possible.

    Why this is correct

    Parquet is a columnar storage format with built-in compression and predicate pushdown, meaning Power Query can read only the specific columns and row groups required instead of scanning the entire 10 GB text file. This dramatically reduces the amount of I/O and memory consumed during the Data Refresh, and also preserves the schema and data types natively, eliminating costly type-inference and parsing steps. By shifting to Parquet, you effectively move the heavy lifting to the storage engine, which is far more efficient for analytical workloads.

  • ✗

    Import all columns and use Power BI to hide unused ones.

    Why it's wrong here

    Hiding columns from the Fields pane only modifies the report's metadata—the hidden columns are still fully loaded into the in-memory Tabular model, consuming RAM and enlarging the .pbit and .pbix files. During refresh, all rows and columns of your 10 GB CSV are parsed, compressed, and stored even if you never create a visual from them, and hidden columns can still be referenced by DAX measures, relationships, or even security rules. A best practice is to remove unused columns in Power Query before loading them into the model, not to import everything and rely on hiding as a cosmetic cleanup.

  • ✓

    Split the large CSV file into multiple smaller files in the same folder.

    Why this is correct

    Power Query's folder-import functionality reads each file individually and in sequence, so processing a 10 GB monolith peaks your memory usage at roughly the size of that one file; splitting it into, say, 100 MB chunks keeps peak memory proportional to the largest component, not the whole dataset. Each file is independently processed by Power Query, which allows for easier parallelization and simpler handling of partial, incremental refreshes, while the same transform steps automatically apply to every file. However, ensure all files share the same schema or your combine step may break, and be aware that you still read all rows if the rows aren't partitioned by date or some other filter.

  • ✓

    Filter rows in Power Query to remove unnecessary data early in the transformation.

    Why this is correct

    By applying row-level filters in the very first steps of your Power Query—before any join, pivot, or expansion—you shrink the number of rows that all subsequent transformations process, directly reducing both the refresh duration and the memory footprint of the query running against your 10 GB CSV. Since flat CSV files offer no server-side query folding, the filtering happens client-side after the raw text is streamed into the engine, but doing it immediately after Source means the engine never materializes an oversized copy of the data in its buffer. This is especially true when splitting files (or using Folder) because filtering there prunes rows as each file enters the pipeline, keeping the engine's working set lean.

  • ✗

    Use the on-premises data gateway to connect to Azure Blob Storage.

    Why it's wrong here

    The on-premises data gateway is exclusively designed to bridge connections from Power Query to data sources that live inside your private network, such as SQL Server or a local file share; Azure Blob Storage is a public cloud endpoint accessed over HTTPS, which the Power BI service can connect to directly from its own virtual networks. Forcing that direct connection through a gateway adds unnecessary latency, requires an always-on VM, and introduces a single point of failure, and also violates the principle of least complexity for cloud-to-cloud connections. The gateway is simply the wrong tool here; you only need it for cloud sources that reside in a private cloud or are firewalled off.

Visual reference

R1 R2 R3 R4 10 100 10 100 OSPF picks R1→R2→R4 (cost 20) over R1→R3→R4 (cost 200)

Quick reference

Azure Blob Storage Tier Comparison

TierStorage CostRetrieval CostLatencyUse Case
HotHighestLowestImmediateActive data, frequent reads
CoolLowerHigherImmediateData accessed < once / month
ColdLower stillHigherImmediateData accessed < once / quarter
ArchiveLowestHighest + rehydration delayHoursLong-term compliance retention

About these practice questions

Courseiva writes every PL-300 question from scratch — 524 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PL-300 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PL-300 exam.