AI-102 Practice Question: Implement knowledge mining and information extraction solutions
You are using Microsoft Purview to create a knowledge map of your organization's data assets. The solution must automatically scan and classify sensitive data in Azure Blob Storage. You need to configure the scanning and classification. Which THREE actions should you perform?
⚠ Common exam trap
Candidates often confuse the required actions for configuring scanning (registering the source, creating a scan rule set, and running a scan) with optional or subsequent steps like creating custom rules or applying sensitivity labels, leading them to select B or C instead of the correct three.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run a full scan of the Blob Storage to discover and classify data.
Option E is correct because before Microsoft Purview can scan any asset, the Azure Blob Storage account must first be registered as a data source in the Purview governance portal, which establishes the connection and allows the account to be managed and scanned. Option D is correct because a scan rule set defines which file types and classification rules are applied during a scan; you must create or select a rule set that includes the desired system or custom classification rules so sensitive data types are detected. Option A is correct because after registering the source and configuring the rule set, you must run a scan (a full scan for initial discovery and classification) so Purview can crawl the Blob Storage, apply the rule set, and populate the knowledge map with classified assets. Option B is not required because Purview provides built-in system classification rules for common sensitive data types, and custom rules are only needed for organization-specific patterns, which the scenario does not require. Option C is not part of the scanning and classification configuration; sensitivity labels are applied through Microsoft Purview Information Protection and are a separate labeling concern from the data map scanning process.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Run a full scan of the Blob Storage to discover and classify data.
Why this is correct
A full scan is what actually triggers Purview to crawl the registered Blob Storage, apply the chosen scan rule set, and populate the knowledge map with discovered assets and classifications. Without running the scan, no classification occurs, so this action satisfies the requirement for automatic discovery and classification.
- ✗
Create custom classification rules for sensitive data types.
Why it's wrong here
Custom classification rules define your own regex or dictionary patterns, yet the scenario requires automatic detection of built-in sensitive types, which system classifications already cover. It is tempting because custom rules are the correct choice when standard classifiers miss organisation-specific data formats.
- ✗
Apply sensitivity labels to the classified data.
Why it's wrong here
Sensitivity labels are applied after classification to govern usage, not to scan and classify Blob Storage content. It is tempting because labels do tag sensitive assets, but that is the correct action once scanning has already identified the data, not part of configuring the scan itself.
- ✓
Create a scan rule set that includes the desired classification rules.
Why this is correct
A scan rule set defines which classification rules Purview applies during a scan, including built-in sensitive information types and custom regex patterns. Creating one containing the desired rules ensures Blob Storage contents are classified against the correct sensitivity definitions, satisfying the automatic classification requirement.
- ✓
Register the Azure Blob Storage account as a data source in Purview.
Why this is correct
Registering the Blob Storage account as a data source creates the Purview resource entry that scans target. Scanning cannot occur against an unregistered source, so this step establishes the connection and scope needed before a scan rule set can be applied to classify sensitive data.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.