Cost-Effective Storage for Cold Analytics Data on Azure
You are designing a storage strategy for a data analytics solution that processes large volumes of streaming data. The data must be stored in a cost-effective manner with low latency for hot data and infrequent access for cold data after 30 days. The solution must support both batch and interactive queries. Which combination of Azure storage services should you recommend?
⚠ Common exam trap
A common mix-up: candidates confuse Azure Blob Storage with ADLS Gen2, overlooking that the hierarchical namespace is essential for analytics workloads, and that lifecycle management alone on standard Blob Storage does not enable the same query performance or directory-level operations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Azure Data Lake Storage Gen2 with lifecycle management
Azure Data Lake Storage Gen2 (ADLS Gen2) combines the scalability and cost benefits of object storage with a hierarchical namespace, enabling both batch and interactive queries via services like Azure Synapse Analytics and Apache Spark. Lifecycle management policies can automatically transition hot data to cooler tiers (e.g., cool or archive) after 30 days, reducing costs for infrequently accessed cold data while maintaining low-latency access for hot data. This makes ADLS Gen2 the ideal choice for streaming data analytics that requires cost-effective tiered storage and supports diverse query patterns.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Azure Data Lake Storage Gen2 with lifecycle management
Why this is correct
Azure Data Lake Storage Gen2 combines a hierarchical namespace with Azure Blob Storage's durable, scalable storage, giving analytics engines like Synapse, Databricks, and HDInsight native directory structures and POSIX-style access control. Lifecycle management policies can automatically transition data from hot to cool to archive tiers based on age or usage, slashing storage costs while keeping the data queryable for batch and interactive analytics. This makes it purpose-built for high-volume analytics pipelines where both performance and cost governance matter.
- ✗
Azure SQL Database with geo-replication
Why it's wrong here
Azure SQL Database is an online transaction processing (OLTP) relational store that maintains B-tree indexes and row-based storage optimized for single-row lookups and small writes, not for scanning or aggregating massive streaming datasets. Geo-replication duplicates the database to another region for disaster recovery, but it does nothing to make the platform suitable for large-scale analytical workloads or to reduce the cost of storing high-volume raw data. Its fixed schemas and limited storage scalability also clash with the schema-on-read nature of modern data analytics.
- ✗
Azure Blob Storage with hot and cool access tiers
Why it's wrong here
Azure Blob Storage with hot and cool access tiers is a generic object store that supports analytics, but it lacks the hierarchical namespace and atomic directory operations found in Data Lake Storage Gen2. Without a true file-system abstraction, analytics engines must emulate folders with prefixes, complicating partition pruning, rename, and merged-read patterns, and exacerbating costs for high throughput jobs. The access tiers only optimize for how often data is read, not for the file-layout and metadata requirements of columnar engines like Spark or Presto, making it a weaker fit than ADLS Gen2.
- ✗
Azure Cosmos DB with multiple consistency levels
Why it's wrong here
Azure Cosmos DB is a multi-model NoSQL service engineered for low-latency transactional access, distributing data via logical partitions and indexing every property for efficient point queries and small-range reads. Multiple consistency levels control the trade-off between latency and staleness, but they do not address how data is physically organized for analytical scanning; running large aggregations would consume excessive request units against a high RU/s budget. Analytics workloads need columnar storage, data skipping, and write-optimized formats like Parquet, which Cosmos DB does not provide.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
This AZ-305 question is part of Courseiva's 795-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AZ-305 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AZ-305 exam.