Handling Bursts Cost-Effectively with Azure Queue Storage & Throttling
You are architecting an Azure AI solution that uses Azure AI Language to analyze text for sentiment and key phrases. The solution must handle bursts of up to 500 requests per second but average only 50 requests per second. You need to ensure cost efficiency while meeting performance requirements. Which THREE actions should you take?
Quick Answer
The cost-efficiency problem here is that provisioning Azure AI Language to sustain 500 requests per second around the clock, to cover a burst that only actually happens occasionally against a 50-per-second average, means paying for capacity that sits mostly idle. Azure Queue Storage solves this by decoupling ingestion from processing: incoming requests land in the queue the moment they arrive, absorbing the full burst instantly and cheaply since queue storage is inexpensive compared to AI service compute, while the Language service itself keeps consuming from that queue at a steady, lower rate it can sustain without needing burst-level capacity provisioned. This is the general pattern for reconciling a spiky arrival rate with a service that's expensive to over-provision — insert a low-cost buffer between the spike and the processing tier so the processing tier only ever needs to be sized for sustained throughput, not peak throughput. Client-side throttling and retry logic complements this by handling the case where the queue itself is catching up, gracefully backing off rather than hammering the service with duplicate requests. Any scenario describing a mismatch between burst capacity and average load, where over-provisioning for the peak would be wasteful, is describing this decouple-with-a-queue pattern.
⚠ Common exam trap
Candidates often assume a higher pricing tier (like S0) alone can handle bursts, but without queuing and retry logic, the service will still throttle requests, leading to failures or the need for costly over-provisioning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Azure Queue Storage to buffer requests during spikes
Option A is correct because Azure Queue Storage decouples the request producers from the Azure AI Language service, buffering the burst of up to 500 requests per second so the service can process them at a sustainable rate rather than being overwhelmed. Option C is correct because the S0 (Standard) pricing tier is the paid tier that supports the throughput and quota needed for production workloads, whereas the Free tier is limited to 5,000 transactions per month at 20 transactions per minute, which cannot handle 50 requests per second on average. Option D is correct because client-side throttling and retry logic (for example, honoring HTTP 429 responses and using exponential backoff) prevents the application from exceeding the service's rate limits and gracefully handles transient throttling during spikes. Option B is not correct because the Free tier's low transaction and rate limits make it unsuitable for this workload even with queuing. Option E is not correct because deploying in multiple regions increases cost and complexity without addressing the burst-handling and cost-efficiency requirements, since the service's per-region rate limits would still apply.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Azure Queue Storage to buffer requests during spikes
Why this is correct
Queue Storage decouples ingestion from processing, absorbing the 500 requests-per-second bursts so the Language service is called at a sustainable rate. This smooths spikes and avoids provisioning for peak throughput, keeping costs aligned with the 50 requests-per-second average.
- ✗
Select the Free tier and implement queuing
Why it's wrong here
The Free tier caps transactions at 5,000 per month with roughly 20 transactions per minute, so it cannot absorb 500 requests per second even with queuing, and queuing adds latency rather than throughput. It is tempting because it costs nothing, and it would suit low-volume development or prototyping, not a production burst workload.
- ✓
Use the S0 pricing tier
Why this is correct
The S0 tier provides the higher transaction throughput and quota needed to absorb bursts up to 500 requests per second. The free tier's limits would throttle the workload, so S0 satisfies the performance constraint while remaining pay-as-you-go.
- ✓
Implement client-side throttling and retry logic
Why this is correct
Client-side throttling smooths the 500 rps bursts down toward the 50 rps average, keeping consumption within the provisioned tier's rate limit and avoiding 429 responses. Retry logic with exponential backoff then recovers any requests still rejected, so performance holds without paying for capacity sized to peak demand.
- ✗
Deploy the service in multiple regions
Why it's wrong here
Multi-region deployment duplicates the Azure AI Language resource, doubling provisioned capacity and cost without raising the per-resource transaction limit, and the workload is bursty rather than geographically distributed. It is tempting because regional redundancy aids availability and latency for globally dispersed users, but here the requirement is absorbing short spikes cheaply, which autoscale and batching address.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on AI-102
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A developer is configuring an Azure AI Language resource for sentiment analysis. The solution must process social media posts in real-time with a throughput of 1000 requests per minute. After testing, the developer notices that the API returns a 429 (Too Many Requests) error when the load exceeds 500 requests per minute. What is the most likely cause and solution?
easy- A.Scale out the resource by creating multiple Azure AI Language instances and load balancing requests.
- ✓ B.Upgrade the Azure AI Language resource to a higher tier (e.g., Standard S) to increase the rate limit.
- C.Implement retry logic with exponential backoff to handle 429 errors.
- D.Use Azure API Management to cache responses and reduce calls.
Why B: The 429 error indicates the request rate exceeds the resource's allocated tier limit. Azure AI Language resources have predefined rate limits per pricing tier; the Standard S tier offers higher throughput (e.g., 1,000 requests per minute) compared to lower tiers. Upgrading to Standard S directly increases the rate limit to match the required 1,000 requests per minute, making it the correct solution.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.