A marketing company collects real-time clickstream data from their website using Azure Event Hubs. They need to perform two tasks: (1) aggregate the number of clicks per advertising campaign every 5 minutes and display the results in a live dashboard, and (2) run complex historical queries on months of aggregated click data to identify trends. They want to minimize data movement and use serverless compute where possible. Which combination of Azure services should they use?
This is correct because Azure Stream Analytics is a fully managed, serverless stream-processing engine that can run live aggregations—like 5-minute tumbling windows—over clickstream events and push results directly to Power BI for a real-time dashboard. For historical analysis, Azure Synapse Analytics serverless SQL pool can query Parquet files in the data lake without provisioning dedicated compute, enabling on-demand T-SQL queries over the same raw clickstream data. This combination cleanly separates the streaming path from the batch/historical path, which is exactly what the scenario requires.
Why this answer
Azure Stream Analytics is ideal for real-time aggregation of clickstream data from Event Hubs, outputting to Power BI for a live dashboard. Azure Synapse Analytics serverless SQL pool allows querying months of aggregated data stored in Azure Data Lake Storage without provisioning compute, minimizing data movement and using serverless compute.
Exam trap
The trap here is confusing batch processing tools like Azure Data Factory or HDInsight with real-time stream processing, and overlooking that Azure Synapse serverless SQL pool is the serverless option for historical queries, not Azure SQL Database.
Why the other options are wrong
Azure Data Factory is an orchestration and data movement service, not a real-time stream processing engine, so it cannot perform live aggregation of clickstream data. Azure Analysis Services is an OLAP engine for semantic models, not a serverless SQL query service for historical data, and it requires data to be moved into its own store.
HDInsight (Spark) is not serverless and requires cluster management, contradicting the requirement to minimize data movement and use serverless compute. Additionally, it is overkill for simple 5-minute aggregations and live dashboards compared to Stream Analytics.
Azure Functions is not designed for real-time stream aggregation at scale; it lacks native windowing and state management for 5-minute tumbling windows. Azure SQL Database is not serverless and requires manual scaling, increasing data movement for historical queries.
When would these options actually be correct?
A company needs to orchestrate and move data from on-premises SQL Server to Azure Blob Storage on a nightly schedule, and then provide interactive analytics on that data using a tabular model. Azure Data Factory would handle the scheduled data movement, and Azure Analysis Services would host the semantic model for fast, interactive queries.
An exam scenario where the company needs to perform complex, custom machine learning on streaming data (e.g., real-time anomaly detection) and batch processing on large historical datasets, and is willing to manage clusters for full control over the processing environment.
A company needs to process individual click events with custom business logic (e.g., enrichment or transformation) in near real-time, and store results in a relational database for simple historical lookups. The question would emphasize low-latency, event-driven processing and a small data volume where serverless compute is not a priority.
Why candidates pick the wrong answer
Candidates may confuse Data Factory's data movement capabilities with real-time processing, and think Analysis Services is suitable for historical queries because it supports analytics, overlooking the need for serverless SQL querying and minimal data movement.
Candidates may think Spark is a one-size-fits-all solution for both streaming and batch, and overlook the serverless and minimal-management requirements of the question.
Candidates may associate Azure Functions with 'serverless compute' and assume it can handle real-time aggregation, overlooking its lack of built-in stream processing features. Azure SQL Database is a familiar choice for historical data, but they miss the 'minimize data movement' and 'serverless' requirements.