A data engineer runs a long-running aggregation query and inspects the Query Profile. The profile shows a single operator with an output row count roughly 400 times larger than its input row count, and the downstream operator is bottlenecked. Which action should the engineer take to resolve this?
An operator whose output rows massively exceed its input rows indicates an unintended row explosion, usually from a fan-out join or an incorrect grouping key. Because the explosion happens before the aggregation, the downstream operator must process every duplicated row, so fixing the cardinality growth is the only action that addresses the actual bottleneck.
Why this answer
The decisive clue is an operator emitting far more rows than it receives, which is the signature of a cardinality explosion rather than a compute or pruning problem. Adding warehouse capacity, toggling result caching, or enabling search optimization all leave that multiplication intact. Only restructuring the plan so the aggregation occurs before or alongside the fan-out keeps downstream operators from processing duplicated rows.
Exam trap
The trap here is assuming a large intermediate row count always means insufficient warehouse compute, when the profile is actually pointing to a join or grouping that multiplies rows.