If a GPU job consistently fails with 'Out of Memory' despite the model size being significantly smaller than the total VRAM, what is the most likely cause?
Memory fragmentation occurs when the available memory is split into small, non-contiguous blocks. When the model requests a large contiguous block of memory, the allocator fails to find one, resulting in an OOM error even if the sum of all free memory is actually greater than the request.
Why this answer
Memory fragmentation occurs when frequent allocations and deallocations leave holes in the memory address space. Even if total free memory seems sufficient, large contiguous blocks cannot be allocated. Monitoring fragmentation patterns is essential for AI engineers to optimize memory management, such as using memory pools or persistent buffers, to ensure stable and predictable training performance in long-running jobs.
Exam trap
Candidates frequently assume an OOM error always means total available memory is exhausted, overlooking how memory fragmentation prevents large contiguous allocations.