Courseiva

CCNA Troubleshooting Questions

11 of 86 questions · Page 2/2 · Troubleshooting topic · Answers revealed

76
MCQmedium

An AI operations engineer is troubleshooting an inference service on an NVIDIA A100 GPU that shows intermittent stalls. The monitoring dashboard reports GPU utilization at 100%, but request throughput is far below the validated baseline. Running nvidia-smi dmon reveals the SM utilization is high while memory controller utilization is low. Which action should the engineer take first to identify the bottleneck?

A.Switch the inference precision from FP16 to FP32 to improve numerical stability.
B.Enable MPS (Multi-Process Service) to allow concurrent kernel execution from multiple processes.
C.Profile the inference process with Nsight Systems to inspect kernel execution and identify serialization gaps.
D.Increase the batch size in the inference server configuration to improve GPU occupancy.
AnswerC

Nsight Systems captures a timeline of CPU and GPU activity, revealing kernel serialization, gaps, and synchronization stalls that inflate utilization without producing throughput. Because memory controller utilization is low, the bottleneck is likely execution serialization or CPU-side latency, which Nsight Systems can pinpoint. This is the correct first step to diagnose the cause before making configuration changes.

Why this answer

The high SM utilization with low memory controller utilization suggests the GPU is busy but not doing useful work, often due to kernel serialization or CPU-GPU synchronization stalls. Profiling with Nsight Systems provides the timeline needed to see gaps and serialization. Only after identifying the specific stall should configuration changes be considered, making profiling the correct first step.

Exam trap

The trap here is assuming that 100% GPU utilization always means the GPU is efficiently processing work, when it can indicate stalls or serialization.

77
MCQhard

An AI operations team is running a large language model inference service on NVIDIA H100 GPUs using NVIDIA Triton Inference Server. They observe that the first inference request after a period of inactivity takes significantly longer than subsequent requests. The model is loaded and ready, but the GPU shows low utilization during the first request. Which optimization should the team implement to reduce this latency spike?

A.Configure Triton's model warmup to run dummy inference requests during model loading.
B.Enable Triton's dynamic batching with a large maximum batch size.
C.Set the Triton `--pinned-memory-pool-byte-size` to a larger value.
D.Increase the number of model instances per GPU to allow more concurrent executions.
AnswerA

Triton's model warmup feature executes a specified number of inference requests when the model is loaded, ensuring that CUDA kernels are compiled, memory allocations are made, and the GPU is initialized. This eliminates the cold-start penalty for the first real request. Setting warmup with representative input shapes and batch sizes directly reduces the latency spike after periods of inactivity.

Why this answer

The first inference after inactivity is slow because CUDA kernels and memory allocations are not yet initialized on the GPU. Triton's model warmup runs dummy requests at load time to trigger this initialization, so the first real request executes at normal speed. Other options target throughput or memory pooling but do not eliminate the cold-start penalty.

Exam trap

The trap here is confusing cold-start latency with throughput optimization, leading to batching or instance scaling instead of pre-warming the model.

78
MCQhard

Refer to the exhibit. The system has two GPUs. What is the most likely cause of the observed performance discrepancy?

A.The GPUs have different amounts of physical memory.
B.Only one GPU is receiving the workload due to improper data distribution.
C.The system is limited by the PCIe bus bandwidth.
D.The model is too small to be parallelized across two GPUs.
AnswerB

The discrepancy between high GPU utilization on one device and low utilization on the other strongly suggests that only one GPU is performing the compute-intensive training loop. This is typical when the DataParallel or DistributedDataParallel wrapper is not correctly configured across all available devices.

Why this answer

The exhibit shows one GPU heavily utilized while the second is idling despite similar memory consumption. This indicates a data parallelism imbalance, where one process is doing the bulk of the work. This is a common issue in multi-GPU setups where workload distribution is not correctly handled, leading to massive inefficiencies where expensive hardware is under-utilized, significantly increasing the time required for model training or inference tasks.

Exam trap

Candidates often guess hardware failure or driver mismatch, when the most common issue is a simple failure to properly initialize or distribute workloads across both available GPUs in the application code.

79
MCQhard

A team is deploying a large language model for inference using NVIDIA TensorRT-LLM on an H100 GPU. They observe that the first inference request takes several seconds, while subsequent requests are fast. They want to reduce this initial latency. Which technique should they implement?

A.Increase the GPU's power limit to boost clock speeds during the first request.
B.Use TensorRT-LLM's built-in paged KV cache and enable continuous batching.
C.Precompile the TensorRT engine and load it at server startup, then perform a warm-up inference.
D.Reduce the model's precision to INT4 to decrease computation time.
AnswerC

The first request latency includes engine deserialization, CUDA context creation, and kernel loading. Precompiling the engine and loading it during startup, followed by a warm-up inference, ensures that these one-time costs are paid before actual requests arrive. This directly reduces the first-request latency for users.

Why this answer

The first inference request incurs one-time costs such as TensorRT engine deserialization, CUDA context setup, and kernel compilation/loading. By precompiling the engine and loading it at startup, and then running a warm-up inference, these costs are moved to server initialization. Subsequent requests then benefit from a fully initialized environment, reducing the observed initial latency.

Exam trap

The trap here is confusing steady-state optimizations like continuous batching or precision reduction with cold-start latency, which is caused by initialization overhead.

80
MCQhard

An administrator wants to ensure that a training process is limited to a single GPU on a multi-GPU node. Which environment variable should be set?

A.NCCL_DEBUG=INFO
B.CUDA_VISIBLE_DEVICES=0
C.NVIDIA_DRIVER_CAPABILITIES=compute
D.OMP_NUM_THREADS=1
AnswerB

CUDA_VISIBLE_DEVICES is the standard environment variable used to mask specific GPUs from a process. Setting it to a specific index restricts the application to use only that hardware device, which is the standard method for isolating jobs in a multi-GPU system.

Why this answer

Controlling GPU visibility is a fundamental skill for resource management in multi-tenant environments. By using CUDA_VISIBLE_DEVICES, an administrator can restrict a process to a specific device, preventing multiple jobs from competing for the same GPU. This isolation is crucial for maintaining performance stability and ensuring that individual jobs receive consistent, predictable access to compute resources without interference from other concurrent tasks.

Exam trap

Students often mistakenly select command-line flags or code-level device placement arguments instead of the standard operating system environment variable required to restrict GPU visibility globally.

81
MCQmedium

An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The training loop shows frequent "CPU bottleneck" warnings in the logs. Which action should the engineer take first to optimize throughput?

A.Upgrade to a higher-end GPU model to handle the processing load.
B.Increase the batch size significantly to fill the GPU memory.
C.Implement NVIDIA DALI to offload preprocessing tasks from the CPU to the GPU.
D.Reduce the number of training epochs to lower CPU overhead.
AnswerC

NVIDIA DALI is specifically designed to accelerate data preprocessing pipelines by moving them from the CPU to the GPU. This eliminates the bottleneck by ensuring that data augmentation and transformation tasks occur at the same high speed as the training process, maximizing overall system hardware utilization.

Why this answer

Identifying CPU bottlenecks is critical because data pipelines often struggle to keep up with GPU compute speed. By optimizing data preprocessing, specifically increasing the number of workers in the DataLoader or using NVIDIA DALI, the engineer ensures the GPU remains saturated with data. This optimization directly impacts total training time and infrastructure cost efficiency, ensuring that high-performance hardware is not left idling while waiting for I/O operations.

Exam trap

Candidates often suggest upgrading the GPU or increasing the batch size, which exacerbates the CPU bottleneck rather than solving the underlying data ingestion starvation occurring at the preprocessing layer.

82
Multi-Selecthard

An AI researcher is deploying a multi-node training job using NCCL on an InfiniBand network. The job is suffering from intermittent latency spikes. Which TWO steps should the engineer perform to troubleshoot the network configuration?

Select 2 answers
A.Check IB link status and error counters using 'ibstat'.
B.Reset the GPU thermal throttling threshold.
C.Update the OS kernel to the latest version.
D.Verify NCCL_IB_HCA and NCCL_IB_GID_INDEX settings.
E.Increase the system swap space to 512GB.
AnswersA, D

Monitoring InfiniBand link counters is essential for detecting physical layer issues like CRC errors or link retrains. If the physical link is unstable, NCCL collective operations will experience significant latency, causing training to stall or fail. Identifying these errors early isolates the issue to the fabric cabling or switches.

Why this answer

Troubleshooting high-performance networking in distributed training requires checking for physical layer stability and protocol-level misconfigurations. Verifying InfiniBand counters helps identify packet drops or link flapping, while ensuring NCCL environment variables are correctly set for the specific topology avoids suboptimal routing. These steps ensure that the inter-GPU communication remains within the high-bandwidth low-latency envelope required for scaling deep learning workloads effectively.

Exam trap

Candidates often try to debug the training application code or model synchronization logic, failing to check the physical InfiniBand layer and NCCL environment variables that govern high-speed network communication.

83
MCQmedium

If a GPU job consistently fails with 'Out of Memory' despite the model size being significantly smaller than the total VRAM, what is the most likely cause?

A.The GPU clock speed is too low.
B.Memory fragmentation in the CUDA context.
C.An incompatible version of NCCL library.
D.The system bus width is insufficient.
AnswerB

Memory fragmentation occurs when the available memory is split into small, non-contiguous blocks. When the model requests a large contiguous block of memory, the allocator fails to find one, resulting in an OOM error even if the sum of all free memory is actually greater than the request.

Why this answer

Memory fragmentation occurs when frequent allocations and deallocations leave holes in the memory address space. Even if total free memory seems sufficient, large contiguous blocks cannot be allocated. Monitoring fragmentation patterns is essential for AI engineers to optimize memory management, such as using memory pools or persistent buffers, to ensure stable and predictable training performance in long-running jobs.

Exam trap

Candidates frequently assume an OOM error always means total available memory is exhausted, overlooking how memory fragmentation prevents large contiguous allocations.

84
MCQeasy

An operations engineer notices that an inference container on an A100 is intermittently returning stale predictions after a model update. The container mounts the model directory from a host path, and the update process replaces files in place. Which change most reliably prevents the stale predictions?

A.Increase the container's shared memory size so the model can be cached entirely in RAM.
B.Reduce the inference batch size so each prediction reads the model from disk again.
C.Enable read-only mounts for the model directory to block writes from the container.
D.Use a versioned, immutable model artifact and restart or roll the inference container so it loads the new version atomically.
AnswerD

Replacing files in place while a process has them mapped or cached can leave the running server serving old weights or a mix of old and new tensors. Publishing each model as an immutable, versioned artifact and restarting the server ensures the process loads a consistent snapshot. This removes the race between file replacement and model loading, eliminating stale or partially updated predictions.

Why this answer

When model files are overwritten in place, a running inference process keeps using the weights it already loaded, so clients see stale results until the process restarts. Treating each model as an immutable, versioned artifact and performing a container restart or rolling deployment guarantees an atomic switch to a consistent set of weights. This removes the timing race between the update and the load, which is the actual source of the stale predictions.

Exam trap

The trap here is assuming that read-only mounts or larger caches make model updates safe, when the real issue is that in-place file replacement never notifies a process that already loaded the weights.

85
MCQhard

Refer to the exhibit. The training job is showing intermittent "thermal throttling" warnings. Which configuration change is the most appropriate adjustment?

A.Set "compute_mode" to "exclusive_process".
B.Disable ECC mode to lower the power consumption.
C.Lower the "power_limit_watts" value.
D.Set "persistence" to "disabled".
AnswerC

Reducing the power limit directly lowers the heat generated by the GPU. While this may slightly decrease the maximum performance, it prevents the GPU from reaching the thermal trip point, thereby eliminating throttling and ensuring a consistent, albeit slightly lower, performance profile during long-duration training jobs.

Why this answer

Thermal throttling occurs when the GPU reaches its maximum operating temperature and lowers its clock speed to prevent physical damage. While power limits can be adjusted to reduce heat, simply lowering the limit may negatively impact training performance. Adjusting the power limit to a value that balances thermal overhead with compute requirements is a necessary trade-off to ensure stable, consistent performance during long training runs without hardware damage.

Exam trap

Test-takers frequently assume that lowering the power limit will eliminate thermal throttling without side effects, forgetting that overly restricted wattage can severely degrade training compute performance and extend runtime.

86
MCQmedium

A distributed training job using PyTorch DDP across eight GPUs on one DGX A100 node shows GPU utilization oscillating between 20 and 40 percent, while `nvidia-smi dmon` shows low SM activity but sustained high memory-controller utilization. The data loader reads from a local NVMe RAID array and applies heavy CPU augmentation. Which action best improves GPU utilization?

A.Enable NCCL gradient compression to reduce inter-GPU communication volume.
B.Reduce the per-GPU batch size so each step completes faster and utilization appears smoother.
C.Increase the number of DataLoader worker processes and enable pinned memory with non-blocking host-to-device transfers.
D.Switch the job from DDP to a single-process data-parallel loop that iterates over all GPUs in Python.
AnswerC

Low SM activity combined with heavy memory-controller traffic points to input pipeline stalls: the GPUs are idle waiting for batches. More worker processes parallelize CPU augmentation, while pinning host buffers and using non-blocking copies lets the copy engine overlap transfer with compute. This directly shortens the gap between kernel bursts and raises sustained SM occupancy without changing model math.

Why this answer

Oscillating SM utilization with high memory-controller traffic is the classic signature of an input pipeline that cannot keep pace with compute. Adding DataLoader workers parallelizes augmentation across CPU cores, and pinned memory plus non-blocking transfers let DMA copies overlap with kernels. This keeps the GPUs fed continuously, converting idle gaps into productive compute and raising sustained utilization without altering the model or optimizer.

Exam trap

The trap here is blaming inter-GPU communication because the job is distributed, when the profile shows idle SMs rather than busy NVLink, which points to starvation from the host input pipeline instead.

← PreviousPage 2 of 2 · 86 questions total

Ready to test yourself?

Try a timed practice session using only Troubleshooting questions.