An engineer is training a model on Vertex AI using a custom container. The training job fails with an error indicating that the container exited with a non-zero status. The engineer wants to debug the issue. What is the best way to access the logs?
Vertex AI writes custom container training output, including stderr and the non-zero exit trace, to Cloud Logging. Logs Explorer surfaces those entries for the failed job, satisfying the debugging requirement by exposing the container's actual error rather than only the job's status.
Why this answer
Vertex AI automatically streams all container stdout and stderr to Cloud Logging (Logs Explorer). When a custom container exits with a non-zero status, the detailed error messages, stack traces, and application logs are captured there, making it the primary and most comprehensive debugging tool. Cloud Logging provides structured, searchable logs without requiring direct access to the container.
Exam trap
A common trap is the misconception that you can SSH into a training container or that logs are stored in Cloud Storage, when in fact Cloud Logging is the centralized, default logging solution for all Vertex AI training jobs.
How to eliminate wrong answers
Option A is wrong because Vertex AI does not provide an SSH feature for training containers; training jobs run in ephemeral, isolated environments with no interactive shell access. Option B is wrong because Cloud Storage under the job's output directory stores artifacts like model checkpoints and metrics, not real-time container logs or error messages. Option C is wrong because Cloud Debugger is designed for debugging running applications in production by capturing snapshots and variable states, not for inspecting container exit errors or retrieving logs from a failed training job.