Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

An AI platform team is deploying a large language model for internal document summarization. Legal requires that no prompt or document content leaves the company's virtual private cloud, and the security team wants to control the exact model weights and runtime version. The team already has GPU capacity reserved in their own VPC. Which deployment approach best satisfies these constraints?

⚠ Common exam trap

The trap here is treating a private network endpoint or customer-managed encryption key as equivalent to keeping data and model execution inside your own VPC.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Self-host the model weights on the reserved GPUs inside the VPC and expose an internal inference endpoint

The constraints combine data residency, model weight control, and runtime version control, and the team already owns reserved GPU capacity. Running the weights on that capacity inside the VPC keeps all payloads private and lets the team pin both the weight revision and the serving runtime. Managed APIs and serverless services, even with private endpoints or encryption, still process content on provider infrastructure.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a provider's managed private endpoint with a business associate agreement

    Why it's wrong here

    A managed private endpoint creates a private network path to the provider's service, but the provider still hosts and runs the model on its own infrastructure, so document content is processed outside the company VPC. A business associate agreement is a contractual control for regulated data, not a technical guarantee that payloads never leave the VPC, and it does not address weight or runtime control.

  • ✓

    Self-host the model weights on the reserved GPUs inside the VPC and expose an internal inference endpoint

    Why this is correct

    Self-hosting the weights on reserved GPUs keeps all prompt and document data inside the company VPC, satisfying the legal data-residency constraint. The team also controls the exact weight revision and runtime version, which meets the security requirement. Because GPU capacity is already reserved, the incremental cost is operational rather than a new capital outlay, making this approach both compliant and practical.

  • ✗

    Deploy the model to a serverless inference service in a different cloud region

    Why it's wrong here

    A serverless inference service in another region moves both the compute and the data outside the company VPC, directly violating the residency constraint. Serverless platforms also abstract away the runtime, so the team cannot pin the exact model weight revision or runtime version. Region separation adds cross-region data transfer and does not restore the required control or isolation.

  • ✗

    Call a public foundation model API using customer-managed encryption keys

    Why it's wrong here

    A public model API still transmits prompt and document content to the provider's infrastructure, which violates the requirement that no content leaves the company VPC. Customer-managed encryption keys protect data at rest in the provider's environment but do not prevent the provider from processing plaintext payloads during inference. This approach also gives no control over model weights or runtime version.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.