AI0-001 AI Infrastructure and Technologies Practice Question
A hospital wants to run a natural language processing model that summarizes clinical notes. Because of patient privacy regulations, the data cannot leave the hospital's on-premises network, and there is no dedicated GPU available. Which deployment approach best fits these constraints?
⚠ Common exam trap
The trap here is treating de-identification as equivalent to keeping data on-premises, when the stated policy forbids any external transmission of clinical notes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy a quantized open-source model on an on-premises CPU server using a runtime such as ONNX Runtime or llama.cpp.
Running a quantized open-source model locally with a CPU-optimized runtime keeps all clinical data inside the hospital network and avoids GPU requirements. Quantization shrinks model size and speeds inference enough for summarization on standard servers. Cloud APIs, remote hosting, and full-precision large-model training all violate either the privacy boundary or the hardware constraint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Host the model in a colocation facility that is physically separate but connected by a dedicated VPN.
Why it's wrong here
A colocation facility is outside the hospital's on-premises network, so clinical note data would still traverse and reside outside the required boundary. A VPN secures transit but does not satisfy a strict data-residency rule. This option relocates the problem rather than solving it.
- ✗
Fine-tune a large foundation model on the hospital's CPU servers and serve it with full FP32 precision.
Why it's wrong here
Fine-tuning and serving a large foundation model in full FP32 on CPUs is impractical due to memory and compute requirements; inference latency would be unacceptable for clinical use. It also does not address the privacy constraint any better than running a smaller local model. The approach is technically infeasible for the stated environment.
- ✓
Deploy a quantized open-source model on an on-premises CPU server using a runtime such as ONNX Runtime or llama.cpp.
Why this is correct
Quantized open-source models can run entirely on-premises on CPU using runtimes like ONNX Runtime or llama.cpp, keeping patient data inside the network and requiring no GPU. Quantization reduces memory and compute demands enough for CPU inference on summarization tasks. This satisfies both the privacy constraint and the hardware limitation.
- ✗
Use a cloud-based LLM API and send de-identified notes for summarization.
Why it's wrong here
Even with de-identification, sending clinical notes to an external API conflicts with the requirement that data cannot leave the on-premises network. De-identification also carries re-identification risk that regulators scrutinize. This option violates the explicit constraint and is therefore unsuitable regardless of its convenience.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.