AI0-001 AI Security Practice Question
A company is training a model on proprietary data and wants to prevent data poisoning. Which TWO practices are most important? (Select TWO.)
⚠ Common exam trap
The AI0-001 exam often tests the distinction between security controls that prevent attacks (access controls, integrity validation) versus performance tuning (model size, epochs) or privacy techniques (homomorphic encryption), leading candidates to confuse data poisoning prevention with unrelated optimizations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implementing access controls on the training dataset
Option A (Implementing access controls on the training dataset) is correct because data poisoning requires an adversary to inject or modify training samples, and strict authentication/authorization (e.g., IAM roles, least-privilege permissions on the S3 bucket or data lake) prevents unauthorized parties from tampering with the proprietary dataset in the first place. Option B (Validating the integrity of training data) is correct because even with access controls, data can be corrupted or subtly altered, so integrity checks such as cryptographic hashes/checksums, provenance tracking, and outlier or anomaly detection help detect poisoned or tampered samples before they influence the model. Option C (Using a larger model) does not belong because model capacity has no bearing on whether poisoned data enters the pipeline and can even make a model more susceptible to memorizing malicious samples. Option D (Increasing training epochs) does not belong because more training iterations only reinforce whatever data is present, potentially amplifying the effect of poisoned samples rather than preventing them. Option E (Using homomorphic encryption) does not belong because it protects data confidentiality during computation, not the authenticity or integrity of the training data against poisoning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Implementing access controls on the training dataset
Why this is correct
Access controls restrict who can write to or modify the training dataset, preventing unauthorised actors from injecting malicious samples. This directly addresses the data poisoning threat by limiting the attack surface to trusted contributors, satisfying the stem's requirement to protect proprietary training data.
- ✓
Validating the integrity of training data
Why this is correct
Integrity validation detects tampered, corrupted or anomalous samples before training, ensuring poisoned data cannot influence model weights. This satisfies the stem's data poisoning prevention goal by verifying that the dataset matches its expected, trusted state prior to use.
- ✗
Using a larger model
Why it's wrong here
Model size governs capacity and representational power, not training-data integrity. A larger network can still learn poisoned patterns, sometimes more readily. Scaling parameters is chosen when underfitting complex data. Poisoning prevention depends on vetting data sources and detecting anomalous or tampered samples before training begins.
- ✗
Increasing training epochs
Why it's wrong here
More epochs simply repeat exposure to the same training data, amplifying any poisoned samples already present rather than filtering them. Epoch count is a convergence and overfitting control, useful when tuning model fit. Preventing poisoning requires validating and sanitising training data at ingestion, plus monitoring for anomalous samples.
- ✗
Using homomorphic encryption
Why it's wrong here
Homomorphic encryption lets computation occur on encrypted data, protecting confidentiality during processing; it does not detect or remove poisoned samples. It suits privacy-preserving inference or outsourced computation. Poisoning defence instead needs provenance checks and integrity validation of the training set before the model consumes it.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.