hardMultiple Select
AIF-C01 Practice Question: A company uses a text generation model to produce…
A company uses a text generation model to produce legal documents. They want to minimize the environmental impact of training and inference. Which THREE approaches should they consider?
⚠ Common exam trap
AIF-C01 often tests the misconception that redundancy improves sustainability; candidates may choose multi-region storage thinking it's good, but it increases energy use.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Train the model on a smaller but representative dataset
Option A is correct because training on a smaller but representative dataset reduces the total compute cycles, energy consumption, and carbon emissions required for model training while still maintaining acceptable model quality. Option C is correct because distilled or otherwise more efficient architectures require fewer parameters and FLOPs per inference, directly lowering the energy and hardware resources needed for both training and serving. Option E is correct because SageMaker managed spot training uses spare AWS capacity and automatically checkpoints/resumes jobs, which improves utilization and reduces idle compute, thereby lowering the environmental footprint. Option B is not appropriate because replicating training data across multiple AWS Regions increases storage, replication traffic, and energy use without reducing the model's environmental impact. Option D is not appropriate because on-demand instances do not inherently reduce energy consumption and can leave resources idle; they address performance consistency, not environmental efficiency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Train the model on a smaller but representative dataset
Why this is correct
A smaller representative dataset shortens each training epoch and reduces total compute hours, lowering the energy consumed during training. This directly satisfies the stem's environmental-impact goal, provided the sample preserves the legal-domain patterns the model must learn.
- ✗
Store all training data in multiple AWS Regions for redundancy
Why it's wrong here
Replicating training data across multiple AWS Regions multiplies storage and cross-Region transfer, adding energy and embodied hardware cost without improving model quality. It is tempting because multi-Region redundancy protects against regional failure, but durability is not the objective — minimising the environmental impact of training and inference is.
- ✓
Use a more efficient model architecture like a distilled version of a larger model
Why this is correct
A distilled model has fewer parameters and lower computational demand per inference, cutting energy consumption during both training and inference. This directly satisfies the stem's goal of minimising environmental impact while still generating legal text of acceptable quality.
- ✗
Perform all training on on-demand instances to ensure consistent performance
Why it's wrong here
On-demand instances keep capacity running regardless of utilisation, so idle GPU hours still consume energy; reserved or spot capacity with right-sizing and scheduling reduces that waste. It is tempting because on-demand suits unpredictable training bursts, but consistent performance is not the goal here — minimising environmental impact is.
- ✓
Use Amazon SageMaker with managed spot training to reduce idle compute
Why this is correct
Managed spot training uses spare EC2 capacity and automatically resumes interrupted jobs from checkpoints, eliminating idle compute and wasted energy. This directly reduces the environmental impact of training, satisfying the stem's requirement to minimise resource consumption during model development.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.