AIF-C01 Fundamentals of Generative AI Practice Question
A research team is using Amazon SageMaker to fine-tune a large language model. They want to optimize training cost and time without sacrificing model quality. Which THREE strategies should they implement? (Choose 3)
⚠ Common exam trap
The AIF-C01 exam often tests the misconception that simply scaling up hardware (larger instances) or maximizing batch size is the best optimization strategy, when in fact algorithmic efficiency (PEFT, mixed precision) and cost-saving infrastructure (spot instances) are the correct approaches for balancing cost, time, and quality.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply parameter-efficient fine-tuning (PEFT) techniques like LoRA.
Option B is correct because parameter-efficient fine-tuning techniques such as LoRA freeze most of the pretrained model weights and train only small low-rank adapter matrices, which drastically reduces GPU memory consumption, training time, and compute cost while preserving model quality. Option D is correct because SageMaker managed spot training can cut training costs by up to 90% versus on-demand instances, and pairing it with checkpointing to Amazon S3 lets training resume from the last checkpoint after a spot interruption rather than restarting from scratch. Option E is correct because mixed precision training with FP16 reduces memory footprint and leverages GPU tensor cores for faster matrix math, speeding up training and allowing larger effective batch sizes with minimal impact on model accuracy. Option A is not correct because simply moving to a larger multi-GPU instance raises cost and does not by itself optimize time or cost efficiency, and may be unnecessary if PEFT and FP16 already fit the workload. Option C is not correct because maximizing batch size to the GPU memory limit is not universally beneficial; overly large batches can hurt convergence and model quality and may require extensive hyperparameter retuning, so it is not a reliable cost-and-time optimization strategy.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger instance type with more GPUs.
Why it's wrong here
Scaling to more GPUs raises hourly compute cost and does not by itself shorten training time proportionally, so it fails the cost-optimisation requirement. It is tempting because larger instances genuinely help when a model or batch cannot fit in a single GPU's memory, but distributed or spot strategies address this scenario's cost and time goals.
- ✓
Apply parameter-efficient fine-tuning (PEFT) techniques like LoRA.
Why this is correct
Parameter-efficient fine-tuning such as LoRA freezes the base model weights and trains only small low-rank adapter matrices, cutting trainable parameters and GPU memory dramatically. This directly satisfies the stem's constraint of reducing training cost and time while preserving model quality, since the pretrained knowledge remains intact.
- ✗
Increase the batch size to the maximum that fits in GPU memory.
Why it's wrong here
Pushing batch size to the GPU memory limit can reduce convergence quality and force gradient-accumulation workarounds, so it risks the no-quality-loss requirement. It is tempting because larger batches improve throughput and GPU utilisation, but the correct approach tunes batch size alongside learning-rate scaling rather than maximising it blindly.
- ✓
Use managed spot training with checkpointing.
Why this is correct
Managed spot training taps spare AWS capacity at up to 90% lower cost, directly meeting the cost-optimisation constraint. Checkpointing periodically saves model state to Amazon S3, so interrupted jobs resume rather than restart, preserving training progress and quality.
- ✓
Enable mixed precision training (FP16).
Why this is correct
Mixed precision uses FP16 for most computations and FP32 where needed, halving memory bandwidth and enabling tensor-core throughput. This cuts training time and GPU memory use while retaining convergence, satisfying the speed constraint without degrading model quality.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.