An AI engineer is tuning a large language model for a summarization task. The output summaries are too verbose and include irrelevant details. Which technique should be applied to encourage concise outputs?
Few-shot prompting supplies in-context demonstrations that steer the model's output distribution toward the desired style. By including concise summary exemplars in the prompt, the model infers the expected length and level of detail, directly countering the verbosity and irrelevant content described in the stem without retraining.
Why this answer
Providing a few-shot example with concise summaries (Option A) directly demonstrates the desired output format to the model, leveraging in-context learning to bias generation toward brevity and relevance. This is the most effective technique for controlling output style without altering the model's underlying parameters.
Exam trap
CompTIA often tests the misconception that adjusting sampling parameters (top-k, temperature) is the primary way to control output length, when in fact these parameters affect randomness and diversity, not the explicit length or relevance of the generated text.
How to eliminate wrong answers
Option B is wrong because chain-of-thought prompting encourages step-by-step reasoning, which typically increases verbosity and is designed for complex reasoning tasks, not for reducing output length. Option C is wrong because decreasing the top-k value restricts the sampling pool to the k most likely tokens, which can reduce randomness but does not inherently enforce conciseness or relevance; it may even produce repetitive or incomplete summaries. Option D is wrong because increasing the temperature raises the randomness of token selection, often leading to more diverse but also more verbose and irrelevant outputs, the opposite of the desired effect.