A financial services firm uses a fine-tuned Gemini model in Vertex AI for regulatory compliance checks. They notice that token usage is high, increasing costs. They want to reduce costs without sacrificing accuracy. Which approach should they take?
Trap 1: Switch to a smaller base model like PaLM 2 Bison
May reduce accuracy for compliance tasks.
Trap 2: Enable context caching to reuse previous responses
Caching saves on repeated prompts but doesn't reduce per-request token usage.
Trap 3: Reduce temperature to 0.0
Affects randomness, not token count.
- A
Switch to a smaller base model like PaLM 2 Bison
Why it fails: May reduce accuracy for compliance tasks.
- B
Enable context caching to reuse previous responses
Why it fails: Caching saves on repeated prompts but doesn't reduce per-request token usage.
- C
Set max output tokens to a lower value and use more precise prompts
Directly reduces output tokens; precise prompts maintain accuracy.
- D
Reduce temperature to 0.0
Why it fails: Affects randomness, not token count.