A team is building a retrieval-augmented assistant that sends a large document set inside the system prompt on every turn. To reduce cost, they enable prompt caching. They notice cache_read_input_tokens is high on most turns but intermittently drops to zero even though the system prompt text has not changed. Which explanation best fits this behavior?
Prompt caching uses a time-to-live window; if no request references the cached prefix before it expires, the entry is evicted. The next request then misses the cache, shows zero cache reads, and re-creates the entry, which matches the intermittent pattern described. Steady traffic keeps the cache warm, while idle gaps cause the drop to zero.
Why this answer
Prompt caching entries expire after a time-to-live if they are not referenced. Bursty or idle traffic patterns cause the prefix to be evicted between requests, so the next call misses the cache and reports zero cache reads before re-creating the entry. Keeping steady traffic, or reducing gaps between calls, stabilizes cache_read_input_tokens and preserves the intended cost savings.
Exam trap
The trap here is assuming cache misses must be caused by changing content or a broken marker, when the real cause is often the cache entry expiring during an idle gap.