CCAR-P · practice · Claude Models, Prompting & Context Engineering
A team enabled prompt caching on a high-volume endpoint two weeks ago and the expected cost reduction has not appeared. Inspecting responses shows `cache_creation_input_tokens` is consistently non-zero and `cache_read_input_tokens` is consistently zero. What is the most likely cause?
Select one