Two thousand agents, one tool list
Your company gateway serves 2,000 employees' agents, and each calls tools/list at the start of every conversation. That's thousands of identical calls a minute hitting your servers.
With ttlMs: 300000 and cacheScope: "public", the gateway answers from cache. Your server sees one call every five minutes. Because the order is deterministic, the model's prompt prefix stays the same, so LLM prompt caching kicks in too.
The cautionary tale: a developer marks resources/read of app://me/payroll as "public". The gateway caches Alice's salary and serves it to Bob. cacheScope is a promise about the data, and the spec says plainly that it is not an access control.