Amazon Bedrock now lets developers cache parts of a prompt, shrinking token-usage bills by as much as 90 % and cutting response latency up to 85 % for applications that reuse a static prompt prefix.

Why the change matters

Running large language models on demand costs money every time tokens are sent to the model. Chatbots, code assistants, and document-search tools often resend the same system instructions or reference material, inflating spend and slowing responses.

How prompt caching works

Bedrock adds a “cacheable” flag that developers attach to any segment of a prompt—typically system-level instructions, long background documents, or tool definitions that never change during a session. When a request arrives, Bedrock checks whether the flagged segment matches a stored entry. If it does, the service skips re-encoding and re-running that part through the model and pulls the pre-computed representation from the cache instead.

The numbers

  • Input-token cost: up to 90 % reduction because the cached prefix no longer consumes tokens on each call.
  • Latency: up to 85 % faster since the model’s heavy lifting is avoided for the static portion.

Best-fit scenarios

The feature shines when the prompt contains a large, unchanging block followed by a short, variable user query. Common patterns include:

  • Retrieval-augmented generation (RAG) pipelines that prepend a fetched document to every query.
  • Customer-support bots that always start with the same policy statement or tone-setting text.
  • Coding assistants that load a fixed language-tool definition before the developer’s snippet.

What developers need to change

Developers must reorder the prompt so the static content sits at the very beginning and stays byte-for-byte identical across calls. The variable user input follows the cached prefix. No model switch is required; the same Bedrock endpoints handle the request.

Who gains, who watches out

The upside applies only to workloads where the prefix truly stays static. Applications that personalize system instructions per user or frequently alter the context will see little benefit and must weigh the added prompting complexity against modest gains.

Bottom line: Prompt caching gives Bedrock users a straightforward lever to trim AI operating costs and improve response times, provided their applications can isolate a reusable prompt prefix.