Prompt caching reuses the processed prefix of a request. Cached input tokens cost less and return faster.

Mark the stable prefix

Put cache_control on the last block of the part that stays the same. Everything before the marker is cached.

Jibu.messages(client, %{
  model: "claude-opus-5-5",
  max_tokens: 1024,
  system: [
    %{type: "text", text: instructions},
    %{type: "text", text: project_list, cache_control: %{type: "ephemeral"}}
  ],
  messages: [%{role: "user", content: todays_events}]
})

The order of a request is tools, then system, then messages. A change anywhere in the prefix invalidates the cache from that point on. Keep timestamps and per-request data after the marker.

Check that it works

response.usage reports the cached part.

response.usage["cache_creation_input_tokens"]
response.usage["cache_read_input_tokens"]

The first request writes the cache. Later requests with the same prefix read it. If cache_read_input_tokens stays at zero, the prefix changes between requests. A prefix below the model's minimum length is not cached.

The telemetry event [:jibu, :request, :stop] carries the same usage map. See Errors, retries and telemetry.