# Prompt Caching

> Pay less and get faster first tokens when prompts share a prefix.

When consecutive requests start with the same prefix — a long system prompt, a document, tool definitions, or earlier
turns of a conversation — Kurrens can reuse the model state computed for that prefix. Cached input tokens are billed at a
lower rate and reduce time to first token.

## How it works

Caching is **automatic** for models that support it; you don't send extra parameters. The response's
`usage.prompt_tokens_details.cached_tokens` shows how many input tokens were served from cache.

## Getting cache hits

- Put stable content first (system prompt, tools, documents) and variable content last.
- Keep the prefix byte-for-byte identical — even whitespace or reordering tools breaks the match.
- Cached state expires after a short idle period, typically minutes.

## Privacy

Cached state lives in GPU or system memory, is scoped to your account, is never written to persistent storage, and is
evicted automatically. It is covered by our [Zero data retention](/docs/data-security/zero-data-retention) commitment.
