Pre-launchAPI access is by invitation. Prices shown are estimates.Request access →
Contact sales
ModelsPricingDedicatedDocsEnterpriseContact sales
Inference

Docs / Inference

Prompt Caching

Pay less and get faster first tokens when prompts share a prefix.

View .md

When consecutive requests start with the same prefix — a long system prompt, a document, tool definitions, or earlier turns of a conversation — Kurrens can reuse the model state computed for that prefix. Cached input tokens are billed at a lower rate and reduce time to first token.

How it works

Caching is automatic for models that support it; you don’t send extra parameters. The response’s usage.prompt_tokens_details.cached_tokens shows how many input tokens were served from cache.

Getting cache hits

  • Put stable content first (system prompt, tools, documents) and variable content last.
  • Keep the prefix byte-for-byte identical — even whitespace or reordering tools breaks the match.
  • Cached state expires after a short idle period, typically minutes.

Privacy

Cached state lives in GPU or system memory, is scoped to your account, is never written to persistent storage, and is evicted automatically. It is covered by our Zero data retention commitment.

    Type to search titles, headings, and page text.

    ↑↓ to move · ↵ to open