Pre-launchAPI access is by invitation. Prices shown are estimates.Request access →
Contact sales
ModelsPricingDedicatedDocsEnterpriseContact sales
Inference

Docs / Inference

Reasoning

Use models that think before answering, and control reasoning effort.

View .md

Models marked Reasoning produce internal reasoning before the final answer.

completion = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4.1-Flash",
    messages=[{"role": "user", "content": "How many primes are below 100?"}],
    reasoning_effort="medium",
)
message = completion.choices[0].message
print(message.content)

What to know

  • Billing. Reasoning tokens are billed at the output rate and reported in usage.completion_tokens_details.reasoning_tokens.
  • max_tokens includes reasoning. Leave enough headroom, or the answer may be cut off before it starts.
  • Effort. Where supported, reasoning_effort (low, medium, high) trades latency and cost for depth.
  • Reasoning content. Some models return their reasoning in a reasoning_content field alongside content. It is processed like any other output — not stored.
  • Keep-alives. While a model reasons, streaming responses may include SSE comment lines; ignore them.

    Type to search titles, headings, and page text.

    ↑↓ to move · ↵ to open