Models marked Reasoning produce internal reasoning before the final answer.
completion = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V4.1-Flash",
messages=[{"role": "user", "content": "How many primes are below 100?"}],
reasoning_effort="medium",
)
message = completion.choices[0].message
print(message.content)
What to know
- Billing. Reasoning tokens are billed at the output rate and reported in
usage.completion_tokens_details.reasoning_tokens. max_tokensincludes reasoning. Leave enough headroom, or the answer may be cut off before it starts.- Effort. Where supported,
reasoning_effort(low,medium,high) trades latency and cost for depth. - Reasoning content. Some models return their reasoning in a
reasoning_contentfield alongsidecontent. It is processed like any other output — not stored. - Keep-alives. While a model reasons, streaming responses may include SSE comment lines; ignore them.