Set stream: true to receive tokens as server-sent events (SSE) instead of waiting for the full response.
stream = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V4.1-Flash",
messages=[{"role": "user", "content": "Write a haiku about rivers."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print("\n", chunk.usage)
What to expect
- Each event is a
data:line containing a chat completion chunk; the stream ends withdata: [DONE]. - With
stream_options.include_usage, the final chunk carries theusageobject. Kurrens always reports usage for streamed requests. - Reasoning models may send SSE comment lines (
: keep-alive) while they think. Clients should ignore them — they keep the connection from timing out. - If an error occurs mid-stream, a final event contains an
errorobject before the stream closes.
Cancelling
Closing the connection cancels generation. You’re billed only for tokens generated before cancellation.