# Chat Completions

> Generate responses with the OpenAI-compatible Chat Completions endpoint.

<PreLaunch />

`POST /v1/chat/completions` is the main inference endpoint. It follows the OpenAI Chat Completions format, so existing
OpenAI clients work by changing only the base URL, the key, and the model id.

<Tabs syncKey="lang">
  <TabItem label="Python">
    ```python
    from openai import OpenAI

    client = OpenAI(base_url="https://api.kurrens.ai/v1", api_key="<KURRENS_API_KEY>")

    completion = client.chat.completions.create(
        model="Qwen/Qwen3.8-27B",
        messages=[
            {"role": "system", "content": "You are a concise assistant."},
            {"role": "user", "content": "Explain KV caching in two sentences."},
        ],
        temperature=0.6,
        max_tokens=512,
    )
    print(completion.choices[0].message.content)
    print(completion.usage)
    ```
  </TabItem>
  <TabItem label="TypeScript">
    ```ts
    import OpenAI from 'openai';

    const client = new OpenAI({ baseURL: 'https://api.kurrens.ai/v1', apiKey: process.env.KURRENS_API_KEY });

    const completion = await client.chat.completions.create({
      model: 'Qwen/Qwen3.8-27B',
      messages: [
        { role: 'system', content: 'You are a concise assistant.' },
        { role: 'user', content: 'Explain KV caching in two sentences.' },
      ],
      temperature: 0.6,
      max_tokens: 512,
    });
    console.log(completion.choices[0].message.content, completion.usage);
    ```
  </TabItem>
  <TabItem label="cURL">
    ```bash
    curl https://api.kurrens.ai/v1/chat/completions \
      -H "Authorization: Bearer $KURRENS_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "Qwen/Qwen3.8-27B",
        "messages": [
          {"role": "system", "content": "You are a concise assistant."},
          {"role": "user", "content": "Explain KV caching in two sentences."}
        ],
        "temperature": 0.6,
        "max_tokens": 512
      }'
    ```
  </TabItem>
</Tabs>

## Messages

| Role | Purpose |
|---|---|
| `system` / `developer` | Instructions that shape the model's behavior |
| `user` | Your input — text, or text plus images for [vision](/docs/inference/vision) models |
| `assistant` | Earlier model turns, including `tool_calls` |
| `tool` | Results of a tool call, with `tool_call_id` |

## Common parameters

| Parameter | Notes |
|---|---|
| `model` | Required. A model id from the [Model catalog](/docs/getting-started/models) |
| `max_tokens` | Upper bound on generated tokens, including reasoning tokens |
| `temperature`, `top_p` | Sampling controls |
| `stop` | Up to 4 stop sequences |
| `stream` | Stream tokens as server-sent events — see [Streaming](/docs/inference/streaming) |
| `tools`, `tool_choice` | [Tool calling](/docs/inference/tool-calling) |
| `response_format` | [Structured outputs](/docs/inference/structured-outputs) |
| `seed` | Best-effort reproducibility |
| `service_tier` | `priority` or `flex` where offered |

Supported parameters and their ranges vary by model; the full per-model list is published in
[models.json](/models.json). Unsupported parameters return `400`.

## Usage

Every response includes a `usage` object with `prompt_tokens`, `completion_tokens`, and — where applicable —
cached and reasoning token counts. This is what you're billed for.

## Data handling

The request and response are processed in memory and discarded when the response completes. See
[Zero data retention](/docs/data-security/zero-data-retention).
