Pre-launchAPI access is by invitation. Prices shown are estimates.Request access →
Contact sales
ModelsPricingDedicatedDocsEnterpriseContact sales
Inference

Docs / Inference

Chat Completions

Generate responses with the OpenAI-compatible Chat Completions endpoint.

View .md

POST /v1/chat/completions is the main inference endpoint. It follows the OpenAI Chat Completions format, so existing OpenAI clients work by changing only the base URL, the key, and the model id.

from openai import OpenAI

client = OpenAI(base_url="https://api.kurrens.ai/v1", api_key="<KURRENS_API_KEY>")

completion = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Explain KV caching in two sentences."},
    ],
    temperature=0.6,
    max_tokens=512,
)
print(completion.choices[0].message.content)
print(completion.usage)
import OpenAI from 'openai';

const client = new OpenAI({ baseURL: 'https://api.kurrens.ai/v1', apiKey: process.env.KURRENS_API_KEY });

const completion = await client.chat.completions.create({
  model: 'Qwen/Qwen3.8-27B',
  messages: [
    { role: 'system', content: 'You are a concise assistant.' },
    { role: 'user', content: 'Explain KV caching in two sentences.' },
  ],
  temperature: 0.6,
  max_tokens: 512,
});
console.log(completion.choices[0].message.content, completion.usage);
curl https://api.kurrens.ai/v1/chat/completions \
  -H "Authorization: Bearer $KURRENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.8-27B",
    "messages": [
      {"role": "system", "content": "You are a concise assistant."},
      {"role": "user", "content": "Explain KV caching in two sentences."}
    ],
    "temperature": 0.6,
    "max_tokens": 512
  }'

Messages

Role Purpose
system / developer Instructions that shape the model’s behavior
user Your input — text, or text plus images for vision models
assistant Earlier model turns, including tool_calls
tool Results of a tool call, with tool_call_id

Common parameters

Parameter Notes
model Required. A model id from the Model catalog
max_tokens Upper bound on generated tokens, including reasoning tokens
temperature, top_p Sampling controls
stop Up to 4 stop sequences
stream Stream tokens as server-sent events — see Streaming
tools, tool_choice Tool calling
response_format Structured outputs
seed Best-effort reproducibility
service_tier priority or flex where offered

Supported parameters and their ranges vary by model; the full per-model list is published in models.json. Unsupported parameters return 400.

Usage

Every response includes a usage object with prompt_tokens, completion_tokens, and — where applicable — cached and reasoning token counts. This is what you’re billed for.

Data handling

The request and response are processed in memory and discarded when the response completes. See Zero data retention.

    Type to search titles, headings, and page text.

    ↑↓ to move · ↵ to open