Pre-launchAPI access is by invitation. Prices shown are estimates.Request access →
Contact sales
ModelsPricingDedicatedDocsEnterpriseContact sales
Getting started

Docs

Kurrens documentation

Fast, private inference for open models through one OpenAI-compatible API.

View .md

Kurrens is an inference provider for open-weight models. We run models such as DeepSeek, Kimi, GLM, Qwen, and MiniMax on GPUs we operate in Malaysia, Indonesia, and Ireland, and serve them through one OpenAI-compatible API, with zero data retention: prompts and completions are processed in memory and never stored or used for training.

What you can do

Serverless inference

Call open models per token with the OpenAI SDK you already use. Streaming, tool calling, structured outputs, and prompt caching where the model supports them.

Dedicated endpoints

Reserve GPUs for a single model and get a private base URL with consistent latency and custom rate limits.

Zero data retention

Inputs and outputs live in memory for the length of the request. We keep only request metadata for billing.

Declared precision

Every model lists its Hugging Face weights, quantization, context window, and price.

Start here

Base URL

https://api.kurrens.ai/v1

Any client that accepts an OpenAI base URL works with Kurrens.

    Type to search titles, headings, and page text.

    ↑↓ to move · ↵ to open