Kurrens is an inference provider for open-weight models. We run models such as DeepSeek, Kimi, GLM, Qwen, and MiniMax on GPUs we operate in Malaysia, Indonesia, and Ireland, and serve them through one OpenAI-compatible API, with zero data retention: prompts and completions are processed in memory and never stored or used for training.
What you can do
Serverless inference
Call open models per token with the OpenAI SDK you already use. Streaming, tool calling, structured outputs, and prompt caching where the model supports them.
Dedicated endpoints
Reserve GPUs for a single model and get a private base URL with consistent latency and custom rate limits.
Zero data retention
Inputs and outputs live in memory for the length of the request. We keep only request metadata for billing.
Declared precision
Every model lists its Hugging Face weights, quantization, context window, and price.
Start here
Base URL
https://api.kurrens.ai/v1
Any client that accepts an OpenAI base URL works with Kurrens.