A dedicated endpoint runs one model on GPUs reserved for you. It uses the same API as serverless inference, on a private base URL.
When to use it
- You need consistent latency at high, steady throughput.
- Your traffic exceeds serverless rate limits.
- You want a model we don’t serve on serverless, or a specific precision.
What you get
| Serverless | Dedicated | |
|---|---|---|
| Capacity | Shared, fair-use limits | Reserved GPUs, single tenant |
| Rate limits | Per-account defaults | Sized to your traffic |
| Billing | Per token | Per GPU-hour |
| API | https://api.kurrens.ai/v1 |
Private base URL, same API |
| Data retention | None | None |
Getting started
Dedicated endpoints are provisioned with our team. Tell us about your workload — model, expected tokens per day or peak requests per second, latency target, and region — Johor, Jakarta, or Dublin (see Regions) — and we’ll size a deployment. A dedicated endpoint stays in the region you choose.
Using it is the same as serverless — only the base URL changes:
client = OpenAI(base_url="https://<your-endpoint>.kurrens.ai/v1", api_key="<KURRENS_API_KEY>")