Simple, per-token pricing.
Pay only for what you use. No subscriptions, no minimums, no platform fees.
EstimatesEstimated — may change before launch. Final pricing will be published at launch. Request access
Serverless
| Model | Context | Input | Cached input | Output |
|---|---|---|---|---|
| gpt-oss-20b | 128K | $0.035 | $0.025 | $0.15 |
| gpt-oss-120b | 128K | $0.12 | $0.065 | $0.60 |
| Qwen3.6 35B A3B | 256K | $0.13 | $0.050 | $1.00 |
| DeepSeek V4 Flash 0731 | 1.3M | $0.14 | $0.025 | $0.37 |
| MiMo-V2.6-Flash | 1M | $0.14 | $0.003 | $0.28 |
| Gemma 4 31B | 256K | $0.14 | $0.095 | $0.40 |
| GLM 5.3 Flash | 1.3M | $0.15 | $0.030 | $0.50 |
| Qwen3.8 Flash | 1M | $0.15 | $0.016 | $0.47 |
| Qwen3 Coder Next | 256K | $0.19 | $0.053 | $1.20 |
| Qwen3.8 27B | 1M | $0.21 | $0.050 | $2.52 |
| DeepSeek V4.1 Flash | 1M | $0.22 | $0.006 | $0.84 |
| MiniMax M3 | 1M | $0.30 | $0.060 | $1.20 |
| MiniMax M2.7 | 200K | $0.30 | $0.060 | $1.20 |
| MiMo-V2.6-Pro | 1M | $0.43 | $0.004 | $0.87 |
| Kimi K2.6 | 256K | $0.75 | $0.15 | $3.50 |
| Kimi K2.7 Code | 256K | $0.91 | $0.18 | $3.84 |
| GLM 5.3 | 1.3M | $1.19 | $0.19 | $4.40 |
| DeepSeek V4 Pro 0813 | 1M | $1.31 | $0.044 | $3.96 |
| GLM 5.2 | 1M | $1.40 | $0.21 | $4.40 |
| Qwen3.8 2.4T A95B | 1M | $2.00 | $0.25 | $6.00 |
| Kimi K3 | 1M | $3.00 | $0.30 | $15.00 |
Estimated prices in USD per 1M tokens; may change before launch. Taxes may apply depending on your location.
Service tiers
| Tier | Scheduling | Price |
|---|---|---|
| Standard | Default scheduling | 1× |
| Priority | Scheduled ahead of Standard during peak demand | 1.5× |
| Flex | Lower cost for non-urgent traffic; may be slower or briefly unavailable | 0.8× |
Set "service_tier": "priority" or "flex" on any request. Availability varies by model.
Dedicated
| GPU | Memory | Price |
|---|---|---|
| NVIDIA H100 | 80 GB | Contact us |
| NVIDIA H200 | 141 GB | Contact us |
| NVIDIA B200 | 180 GB | Contact us |
Billed per minute. Minimum commitment may apply.
Enterprise
Volume discounts · Monthly invoicing · Signed DPA · SLA · Dedicated support · Custom rate limits
FAQ
How does billing work?
Add prepaid credits to your account. Usage is deducted in real time. You can enable auto-recharge so you never run out.
Do credits expire?
Purchased credits expire 12 months after purchase. We remind you by email before a material amount expires.
Can I get a refund?
Unused purchased credits can be refunded within 14 days of purchase. Billing errors and failed requests on our side are always refunded or credited.
How are cached tokens billed?
When a prompt prefix is reused, those input tokens are billed at the cached rate.
Are reasoning tokens billed?
Yes, reasoning tokens are billed at the output rate.
Do you offer invoicing?
Yes, for Enterprise customers with monthly invoicing terms.
Are there rate limits?
Yes. Default limits scale with your account history. Contact us for higher limits or dedicated capacity.
Are taxes included?
Prices exclude taxes. Applicable taxes such as Singapore GST are added at checkout where required.
See the Billing, Credits & Refund Policy for full terms.