Grok 4.6 is liveup to 80% off
Cut your AI bill · Keep your quality

Route the best models.
Own your custom model.

Lightning-fast frontier models at 0% markup, through one OpenAI-compatible API.

Custom models that replace GPT-5.6 Sol workflows at 10× lower cost.

0% token markupTPM & rate limits sized to your workloadLow-latency streaming pathPer-request usage receiptsHuman onboarding

01

The router

Route the newest frontier models. Lightning-fast, radically cheap.

Easy integration — change one URL0% token markupLow-latency streaming
Launch pricing

Run more agent loops for every dollar.

Pay once, get matched Grok usage on top — up to 5× total value and up to 80% lower effective cost at full use. Base tokens always meter at the provider-list rate, 0% markup.

max usage value

80%

max effective savings

Builder

2×

Self-serve after verification

Pay $100 → up to $200

5-day window · up to 50% off at full use

Startup

3×

Best for active agent teams

Pay $500 → up to $1,500

10-day window · up to 66.7% off at full use

Scale

4×

Invoice only · Workload verification

Pay $5,000 → up to $20,000

14-day window · up to 75% off at full use

Partner

5×

Limited slots · Invoice only

Pay $20,000 → up to $100,000

21-day window · up to 80% off at full use

Availability

Start with Grok. More on the way.

A model appears here exactly when it is callable, priced, and passing production checks.

grok-4.6Newest

Available
Input
$2.50 / 1M tokens
Output
$10.00 / 1M tokens
Max output
32,768 tokens

Streaming Chat Completions · agent-ready throughput

Checked 2026-08-17 00:00 UTC

grok-4.3

Available
Input
$1.25 / 1M tokens
Output
$2.50 / 1M tokens
Max output
32,768 tokens

Streaming Chat Completions · agent-ready throughput

Checked 2026-08-17 00:00 UTC

smart routing

Coming soon
Cost-optimal routing
Latency-optimal routing
Reliability failover

Unlocks at two measured providers

openaiGPT-5.6 family

Coming soon

googleGemini family

Coming soon

anthropicClaude family

Coming soon

Registry verified 2026-08-17 00:00 UTC · GET /v1/models returns callable IDs only

02

The platform

One platform to train, run, and improve the model you own.

Qwen3.8 27B · custom RLParity with GPT-5.6 Sol, proven on evals10× cheaper tokens

Replace your expensive GPT-5.6 Sol workflow with a model you own. Same quality on your tasks, proven on evals — at a tenth of the token price.

10×

cheaper than GPT-5.6 Sol

Step 1

Train

Qwen3.8 27B, post-trained with RL on your workflow data — it learns your job, not the internet's.

Training

run E1090 · verified

1.51.20.906001,176

Mean NLL by step · qwen3.8-27b · LoRA r16 · 3 epochs

Step 2

Prove parity

Held-out evals against GPT-5.6 Sol on your real tasks. It ships only when it meets or beats it.

Evals

beats GPT-5.6 Sol

qwen3.8-27b · tuned

72.5%

gpt-5.6-sol

71.8%

qwen3.8-27b · base

46.0%

Held-out eval · 200 rows / arm · 95% CI

Step 3

Run 10× cheaper

Served on our managed stack at $0.40/$3 per 1M tokens — versus Sol's $5/$30.

Inference

serving · 78% traffic

TTFT p50

140ms

Decode

75tok/s

KV-cache hit

94.9%

vs Sol tokens

10×cheaper

qwen3.8-27b · managed serving · last 24h

FAQ

Plain answers.

Which models are available right now?
Exactly the model IDs returned by GET /v1/models — today that is grok-4.6 and grok-4.3. A model is never selectable before it passes production canary, billing, compatibility, and capacity checks. OpenAI, Google, and Anthropic routes are coming soon; smart routing unlocks once two measured providers are live.
What is the difference between paid credits and promotional usage?
Paid credits are a USD balance that works across every model marked Available, with their own validity and refund policy. Promotional usage is a launch match restricted to eligible Grok models: it settles paired with paid usage, has no cash value, cannot be spent alone, and expires at the published UTC cutoff.
When does my promotion expire?
Each tier has a fixed window — Builder 5 days, Startup 10, Scale 14, Partner 21 — with the exact UTC cutoff shown before you pay. Expiry cancels unused promotional entitlement only; it never touches your paid credits.
What fees and taxes apply?
Tokens meter at the provider-list base rate with 0% markup. ACH top-ups carry a 0% launch fee; card processing is disclosed at payment. Your first payment starts a 90-day founding fee lock; afterwards a 3% managed-credit fee applies to new top-ups only (waived for ACH commitments of $5k+ and enterprise contracts). Taxes are always shown separately.
What do you retain, and what happens in an outage?
We do not log prompt or completion content on the router path — receipts store request ID, model, tokens, policy, timing, and the paid/promo debit. And we fail closed: if the ledger or settlement authority is unavailable, requests are rejected before dispatch, and an ambiguous request holds its reservation until reconciliation settles it by request ID. No silent charges.
How does the custom model track work?
It starts with a two-week eval sprint: we build a task set from your real workflow and baseline candidate models against it. If the evidence supports it, we post-train an open model like Qwen3.8 27B (SFT/RL) on your data, prove it meets or beats GPT-5.6 Sol on held-out evals, and serve it on our managed stack. On list prices that is about 10× cheaper per token — GPT-5.6 Sol is $5/$30 per 1M tokens, Qwen3.8 27B serves at $0.40/$3. Separate SOW and data permissions; independent of your API balance.

Route the best models today. Own yours tomorrow.