Grok 4.6 launch pricingUp to 80% off
AI model routing

Spend more token, Pay less money

Start with discounted Grok today. Use the same endpoint as more models go live.

For specialized workloads, we run custom evals, train models, and manage serving.

Change the base URL. Keep your SDK.

from openai import OpenAI

client = OpenAI(
    api_key="$THINWIRE_API_KEY",
    base_url="https://api.thinwire.ai/v1",
)

response = client.chat.completions.create(
    model="grok-4.6",
    messages=[{"role": "user", "content": "Hello"}],
)

Live request path

Example request

Your SDK
Thinwire
grok-4.6
Receipt

Input

1,284

Output

342

Status

Settled

OpenAI compatible API0% token markupUsage receiptsCustom model evals

01

Model routing

Start with Grok. Add more models through the same API.

Change one URL0% token markupPer-request receipts
Grok launch offer

Buy Grok credits. Get more usage.

Grok tokens meter at the provider list rate with 0% markup. Launch packages add promotional usage for a limited time.

Max usage value

80%

Max savings

Builder

2×

Buy in the console

Pay $100

$200 total Grok usage

$100 promotional · 5 days · up to 50% off when fully used

Startup

3×

Buy in the console

Pay $500

$1,500 total Grok usage

$1,000 promotional · 10 days · up to 66.7% off when fully used

Scale

4×

Sales assisted

Pay $5,000

$20,000 total Grok usage

$15,000 promotional · 14 days · up to 75% off when fully used

Partner

5×

Sales assisted

Pay $20,000

$100,000 total Grok usage

$80,000 promotional · 21 days · up to 80% off when fully used

Promotional usage eligibility: grok-4.6, grok-4.5. Fixed package windows are shown on each card; all promotional usage remains subject to the campaign cutoff. Paid credits remain separate.

Models

Grok is live.

Use grok-4.6 or grok-4.5 now. New providers appear after validation.

grok-4.6Newest

Available
Input
$2.00 / 1M tokens
Cached input
$0.50 / 1M tokens
Output
$6.00 / 1M tokens
Max output
32,768 tokens

Chat Completions, streaming

Checked 2026-08-18 00:00 UTC

grok-4.5

Available
Input
$2.00 / 1M tokens
Cached input
$0.30 / 1M tokens
Output
$6.00 / 1M tokens
Max output
32,768 tokens

Chat Completions, streaming

Checked 2026-08-18 00:00 UTC

smart routing

Coming soon
Cost-optimal routing
Planned
Latency-optimal routing
Planned
Reliability failover
Planned

Available after two providers are measured

openaiGPT-5.6 family

Coming soon

googleGemini family

Coming soon

anthropicClaude family

Coming soon

Beyond text

Grok Imagine Image 2.0Grok Imagine Video 1.5Grok VoiceText-to-Speech
Full catalog in the console

Registry verified 2026-08-18 00:00 UTC. GET /v1/models returns callable IDs only

02

Custom models and evals

Test models on your work. Build one when the evidence supports it.

Workload evalsCustom trainingManaged servingProduction rollout

We test candidate models on your workflow, then train and serve a custom model when the results justify it.

10×

cheaper than GPT-5.6 Sol

Step 1

Train

Qwen3.8 27B, post-trained with RL on your workflow data. It learns your job, not the internet's.

Training

run E1090 · verified

1.51.20.906001,176

Mean NLL by step · qwen3.8-27b · LoRA r16 · 3 epochs

Step 2

Prove parity

Held-out evals against GPT-5.6 Sol on your real tasks. It ships only when the results support it.

Evals

beats GPT-5.6 Sol

qwen3.8-27b · tuned

72.5%

gpt-5.6-sol

71.8%

qwen3.8-27b · base

46.0%

Held-out eval · 200 rows / arm · 95% CI

Step 3

Run 10× cheaper

Served on our managed stack at $0.40/$3 per 1M tokens. Sol is $5/$30.

Inference

serving · 78% traffic

TTFT p50

140ms

Decode

75tok/s

KV-cache hit

94.9%

vs Sol tokens

10×cheaper

qwen3.8-27b · managed serving · last 24h

Step 4

Switch without the leap

Shadow traffic proves the model is operationally sound. Endpoint A/B tests the result with the same URL and keys. No client release is required to ramp or roll back.

Rollout

A/B live · 5% variant

live traffic

one endpoint

95%control

gpt-5.6-sol

5%variant

qwen3.8-27b · tuned

Ramp

5 → 25 → 100%

one call per step

Rollback

one call

traffic snaps to control

Client changes

zero

no flags, no redeploys

Same endpoint, API, and keys · shadow mirrors traffic with responses discarded

FAQ

Plain answers.

Which models are available right now?
grok-4.6 and grok-4.5 are live today. The console lists every callable model. We add a provider only after its pricing, billing, and compatibility checks pass.
What is the difference between paid credits and promotional usage?
Paid credits are your USD balance. Promotional usage applies only to grok-4.6, grok-4.5, has no cash value, and expires at the published UTC cutoff.
When does my promotion expire?
Each package shows its fixed usage window when one applies. Partner has no fixed tier window and remains subject to the campaign cutoff. Expiry removes unused promotional usage, not paid credit.
What fees and taxes apply?
Tokens meter at the provider list rate with 0% markup. Payment fees and taxes appear before checkout.
What do you retain, and what happens in an outage?
We do not log prompt or completion content on the routing path. Receipts store request ID, model, tokens, policy, timing, and debit. If accounting is unavailable, the request does not dispatch.
How does the custom model track work?
We build a test set from your workflow and compare candidate models. If customization has a clear benefit, we train, evaluate, and serve the model for you. Book an eval to scope the work.
How do we switch to the custom model without breaking production?
We shadow the candidate, then run an endpoint A/B test. The URL and API keys stay the same. You can ramp traffic or return every request to the control model without a client release.

Buy Grok credits today. Scope a custom model next.