Spend more token, Pay less money
Start with discounted Grok today. Use the same endpoint as more models go live.
For specialized workloads, we run custom evals, train models, and manage serving.
Change the base URL. Keep your SDK.
from openai import OpenAI
client = OpenAI(
api_key="$THINWIRE_API_KEY",
base_url="https://api.thinwire.ai/v1",
)
response = client.chat.completions.create(
model="grok-4.6",
messages=[{"role": "user", "content": "Hello"}],
)Live request path
Example request
Input
1,284
Output
342
Status
Settled
01
Model routing
Start with Grok. Add more models through the same API.
Buy Grok credits. Get more usage.
Grok tokens meter at the provider list rate with 0% markup. Launch packages add promotional usage for a limited time.
5×
Max usage value
80%
Max savings
Builder
2×
Buy in the console
Pay $100
$200 total Grok usage
$100 promotional · 5 days · up to 50% off when fully used
Startup
3×
Buy in the console
Pay $500
$1,500 total Grok usage
$1,000 promotional · 10 days · up to 66.7% off when fully used
Scale
4×
Sales assisted
Pay $5,000
$20,000 total Grok usage
$15,000 promotional · 14 days · up to 75% off when fully used
Partner
5×
Sales assisted
Pay $20,000
$100,000 total Grok usage
$80,000 promotional · 21 days · up to 80% off when fully used
Promotional usage eligibility: grok-4.6, grok-4.5. Fixed package windows are shown on each card; all promotional usage remains subject to the campaign cutoff. Paid credits remain separate.
Grok is live.
Use grok-4.6 or grok-4.5 now. New providers appear after validation.
grok-4.6Newest
Available- Input
- $2.00 / 1M tokens
- Cached input
- $0.50 / 1M tokens
- Output
- $6.00 / 1M tokens
- Max output
- 32,768 tokens
Chat Completions, streaming
Checked 2026-08-18 00:00 UTC
grok-4.5
Available- Input
- $2.00 / 1M tokens
- Cached input
- $0.30 / 1M tokens
- Output
- $6.00 / 1M tokens
- Max output
- 32,768 tokens
Chat Completions, streaming
Checked 2026-08-18 00:00 UTC
smart routing
Coming soon- Cost-optimal routing
- Planned
- Latency-optimal routing
- Planned
- Reliability failover
- Planned
Available after two providers are measured
openaiGPT-5.6 family
Coming soongoogleGemini family
Coming soonanthropicClaude family
Coming soonBeyond text
Grok Imagine Image 2.0Grok Imagine Video 1.5Grok VoiceText-to-SpeechRegistry verified 2026-08-18 00:00 UTC. GET /v1/models returns callable IDs only
02
Custom models and evals
Test models on your work. Build one when the evidence supports it.
We test candidate models on your workflow, then train and serve a custom model when the results justify it.
10×
cheaper than GPT-5.6 Sol
Step 1
Train
Qwen3.8 27B, post-trained with RL on your workflow data. It learns your job, not the internet's.
Training
run E1090 · verified
Mean NLL by step · qwen3.8-27b · LoRA r16 · 3 epochs
Step 2
Prove parity
Held-out evals against GPT-5.6 Sol on your real tasks. It ships only when the results support it.
Evals
beats GPT-5.6 Sol
qwen3.8-27b · tuned
72.5%
gpt-5.6-sol
71.8%
qwen3.8-27b · base
46.0%
Held-out eval · 200 rows / arm · 95% CI
Step 3
Run 10× cheaper
Served on our managed stack at $0.40/$3 per 1M tokens. Sol is $5/$30.
Inference
serving · 78% traffic
TTFT p50
140ms
Decode
75tok/s
KV-cache hit
94.9%
vs Sol tokens
10×cheaper
qwen3.8-27b · managed serving · last 24h
Step 4
Switch without the leap
Shadow traffic proves the model is operationally sound. Endpoint A/B tests the result with the same URL and keys. No client release is required to ramp or roll back.
Rollout
A/B live · 5% variant
live traffic
one endpoint
95%control
gpt-5.6-sol
5%variant
qwen3.8-27b · tuned
Ramp
5 → 25 → 100%
one call per step
Rollback
one call
traffic snaps to control
Client changes
zero
no flags, no redeploys
Same endpoint, API, and keys · shadow mirrors traffic with responses discarded