Route the best models.
Own your custom model.
Lightning-fast frontier models at 0% markup, through one OpenAI-compatible API.
Custom models that replace GPT-5.6 Sol workflows at 10× lower cost.
Change the base URL. Keep the SDK you know.
from openai import OpenAI
client = OpenAI(
api_key="$ROUTER_API_KEY",
base_url="https://api.thinwire.ai/v1",
)
r = client.chat.completions.create(
model="grok-4.6",
messages=[{"role": "user", "content": "Hi"}],
stream=True,
)OpenAI-compatible · SSE streaming · one hop to the provider
01
The router
Route the newest frontier models. Lightning-fast, radically cheap.
Run more agent loops for every dollar.
Pay once, get matched Grok usage on top — up to 5× total value and up to 80% lower effective cost at full use. Base tokens always meter at the provider-list rate, 0% markup.
5×
max usage value
80%
max effective savings
Builder
2×
Self-serve after verification
Pay $100 → up to $200
5-day window · up to 50% off at full use
Startup
3×
Best for active agent teams
Pay $500 → up to $1,500
10-day window · up to 66.7% off at full use
Scale
4×
Invoice only · Workload verification
Pay $5,000 → up to $20,000
14-day window · up to 75% off at full use
Partner
5×
Limited slots · Invoice only
Pay $20,000 → up to $100,000
21-day window · up to 80% off at full use
Start with Grok. More on the way.
A model appears here exactly when it is callable, priced, and passing production checks.
grok-4.6Newest
Available- Input
- $2.50 / 1M tokens
- Output
- $10.00 / 1M tokens
- Max output
- 32,768 tokens
Streaming Chat Completions · agent-ready throughput
Checked 2026-08-17 00:00 UTC
grok-4.3
Available- Input
- $1.25 / 1M tokens
- Output
- $2.50 / 1M tokens
- Max output
- 32,768 tokens
Streaming Chat Completions · agent-ready throughput
Checked 2026-08-17 00:00 UTC
smart routing
Coming soon- Cost-optimal routing
- —
- Latency-optimal routing
- —
- Reliability failover
- —
Unlocks at two measured providers
openaiGPT-5.6 family
Coming soongoogleGemini family
Coming soonanthropicClaude family
Coming soonRegistry verified 2026-08-17 00:00 UTC · GET /v1/models returns callable IDs only
02
The platform
One platform to train, run, and improve the model you own.
Replace your expensive GPT-5.6 Sol workflow with a model you own. Same quality on your tasks, proven on evals — at a tenth of the token price.
10×
cheaper than GPT-5.6 Sol
Step 1
Train
Qwen3.8 27B, post-trained with RL on your workflow data — it learns your job, not the internet's.
Training
run E1090 · verified
Mean NLL by step · qwen3.8-27b · LoRA r16 · 3 epochs
Step 2
Prove parity
Held-out evals against GPT-5.6 Sol on your real tasks. It ships only when it meets or beats it.
Evals
beats GPT-5.6 Sol
qwen3.8-27b · tuned
72.5%
gpt-5.6-sol
71.8%
qwen3.8-27b · base
46.0%
Held-out eval · 200 rows / arm · 95% CI
Step 3
Run 10× cheaper
Served on our managed stack at $0.40/$3 per 1M tokens — versus Sol's $5/$30.
Inference
serving · 78% traffic
TTFT p50
140ms
Decode
75tok/s
KV-cache hit
94.9%
vs Sol tokens
10×cheaper
qwen3.8-27b · managed serving · last 24h