China's best models. One wire.
Frontier-class open weights — DeepSeek V4, GLM-5.3 / 5.2 / 5.1, Kimi K3, MiniMax M3, Qwen 3.8 — served through one endpoint at China-domestic prices, in dollars. Native Anthropic / OpenAI / Gemini protocols. Zero code changes.
Eight frontier models,
one catalog.
Every headline open-weight release from China's top labs, unified behind one API — priced at what they cost at home.
USD per 1M tokens, cache-miss input / output. GLM-5.3 served at GLM-5.2 tier pricing. * promotional tier as listed Aug 2026.
Close the gap.
Then some.
Independently evaluated, not vendor folklore. On Vals AI's uniform bash-only harness, this catalog holds #2, #5, #7 and #9 of the world's top ten on SWE-bench Verified — GLM-5.3 at 95.4 passes Claude Fable 5 outright, with no vendor number in sight.
SWE-bench Verified — the whole frontier
Vals AI · independent · Aug 2026Percent of 500 real GitHub issues resolved — every model on the same bash-only agent harness. Gradient bars = open-weight models served here: #2, #5, #7 and #9 of the world's top ten. All 8 catalog models present.
Terminal-Bench 2.1
vendor-reportedOurs hold #2–#4: GLM-5.3 debuts 0.6 behind Sol; K3 misses #1 by 0.5 — at one-fifth of Sol's output price.
Terminal-Bench 3.0
Z.ai evaluationThe newest, hardest terminal bench. GLM-5.3 = open-weight SOTA, 6× its predecessor's score.
| Benchmark | GLM-5.3 | Kimi K3 | DeepSeek V4 Pro | GLM-5.2 | Qwen 3.8 | Claude | OpenAI | Source |
|---|---|---|---|---|---|---|---|---|
| SWE-bench Verified | 95.4 | 93.4 | 96.4 | 82.8 | 85.6 | Opus 5 · 97.0 | Sol · 96.2 | Vals AI · independent |
| Terminal-Bench 2.1 | 88.2 | 88.3 | — | 81.0 | 86.6 | Opus 4.8 · 84.6 | Sol · 88.8 | vendor |
| Terminal-Bench 3.0 | 28.3 | 17.4 | — | 4.6 | — | Fable 5 · 33.7 | Sol · 34.6 | Z.ai |
| DeepSWE v1.1 | 66.9 | 67.5 | — | 46.2 | — | Fable 5 · 69.7 | Sol · 72.7 | Z.ai |
| SWE-Marathon v1.1 | 42.5 | 48.1 | — | 19.4 | — | — | — | Z.ai |
| FrontierSWE | 78.1 | — | — | 67.5 | — | Fable 5 · 88.2 | — | Z.ai / Proximal |
| Toolathlon Verified | 73.0 | 76.5 | 74.1 | 59.9 | — | Fable 5 · 74.7 | Sol · 74.9 | Z.ai |
| NL2Repo | 58.0 | — | 61.1 | 48.9 | — | — | — | Z.ai |
| SWE-bench Pro | — | — | — | 62.1 | 67.7 | Opus 4.8 · 69.2 | Sol · 64.6 | vendor |
| AA Intelligence Index | — | 57 (#4) | — | — | — | Opus 5 · #1 overall | — | Artificial Analysis |
Green column = GLM-5.3, the newest catalog model — independently scored 95.4 on SWE-bench Verified. Bold gold = best open-weight model per row. DSV4 Flash / GLM-5.1 / MiniMax M3: see the SWE-bench chart above. — = not published.
That's the entire gap between DeepSeek V4 Pro and Claude Opus 5 on SWE-bench Verified — independently measured by Vals AI on a uniform harness. The capability gap is a rounding error; the invoice gap is 29× on output tokens. And Kimi K3 doesn't chase closed models — at 93.4, it passes Claude Opus 4.8 outright.
Sources — Vals AI SWE-bench Verified leaderboard (independent evaluation, Aug 12 2026); Artificial Analysis Intelligence Index; Frontend Code Arena (community blind votes); vendor technical reports: Z.ai (GLM-5.2 / GLM-5.3, DeepSWE v1.1), Moonshot (Kimi K3), Alibaba (Qwen 3.8), OpenAI (GPT-5.6), Google (Gemini 3.1 Pro), Anthropic. ~ derived from Vals AI per-difficulty resolution data. † vendor-reported. Closed-model comparators are each lab's latest flagship on the named bench.
Frontier performance.
Without the frontier invoice.
Western frontier APIs price like scarcity. Chinese labs price like competition — and we pass that pricing straight through, converted at source.
| Model | Input $/1M | Output $/1M | Context | Weights |
|---|
What $100 buys you
Prices as published by each vendor, Aug 2026; converted from CNY domestic list where applicable. Claude Opus 5 $5/$25 · GPT-5.6 Sol $5/$30 · Claude Fable 5 $10/$50 per official pricing pages. Every priced catalog model shown; GLM-5.1 is tiered, priced on request. Benchmark figures: Vals AI (SWE-bench Verified), vendor technical reports, Artificial Analysis. DeepSeek V4 Pro at promotional tier.
Routed for latency,
not geography.
Your users' requests enter at the edge nearest them, ride dedicated backbone capacity into Chinese inference regions, and stream back before anyone notices the distance. Three backbone paths, probed every second — when one degrades, traffic shifts by itself.
Nearest-edge ingress
Users enter at the closest healthy edge — one hop, no tromboning across oceans — then get steered to the fastest inference path in real time.
Lowest-latency path, always
Anycast entry plus dedicated transpacific and Eurasian backbone capacity keep round-trips at internet-physics minimum — the fastest path wins every time.
Self-healing failover
Every route is probed around the clock. A degraded backbone is drained and its traffic re-routed in seconds — multiple upstreams per model, no ticket, no waiting.
Encrypted in transit.
Nothing at rest.
No sign-up form, no email, no OAuth — an API key is your entire identity here. Requests ride TLS from your process to the inference region, and the only thing we keep is what billing needs: token counts. No prompt or completion bodies are ever written to disk — nothing to leak, nothing to sell.
TLS 1.2/1.3 on every hop
Your process to our gateway, our gateway to the inference region — payloads cross every network we control as ciphertext. No plaintext hop, no exceptions.
Zero content retention
Prompts and completions are streamed through, never written to disk. Billing reads token counts, timestamps and model IDs — that is the whole log line.
No accounts, no trackers
Keys are handed out manually — no sign-up form, no email database, no OAuth. This site serves every font and script from its own origin: zero third-party requests.
GDPR / CCPA-aligned by design
We collect no personal data, so there is nothing to mishandle, nothing to export, nothing to breach. Compliance by architecture — not by a policy PDF nobody reads.
Data practices align with GDPR and CCPA data-minimization principles. No personal data is collected, stored, or processed; traffic metadata is limited to billing counters. We claim no certifications — there is nothing to certify.
Pay per token.
Or lock in a plan.
Two ways to pay, one balance across every model. Switch between them whenever your workload changes.
Per-token, to the penny
No minimums, no monthly fee, no seat pricing. Metered per request with minute-level detail.
- Exact metering — input, output and cached tokens billed separately
- Cache discounts — repeated prompts cost up to 90% less on supported models
- One balance — top up once, spend across all eight models
Commit volume, pay less
Pre-purchase large token blocks at tiered discounts. Built for steady production workloads.
- Tiered pricing — bigger blocks, steeper rates
- Credits don't expire — use them on your schedule
- Any model — a plan covers the whole catalog, not one model
Early access feedback.
From builders in our private beta.
Quotes from private-beta participants, shared with permission. Handles withheld by request.
Your stack already works.
Three native protocol surfaces on one endpoint. Point your existing SDK, agent or harness at a new base URL — that's the whole migration.
Get your key.
Keys are provisioned personally during beta — tell us about your workload and you're usually running within a day.