GLOBAL EDGE → CHINA FRONTIER INFERENCE

China's best models. One wire.

Frontier-class open weights — DeepSeek V4, GLM-5.3 / 5.2 / 5.1, Kimi K3, MiniMax M3, Qwen 3.8 — served through one endpoint at China-domestic prices, in dollars. Native Anthropic / OpenAI / Gemini protocols. Zero code changes.

36× cheapervs GPT-5.6 input
8 modelsfrontier open weights
3 protocolsnative, drop-in
1 endpointapi.tokyoyonars.com
The Lineup

Eight frontier models,
one catalog.

Every headline open-weight release from China's top labs, unified behind one API — priced at what they cost at home.

USD per 1M tokens, cache-miss input / output. GLM-5.3 served at GLM-5.2 tier pricing. * promotional tier as listed Aug 2026.

Benchmarks

Close the gap.
Then some.

Independently evaluated, not vendor folklore. On Vals AI's uniform bash-only harness, this catalog holds #2, #5, #7 and #9 of the world's top ten on SWE-bench Verified — GLM-5.3 at 95.4 passes Claude Fable 5 outright, with no vendor number in sight.

SWE-bench Verified — the whole frontier

Vals AI · independent · Aug 2026
Claude Opus 5 · closed
97.0
DeepSeek V4 Pro · ours
96.4
GPT-5.6 Sol · closed
96.2
Grok 4.6 · closed
95.6
GLM-5.3 · ours · new
95.4
Claude Fable 5 · closed
95.0
Kimi K3 · ours
93.4
GPT-5.6 Luna · closed
93.0
DeepSeek V4 Flash · ours
88.8
Claude Opus 4.8 · closed
88.6
Grok 4.5 · closed
86.6
Muse Spark 1.2 · closed
86.6
Opus 4.8 · Claude Code harness
85.8
Qwen 3.8 Max · ours
85.6
GLM-5.2 · ours
82.8
GPT 5.5 · closed
82.6
Gemini 3.1 Pro · closed
80.6†
GLM-5.1 · ours
76.6~
MiniMax M3 · ours
75.3~

Percent of 500 real GitHub issues resolved — every model on the same bash-only agent harness. Gradient bars = open-weight models served here: #2, #5, #7 and #9 of the world's top ten. All 8 catalog models present.

Terminal-Bench 2.1

vendor-reported
GPT-5.6 Sol · closed
88.8
Kimi K3 · ours
88.3
GLM-5.3 · ours · new
88.2
Qwen 3.8 Max · ours
86.6
Claude Opus 4.8 · closed
84.6
GLM-5.2 · ours
81.0

Ours hold #2–#4: GLM-5.3 debuts 0.6 behind Sol; K3 misses #1 by 0.5 — at one-fifth of Sol's output price.

Terminal-Bench 3.0

Z.ai evaluation
GPT-5.6 Sol · closed
34.6
Claude Fable 5 · closed
33.7
GLM-5.3 · ours · new
28.3
Opus 4.8 · closed
21.1
Kimi K3 · ours
17.4
GLM-5.2 · ours
4.6

The newest, hardest terminal bench. GLM-5.3 = open-weight SOTA, 6× its predecessor's score.

BenchmarkGLM-5.3Kimi K3DeepSeek V4 ProGLM-5.2Qwen 3.8ClaudeOpenAISource
SWE-bench Verified95.493.496.482.885.6Opus 5 · 97.0Sol · 96.2Vals AI · independent
Terminal-Bench 2.188.288.381.086.6Opus 4.8 · 84.6Sol · 88.8vendor
Terminal-Bench 3.028.317.44.6Fable 5 · 33.7Sol · 34.6Z.ai
DeepSWE v1.166.967.546.2Fable 5 · 69.7Sol · 72.7Z.ai
SWE-Marathon v1.142.548.119.4Z.ai
FrontierSWE78.167.5Fable 5 · 88.2Z.ai / Proximal
Toolathlon Verified73.076.574.159.9Fable 5 · 74.7Sol · 74.9Z.ai
NL2Repo58.061.148.9Z.ai
SWE-bench Pro62.167.7Opus 4.8 · 69.2Sol · 64.6vendor
AA Intelligence Index57 (#4)Opus 5 · #1 overallArtificial Analysis

Green column = GLM-5.3, the newest catalog model — independently scored 95.4 on SWE-bench Verified. Bold gold = best open-weight model per row. DSV4 Flash / GLM-5.1 / MiniMax M3: see the SWE-bench chart above. — = not published.

Beyond the charts — three more independent boards
Kimi K3 — #1 on Frontend Code Arena · a human blind-vote arena for real UI builds: its code was picked over Claude Fable 5's in 76% of head-to-heads
DeepSeek V4 Pro — top open model on LiveCodeBench · competitive-programming board refreshed with brand-new problems, so scores can't come from memorized training data
Kimi K3 — 67.5 on DeepSWE v1.1, GLM-5.3 right behind at 66.9 · long-horizon agentic work: real GitHub issues resolved end-to-end, unattended
0.6 pts

That's the entire gap between DeepSeek V4 Pro and Claude Opus 5 on SWE-bench Verified — independently measured by Vals AI on a uniform harness. The capability gap is a rounding error; the invoice gap is 29× on output tokens. And Kimi K3 doesn't chase closed models — at 93.4, it passes Claude Opus 4.8 outright.

Sources — Vals AI SWE-bench Verified leaderboard (independent evaluation, Aug 12 2026); Artificial Analysis Intelligence Index; Frontend Code Arena (community blind votes); vendor technical reports: Z.ai (GLM-5.2 / GLM-5.3, DeepSWE v1.1), Moonshot (Kimi K3), Alibaba (Qwen 3.8), OpenAI (GPT-5.6), Google (Gemini 3.1 Pro), Anthropic. ~ derived from Vals AI per-difficulty resolution data. † vendor-reported. Closed-model comparators are each lab's latest flagship on the named bench.

The Ledger

Frontier performance.
Without the frontier invoice.

Western frontier APIs price like scarcity. Chinese labs price like competition — and we pass that pricing straight through, converted at source.

5.7×
Cheaper output
GLM-5.2 vs Claude Opus 5 ($4.40 vs $25)
36×
Cheaper input
DeepSeek V4 Flash vs GPT-5.6 Sol ($0.14 vs $5)
107×
Cheaper output
DeepSeek V4 Flash vs GPT-5.6 Sol ($0.28 vs $30)
ModelInput $/1MOutput $/1MContextWeights

What $100 buys you

OUTPUT TOKENS PURCHASED WITH A FLAT $100 — HIGHER IS BETTER
DeepSeek V4 Flash
357M
DeepSeek V4 Pro
114.9M
MiniMax M3
83.3M
GLM-5.3
22.7M
GLM-5.2
22.7M
Qwen 3.8
16.7M
Kimi K3
6.7M
Claude Opus 5
4.0M
GPT-5.6 Sol
3.3M
Claude Fable 5
2.0M

Prices as published by each vendor, Aug 2026; converted from CNY domestic list where applicable. Claude Opus 5 $5/$25 · GPT-5.6 Sol $5/$30 · Claude Fable 5 $10/$50 per official pricing pages. Every priced catalog model shown; GLM-5.1 is tiered, priced on request. Benchmark figures: Vals AI (SWE-bench Verified), vendor technical reports, Artificial Analysis. DeepSeek V4 Pro at promotional tier.

The Wire

Routed for latency,
not geography.

Your users' requests enter at the edge nearest them, ride dedicated backbone capacity into Chinese inference regions, and stream back before anyone notices the distance. Three backbone paths, probed every second — when one degrades, traffic shifts by itself.

Nearest-edge ingress

Users enter at the closest healthy edge — one hop, no tromboning across oceans — then get steered to the fastest inference path in real time.

1 hopTO THE EDGE

Lowest-latency path, always

Anycast entry plus dedicated transpacific and Eurasian backbone capacity keep round-trips at internet-physics minimum — the fastest path wins every time.

43–164MS GLOBAL RTT

Self-healing failover

Every route is probed around the clock. A degraded backbone is drained and its traffic re-routed in seconds — multiple upstreams per model, no ticket, no waiting.

99.99%UPTIME
All systems operational
ALL PATHS NOMINAL — 3/3 BACKBONES UP · 0 DROPPED
user cityedge POPbackboneinference region
Privacy

Encrypted in transit.
Nothing at rest.

No sign-up form, no email, no OAuth — an API key is your entire identity here. Requests ride TLS from your process to the inference region, and the only thing we keep is what billing needs: token counts. No prompt or completion bodies are ever written to disk — nothing to leak, nothing to sell.

TLS 1.2/1.3 on every hop

Your process to our gateway, our gateway to the inference region — payloads cross every network we control as ciphertext. No plaintext hop, no exceptions.

Zero content retention

Prompts and completions are streamed through, never written to disk. Billing reads token counts, timestamps and model IDs — that is the whole log line.

No accounts, no trackers

Keys are handed out manually — no sign-up form, no email database, no OAuth. This site serves every font and script from its own origin: zero third-party requests.

GDPR / CCPA-aligned by design

We collect no personal data, so there is nothing to mishandle, nothing to export, nothing to breach. Compliance by architecture — not by a policy PDF nobody reads.

Data practices align with GDPR and CCPA data-minimization principles. No personal data is collected, stored, or processed; traffic metadata is limited to billing counters. We claim no certifications — there is nothing to certify.

CUSTOMER FILE · S-4471ZERO PII
NAMEJordan Miller
EMAILjordan@acme.dev
COMPANYAcme GmbH
LOCATIONBerlin · DE
PAYMENTVISA ·· 4242
DEVICEmacOS · Chrome
API KEY sk-suh-···7f2a
MODEL glm-5.3
TOKENS 1,204,118
THE ENTIRE RECORD A REQUEST LEAVES BEHIND
Billing

Pay per token.
Or lock in a plan.

Two ways to pay, one balance across every model. Switch between them whenever your workload changes.

Pay as you go

Per-token, to the penny

No minimums, no monthly fee, no seat pricing. Metered per request with minute-level detail.

  • Exact metering — input, output and cached tokens billed separately
  • Cache discounts — repeated prompts cost up to 90% less on supported models
  • One balance — top up once, spend across all eight models
Token plans

Commit volume, pay less

Pre-purchase large token blocks at tiered discounts. Built for steady production workloads.

  • Tiered pricing — bigger blocks, steeper rates
  • Credits don't expire — use them on your schedule
  • Any model — a plan covers the whole catalog, not one model
Voices

Early access feedback.

From builders in our private beta.

Quotes from private-beta participants, shared with permission. Handles withheld by request.

Connect

Your stack already works.

Three native protocol surfaces on one endpoint. Point your existing SDK, agent or harness at a new base URL — that's the whole migration.

api.tokyoyonars.com · chat completions

    
Onboarding

Get your key.

Keys are provisioned personally during beta — tell us about your workload and you're usually running within a day.