OpenAI compatible API · Attested · Public status

Z.ai: GLM 5.2 Performance

Compare measured TTFT, TTFB, throughput, uptime, and route health for Z.ai: GLM 5.2 across TrustedRouter providers using metadata-only production probes.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

z-ai/glm-5.2

open weights Performance

All models

AI IQ IQ 120 #16 public AI IQ rank for glm-5.2
View AI IQ profile

Measured performance

Continuously sampled p50/p95 time-to-first-token (TTFT), time-to-first-byte (TTFB), effective throughput, and success rate for Z.ai: GLM 5.2. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime, and no prompt or output content is stored.

Providerp50 TTFTp95 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
wafer 842 ms 6793 ms 842 ms 100.00% 12
friendli 1173 ms 1587 ms 1173 ms 169 tok/s n=1 100.00% 5
digitalocean 1593 ms 3861 ms 1593 ms 69 tok/s n=1 100.00% 6
together 1662 ms 9245 ms 1662 ms 100.00% 4
telnyx 1797 ms 2722 ms 1797 ms 157 tok/s n=1 100.00% 4
siliconflow 1820 ms 2917 ms 1820 ms 31 tok/s n=1 100.00% 3
makora 1862 ms 3720 ms 1862 ms 100.00% 12
engy 1922 ms 3411 ms 1921 ms 36 tok/s n=1 100.00% 25
phala 2185 ms 4562 ms 2184 ms 32 tok/s n=1 100.00% 18
fireworks 2259 ms 3516 ms 2259 ms 36 tok/s n=1 100.00% 6
morph 2836 ms 4318 ms 2835 ms 28 tok/s n=1 100.00% 11
zero-g 2881 ms 5285 ms 2880 ms 50 tok/s n=1 100.00% 20
baseten 2918 ms 4034 ms 2917 ms 49 tok/s n=1 100.00% 7
venice 3101 ms 3666 ms 3101 ms 55 tok/s n=1 100.00% 3
zai 3444 ms 14580 ms 2829 ms 50 tok/s n=1 100.00% 8
atlas-cloud 3453 ms 3453 ms 3453 ms 100.00% 1
gmi 3469 ms 4535 ms 3468 ms 37 tok/s n=1 100.00% 12
crusoe 3908 ms 4417 ms 3908 ms 100.00% 4
inceptron 4057 ms 7975 ms 4056 ms 100.00% 12
deepinfra 27546 ms 27546 ms 27546 ms 34 tok/s n=1 100.00% 1
parasail 4152 ms 8054 ms 1327 ms 29 tok/s n=1 97.44% 39
tinfoil 5771 ms 21440 ms 5771 ms 90.00% 10
chutes 1438 ms 2303 ms 1438 ms 33.33% 6
alibaba 31 tok/s n=1 0
novita 22 tok/s n=1 0

Full provider & model leaderboard.

Provider diversity

45 routes.

More routes give the auto router more room to fail over around provider 429 and 5xx responses.

Streaming

Gateway overhead is measured separately.

Public status separates TLS/health overhead from full model latency so slow LLMs do not inflate the router metric.

Status

Metadata rollups.

Status samples store latency, outcome, provider, model, route, cost, and region metadata only.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.