OpenAI compatible API · Attested · Public status

Z.ai: GLM 5.2 Performance

Compare measured TTFT, TTFB, throughput, uptime, and route health for Z.ai: GLM 5.2 across TrustedRouter providers using metadata-only production probes.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

z-ai/glm-5.2

open weights Performance

All models

AI IQ IQ 120 #16 public AI IQ rank for glm-5.2
View AI IQ profile

Measured performance

Continuously sampled p50/p95 time-to-first-token (TTFT), time-to-first-byte (TTFB), effective throughput, and success rate for Z.ai: GLM 5.2. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime, and no prompt or output content is stored.

Providerp50 TTFTp95 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
friendli 1173 ms 1587 ms 1173 ms 169 tok/s n=1 100.00% 7
wafer 1218 ms 8411 ms 1218 ms 100.00% 14
together 1662 ms 2782 ms 1662 ms 100.00% 2
engy 1734 ms 3411 ms 1733 ms 36 tok/s n=1 100.00% 24
deepinfra 1777 ms 27546 ms 1777 ms 34 tok/s n=1 100.00% 2
telnyx 1797 ms 2722 ms 1797 ms 157 tok/s n=1 100.00% 4
siliconflow 1820 ms 2917 ms 1820 ms 31 tok/s n=1 100.00% 3
makora 1862 ms 3720 ms 1862 ms 100.00% 10
phala 2244 ms 4392 ms 2244 ms 32 tok/s n=1 100.00% 23
fireworks 2266 ms 3516 ms 2266 ms 36 tok/s n=1 100.00% 7
digitalocean 2558 ms 3861 ms 2557 ms 69 tok/s n=1 100.00% 5
morph 2836 ms 4318 ms 2835 ms 28 tok/s n=1 100.00% 11
zero-g 2881 ms 5285 ms 2880 ms 50 tok/s n=1 100.00% 20
baseten 2918 ms 4034 ms 2917 ms 49 tok/s n=1 100.00% 7
novita 3029 ms 3029 ms 3028 ms 22 tok/s n=1 100.00% 1
venice 3101 ms 3666 ms 3101 ms 55 tok/s n=1 100.00% 3
zai 3444 ms 14580 ms 2829 ms 50 tok/s n=1 100.00% 8
atlas-cloud 3453 ms 3453 ms 3453 ms 100.00% 1
gmi 3469 ms 4535 ms 3468 ms 37 tok/s n=1 100.00% 13
crusoe 3908 ms 4417 ms 3908 ms 100.00% 4
inceptron 4519 ms 7975 ms 4519 ms 100.00% 11
parasail 3597 ms 7445 ms 1327 ms 29 tok/s n=1 97.50% 40
tinfoil 5771 ms 21440 ms 5771 ms 88.89% 9
chutes 1438 ms 10685 ms 1438 ms 33.33% 6
alibaba 31 tok/s n=1 0

Full provider & model leaderboard.

Provider diversity

45 routes.

More routes give the auto router more room to fail over around provider 429 and 5xx responses.

Streaming

Gateway overhead is measured separately.

Public status separates TLS/health overhead from full model latency so slow LLMs do not inflate the router metric.

Status

Metadata rollups.

Status samples store latency, outcome, provider, model, route, cost, and region metadata only.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.