Z.ai: GLM 5.2 Performance
Compare measured TTFT, TTFB, throughput, uptime, and route health for Z.ai: GLM 5.2 across TrustedRouter providers using metadata-only production probes.
z-ai/glm-5.2
Measured performance
Continuously sampled p50/p95 time-to-first-token (TTFT), time-to-first-byte (TTFB), effective throughput, and success rate for Z.ai: GLM 5.2. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime, and no prompt or output content is stored.
| Provider | p50 TTFT | p95 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|---|
| wafer | 842 ms | 6793 ms | 842 ms | — | 100.00% | — | 12 |
| friendli | 1173 ms | 1587 ms | 1173 ms | 169 tok/s n=1 | 100.00% | — | 5 |
| digitalocean | 1593 ms | 3861 ms | 1593 ms | 69 tok/s n=1 | 100.00% | — | 6 |
| together | 1662 ms | 9245 ms | 1662 ms | — | 100.00% | — | 4 |
| telnyx | 1797 ms | 2722 ms | 1797 ms | 157 tok/s n=1 | 100.00% | — | 4 |
| siliconflow | 1820 ms | 2917 ms | 1820 ms | 31 tok/s n=1 | 100.00% | — | 3 |
| makora | 1862 ms | 3720 ms | 1862 ms | — | 100.00% | — | 12 |
| engy | 1922 ms | 3411 ms | 1921 ms | 36 tok/s n=1 | 100.00% | — | 25 |
| phala | 2185 ms | 4562 ms | 2184 ms | 32 tok/s n=1 | 100.00% | — | 18 |
| fireworks | 2259 ms | 3516 ms | 2259 ms | 36 tok/s n=1 | 100.00% | — | 6 |
| morph | 2836 ms | 4318 ms | 2835 ms | 28 tok/s n=1 | 100.00% | — | 11 |
| zero-g | 2881 ms | 5285 ms | 2880 ms | 50 tok/s n=1 | 100.00% | — | 20 |
| baseten | 2918 ms | 4034 ms | 2917 ms | 49 tok/s n=1 | 100.00% | — | 7 |
| venice | 3101 ms | 3666 ms | 3101 ms | 55 tok/s n=1 | 100.00% | — | 3 |
| zai | 3444 ms | 14580 ms | 2829 ms | 50 tok/s n=1 | 100.00% | — | 8 |
| atlas-cloud | 3453 ms | 3453 ms | 3453 ms | — | 100.00% | — | 1 |
| gmi | 3469 ms | 4535 ms | 3468 ms | 37 tok/s n=1 | 100.00% | — | 12 |
| crusoe | 3908 ms | 4417 ms | 3908 ms | — | 100.00% | — | 4 |
| inceptron | 4057 ms | 7975 ms | 4056 ms | — | 100.00% | — | 12 |
| deepinfra | 27546 ms | 27546 ms | 27546 ms | 34 tok/s n=1 | 100.00% | — | 1 |
| parasail | 4152 ms | 8054 ms | 1327 ms | 29 tok/s n=1 | 97.44% | — | 39 |
| tinfoil | 5771 ms | 21440 ms | 5771 ms | — | 90.00% | — | 10 |
| chutes | 1438 ms | 2303 ms | 1438 ms | — | 33.33% | — | 6 |
| alibaba | — | — | — | 31 tok/s n=1 | — | — | 0 |
| novita | — | — | — | 22 tok/s n=1 | — | — | 0 |
Full provider & model leaderboard.
45 routes.
More routes give the auto router more room to fail over around provider 429 and 5xx responses.
Gateway overhead is measured separately.
Public status separates TLS/health overhead from full model latency so slow LLMs do not inflate the router metric.
Metadata rollups.
Status samples store latency, outcome, provider, model, route, cost, and region metadata only.
View public status or inspect provider routes.