OpenAI compatible API · Attested · Public status
Cloudflare Workers AI performance
Review measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Cloudflare Workers AI on TrustedRouter using metadata-only production probes.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
Cloudflare Workers AIcloudflare-workers-ai
34 samplesContinuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 1441 ms |
|---|---|
| p95 TTFT | 4344 ms |
| p50 TTFB | 1413 ms |
| Effective throughput | 42 tok/s n=3 |
| Uptime | 100.00% |
Measured model routes
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| meta-llama/llama-4-scout-17b-16e-instruct | 598 ms | 598 ms | — | 100.00% | — | 1 |
| meta-llama/llama-3.1-8b-instruct-fp8 | 713 ms | 713 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-120b | 766 ms | 766 ms | 42 tok/s n=2 | 100.00% | — | 2 |
| meta-llama/llama-3.3-70b-instruct-fp8-fast | 1247 ms | 1247 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-20b | 1305 ms | 1305 ms | — | 100.00% | — | 5 |
| meta-llama/llama-3.2-3b-instruct | 1402 ms | 1402 ms | — | 100.00% | — | 3 |
| ibm-granite/granite-4.0-h-micro | 1407 ms | 1406 ms | — | 100.00% | — | 3 |
| nvidia/nemotron-3-120b-a12b | 1441 ms | 1441 ms | — | 100.00% | — | 3 |
| qwen/qwen3-30b-a3b-fp8 | 1466 ms | 1465 ms | — | 100.00% | — | 1 |
| google/gemma-4-26b-a4b-it | 1576 ms | 1576 ms | — | 100.00% | 2 probe_config_error |
5 |
| qwen/qwen2.5-coder-32b-instruct | 2448 ms | 2447 ms | — | 100.00% | — | 5 |
| moonshotai/kimi-k3 | 2923 ms | 2923 ms | 25 tok/s n=1 | 100.00% | — | 2 |
| z-ai/glm-4.7-flash | 3110 ms | 3110 ms | — | 100.00% | — | 1 |
| mistralai/mistral-small-3.1-24b-instruct | 7361 ms | 7361 ms | — | 100.00% | — | 1 |